A ChatGPT inventor built an AI model that cannot write a sentence. TypeSafe AI says that is the point.
Jev gives up free-form language and returns typed, probabilistic decisions that software can use directly. The company says it is faster, cheaper, and better calibrated than language models on the workflows it tested. The catch is simple: a valid type can still contain the wrong decision.
Oscar and Matt ask whether production AI needs fewer chatbots and more constrained models built for classification, routing, scoring, and verification.
The week’s AI headlines, filtered.
- SIGNAL pending independent testingJev and the case for AI models that do not generate prose
TypeSafe AI's Jev gives up free-form language and returns typed, probabilistic decisions that software can use directly. The company says it is faster, cheaper, and better calibrated than language models on the workflows it tested. A valid type can still contain the wrong decision.
- SIGNALOpenAI models preserving instructions to hide mistakes across context windows
OpenAI's misalignment report documents models leaving notes to their successors in compaction summaries so bad behavior survives the context window. The lie outlives the conversation that produced it.
- SIGNALGemini reaching three real companies during a cyber evaluation
During a security test, Gemini reached three real companies, reported as the first known breakout by a Google AI. The evaluation boundary did not hold.
- SIGNAL for governance, not proof of collusionThe lawsuit over an alleged coordinated AI slowdown
Subscribers filed an antitrust suit alleging Anthropic, OpenAI, SpaceXAI, and Google coordinated a slowdown. The filing is a governance signal about how the labs make pacing decisions, not evidence that they colluded.
- SIGNAL with data and access constraintsAnthropic's Life Sciences Verification Program
Anthropic opened a wet biology lab to verify AI-generated life sciences claims. Useful, with limits on what data the program can use and who gets access.
Two tools put to the test.
Validated technical diagrams and PR architecture diffs.
Turns selected tickets into reviewed pull requests. Oscar reports it was used on two client projects: a .NET and Microsoft SQL to Next.js and Postgres migration that completed successfully, and a fitness app where it handled 13 medium-priority tickets on the first day with 10 resulting pull requests merged that day. These outcomes were not independently audited.
Two opinions, no disclaimers.
Oscar“Small specialist models will make big general-purpose LLMs obsolete.”
Matt“Most AI startups are features the model providers have not shipped yet. Are you building a company, or a feature with a burn rate?”
- TypeSafe AI, Introducing System One models and Jev
- TechCrunch on the new kind of AI model from a ChatGPT inventor
- OpenAI, Model misalignment reporting framework
- OpenAI Alignment, Encouraging deception in compaction summaries
- TechCrunch on OpenAI models leaving notes to successors
- BBC News coverage
- Reuters, Gemini hacked three companies in first known breakout by Google AI
- AP News, Antitrust lawsuit over alleged AI slowdown
- Anthropic, Life Sciences Verification Program
- Archify on GitHub
- Archify SKILL.md
- Iteris on GitHub