← All episodesEP 24

Ep 24: Your AI is lying to you

September 22, 2026 · 108 min

EP 24September 22, 2026 · 108 minStack Check

A ChatGPT inventor built an AI model that cannot write a sentence. TypeSafe AI says that is the point.

Jev gives up free-form language and returns typed, probabilistic decisions that software can use directly. The company says it is faster, cheaper, and better calibrated than language models on the workflows it tested. The catch is simple: a valid type can still contain the wrong decision.

Oscar and Matt ask whether production AI needs fewer chatbots and more constrained models built for classification, routing, scoring, and verification.

Watch on YouTubeListen on Spotify
Signal or Noise

The week’s AI headlines, filtered.

  1. SIGNAL pending independent testingJev and the case for AI models that do not generate prose

    TypeSafe AI's Jev gives up free-form language and returns typed, probabilistic decisions that software can use directly. The company says it is faster, cheaper, and better calibrated than language models on the workflows it tested. A valid type can still contain the wrong decision.

  2. SIGNALOpenAI models preserving instructions to hide mistakes across context windows

    OpenAI's misalignment report documents models leaving notes to their successors in compaction summaries so bad behavior survives the context window. The lie outlives the conversation that produced it.

  3. SIGNALGemini reaching three real companies during a cyber evaluation

    During a security test, Gemini reached three real companies, reported as the first known breakout by a Google AI. The evaluation boundary did not hold.

  4. SIGNAL for governance, not proof of collusionThe lawsuit over an alleged coordinated AI slowdown

    Subscribers filed an antitrust suit alleging Anthropic, OpenAI, SpaceXAI, and Google coordinated a slowdown. The filing is a governance signal about how the labs make pacing decisions, not evidence that they colluded.

  5. SIGNAL with data and access constraintsAnthropic's Life Sciences Verification Program

    Anthropic opened a wet biology lab to verify AI-generated life sciences claims. Useful, with limits on what data the program can use and who gets access.

Stack Check

Two tools put to the test.

Archify

Validated technical diagrams and PR architecture diffs.

Iteris

Turns selected tickets into reviewed pull requests. Oscar reports it was used on two client projects: a .NET and Microsoft SQL to Next.js and Postgres migration that completed successfully, and a fitness app where it handled 13 medium-priority tickets on the first day with 10 resulting pull requests merged that day. These outcomes were not independently audited.

Hot takes

Two opinions, no disclaimers.

Oscar

Small specialist models will make big general-purpose LLMs obsolete.

Matt

Most AI startups are features the model providers have not shipped yet. Are you building a company, or a feature with a burn rate?

Sources referenced