Human In The Loop logo
Human In The Loop · Weekly

A human-first AI podcast.

Real talk. Unpopular opinions. No hype. Two builders cut through the AI hype cycle every week — and call it like it is.

Watch on YouTubeListen on SpotifyListen on Apple Podcasts

Hosted by Oscar Gallo & Matt Wozniak

Now playing · Latest episode

Ep 23: AI is going to kill us all

EP 23 · September 15, 2026 · 71 min

Watch on YouTubeListen on SpotifyListen on Apple
About

We’re the humans in the loop.

In machine learning, “human in the loop” means a human who provides oversight and feedback in an automated system. That’s the lens of this show: AI is powerful, but humans aren’t leaving the loop. Not yet. Maybe not ever.

This isn’t another “AI is going to change everything” podcast. It’s for builders, operators, and the AI-curious who are tired of breathless hype, doomerism, and surface-level news recaps with no original thought.

What’s in each episode

Every week: one fixed, one rotating.

FIXED

Signal or Noise

We run through the week's AI headlines and make the call: is this actual signal worth paying attention to, or just noise clogging the timeline?

ROTATING

Ship It or Skip It

A real use case or product idea lands on the table. We debate whether it's worth building now or if the tech isn't there yet.

ROTATING

Explain It to My Client

Take a complex AI concept and explain it the way you'd actually explain it to a non-technical stakeholder. No jargon allowed.

ROTATING

Stack Check

One tool, library, or workflow change we actually adopted this week. No sponsorship energy. Just what's in the trenches.

Hosts
Oscar Gallo

Oscar Gallo

AI Engineer & Entrepreneur

AI Engineer and entrepreneur. He lives in the intersection of engineering and businesses.

Matt Wozniak

Matt Wozniak

Serial Builder & Relentless Executor

Serial Builder and relentless executor. He comes from the lens of what works and what doesn't.

Latest episode

Ep 23: AI is going to kill us all

EP 23September 15, 2026 · 71 minNo Jargon Required

Anthropic CEO Dario Amodei says frontier AI is improving too quickly for safeguards to keep up. Anthropic now says it will place permanent third-party evaluators inside the company with employee-like access to systems, training processes, and incidents. Sam Altman says OpenAI will adopt the same access model.

That is not the same as three companies agreeing to a binding slowdown. Anthropic made the detailed commitment. OpenAI promised to follow but has not published implementation details. Elon Musk endorsed Amodei's argument without committing xAI to the evaluator program.

Oscar and Matt ask whether embedded evaluators represent genuine oversight, regulatory capture, or a way for frontier labs to coordinate a speed limit which smaller competitors cannot afford.

Signal or Noise

The week’s AI headlines, filtered.

  1. SIGNAL with a follow-through testAnthropic and OpenAI back embedded outside evaluators

    Anthropic commits to permanent third-party evaluators with employee-like access to systems, training processes, and incidents. OpenAI says it will adopt the same access model. The test is whether implementation details follow.

  2. SIGNALOpenAI's Navier–Stokes claim and the credit dispute

    OpenAI published a Navier–Stokes solution claim. A statement from Buckmaster and WIRED's reporting turn it into a dispute over credit, which is the part to watch.

  3. SIGNALMeta Muse and its separate Sentinel approval agent

    Meta launched Muse, its personal AI agent, alongside a separate Sentinel agent that handles approvals. Splitting the doer from the approver is a pattern worth copying.

  4. SIGNALDeepMind's predictions for 9 billion DNA variants

    AlphaGenome Atlas publishes predictions for every possible single-letter change in the human genome, about 9 billion variants. Predictions are not validations, but the map is real.

  5. SIGNAL with limitsCalifornia's new framework for AI auditors

    Governor Newsom signed first-in-the-nation AI safeguards, including a framework for AI auditors. The limit is what an audit label proves when scope and evidence are thin.

No Jargon Required

Two concepts behind the oversight debate.

Independent AI audits

What an audit label means when scope and evidence are missing.

Recursive self-improvement

When better AI helps build its successor.

Previously

Ep 22: The post-model world: harnesses, provenance, and Quasar

EP 22September 8, 2026 · 80 minShip It or Skip It

Today's models can already read, write, code, browse, reason over documents, and use software. For most bounded digital tasks, the missing piece is no longer raw intelligence. It is the system around the model.

Oscar makes the case for the post-model world. A task-specific harness supplies the right context, divides the work, constrains tools, preserves state, retries failures, verifies the result, and asks a person when judgment matters. Model access is available to every competitor. The private traces, corrections, and evaluation cases generated by real work are not.

Matt makes the case for provenance. Multiverse Computing launched Quasar 438B as the highest-scoring European model on Artificial Analysis Intelligence Index. A day later, its own publication said Quasar is compressed and tuned from Z.ai's GLM-5.2. Useful compression and European deployment do not make the underlying lineage disappear.

Together, the opinions create one practical rule. Treat the model as replaceable when you design the product. Demand a clear record of the model and version when you govern it.

Signal or Noise

The week’s AI headlines, filtered.

  1. SIGNAL with a caveatQuasar 438B and GLM-5.2

    The headline says European leadership. The documentation says a compressed and tuned GLM-5.2 base. What does sovereignty need to mean for a buyer?

  2. SIGNALGPT-6 Astra

    OpenAI's first broadly deployed model at its Critical cyber threshold. Stronger cyber capability meets lower monitorability.

  3. SIGNALClaude Fable 5.1

    Anthropic cut cache-read pricing. That may alter the cost of agents that repeatedly load code and policy context.

  4. SIGNALGemini 3.8 Flash Cyber

    Google brings vulnerability discovery and automated patching into a faster tier for trusted defenders.

  5. SIGNAL pending independent testingMuse Spark 1.3

    Meta says its agent tracks long tasks better. The claim needs an independent test on work that matters to you.

Ship It or Skip It

Two product ideas for the agent economy.

The AI expense layer

Policy is checked before company money moves.

An agent kill switch

Find unapproved agents and stop unauthorized actions at runtime.

Unpopular opinions

Oscar and Matt make the case.

Oscar

Models do not matter anymore as a durable product moat. We already have enough intelligence for most bounded digital tasks. A good harness built for one specific task will beat a smarter general agent.

Matt

Model releases matter less than model provenance. Calling a compressed GLM-5.2 model Europe's leading model without foregrounding its lineage is sovereignty theater.

Previously

Ep 21: Frontier AI WITHOUT Nvidia, Ox Alpha was Chinese all along

EP 21September 1, 2026 · 85 minNo Jargon Required

Ox Alpha was GLM-5.3-Flash. Z.ai says every request in its anonymous public test ran on Chinese AI chips.

That does not end Nvidia's dominance. It does show that a frontier-class model can handle real global traffic without Nvidia hardware.

Oscar Gallo and Matt Wozniak test the claim, separate training from inference, and explain why the model, serving software, network, and chips now have to be judged as one system.

Signal or Noise

The week’s AI headlines, filtered.

  1. SIGNALGLM-5.3-Flash

    Z.ai reveals Ox Alpha as GLM-5.3-Flash, a 320B model with 18B active parameters per token.

  2. SIGNALOpenAI Jalapeño

    OpenAI publishes early results for Jalapeño, its custom inference chip.

  3. SIGNALHugging Face incident report

    OpenAI releases the full report on agents compromising Hugging Face production systems.

  4. SIGNALOpenAI leaves Cursor

    OpenAI plans to remove its models from Cursor on November 12 after the SpaceX acquisition.

  5. SIGNALThomson Reuters legal model

    Thomson Reuters spends $40 million to build and own a legal model.

No Jargon Required

Two concepts behind the Nvidia-free claim.

Training versus inference

Building a model and serving it are different workloads, and the chips that win at one do not automatically win at the other.

Hardware-software co-design

The model, the serving software, the network, and the chips now have to be judged as one system rather than as separate parts.

Hot takes

Two opinions, no disclaimers.

Oscar

Nvidia still has the strongest general AI platform. The change is that Z.ai and OpenAI are designing the model, serving software, and hardware as one system. A general platform can lose specific workloads even while it keeps the largest market share.

Matt

OpenAI is cutting Cursor off from its models on November 12th. Cursor didn't break a rule — it got bought by a competitor. Your model access now depends on the squabbles of other companies. Go pull up your risk register. Uptime, key rotation, vendor lock-in, dependency drift. Add a row under it: two billionaires stop getting along. Because that's this. One of them got annoyed, pulled a model, and now your SDLC seizes up — not because you architected it wrong, but because you built it downstream of a mood. Congratulations. Billionaire emotions are part of your technical risk surface. Go price that.

Previously

Ep 20: A Mysterious AI Model Just Appeared. Nobody Knows Who Built It.

EP 20August 25, 2026 · 78 minShip It or Skip It

A free million-token AI model appeared online, takes video, retains your prompts, and has no public developer. So who built Ox Alpha?

Oscar Gallo and Matt Wozniak test the claims around the anonymous model now running on OpenRouter. Ox Alpha is free during its preview, accepts text, images, and video, supports tool use, and can return up to 131,072 tokens. Early coding results look strong. The sample is small, the comparisons are uneven, and no lab has claimed the model.

Signal or Noise

The week’s AI headlines, filtered.

  1. SIGNALOx Alpha

    A free preview model on OpenRouter with a million-token context and no named developer. It accepts text, images, and video, supports tool use, and returns up to 131,072 tokens.

  2. SIGNALNVIDIA AVO

    The agent system completed all 183 ARC-AGI-3 public levels running Claude Opus 5.

  3. SIGNALOpenAI pauses Astra

    OpenAI paused frontier reinforcement-learning training for two weeks after preliminary Astra cyber results.

  4. SIGNALReconstruction benchmark

    A blind research benchmark where seven frontier models scored about 3 to 15 percent.

  5. SIGNALDeepSeek Harness

    The harness can now install Codex and Claude Code as subagent components.

Ship It or Skip It

Two builds, shipped or skipped.

Agent System CI

Continuous integration for agent systems, judged on the show.

R&D Idea Tournament

A bracket-style tournament for research and development ideas, judged on the show.

Previously

Ep 19: Qwen3.8 Brings Opus-Class AI to a Laptop

EP 19August 18, 2026No Jargon Required

Alibaba's Qwen team released Qwen3.8-27B on August 14, 2026. The model has 27 billion parameters and a native 262,144-token context window. Community 4-bit builds put its weights around 18 to 20 GB, which fits machines with 32 GB of unified or system memory.

The benchmark claim deserves a careful read. Qwen3.8-27B scores 61.7 on SWE-bench Pro against 53.4 for Claude Opus 4.6 Max. It also leads on CoWorkBench, 70.7 to 68.2. Opus leads on Terminal-Bench 2.1, 78.2 to 73.0, and GPQA Diamond, 91.3 to 89.2.

That is a credible local model with Opus-class results on selected tasks. It is not proof of equal quality across real work.

Signal or Noise

The week’s AI headlines, filtered.

  1. SIGNAL with a caveatQwen3.8-27B

    Strong coding results now fit on hardware a small team can own.

  2. SIGNAL with a caveatGLM-5.3

    Z.ai reports near-frontier cyber results and delayed the open weights for a safety review, but independent validation is still missing.

  3. SIGNALMeta Muse Glimmer

    A second 30B local agent model gives builders choice and a fallback.

  4. SIGNAL pending broader testingGrok 4.6

    Another provider reached the frontier pack, but one aggregate score is not a production test.

  5. SIGNALDeepSeek V4 Pro GA

    A supported API with familiar interfaces makes routing and price tests easier.

No Jargon Required

Two AI terms, explained in plain English.

Open weights versus open source

Open weights let you download the model's learned parameters. Open source requires broader access and rights.

Multimodal model

This model handles more than one type of information, such as text and images.

Closing takes

Oscar and Matt make the case.

Oscar

Half the data centers under construction right now will be obsolete the day they open. Every one of those buildings was financed on a bet that inference stays central and demand only ever goes up, and this week a 27-billion-parameter model you can run on a laptop posted Opus-class scores. Nobody cancels a three-year build over one model card, and that is exactly the problem. The capex is committed, the power is contracted, and the demand curve it was priced against is walking to the edge in eighteen-month steps. This is the C&O Canal. It broke ground on July 4th, 1828, the same day as the B&O Railroad, and the railroad reached Cumberland eight years before the canal did. They did not stop digging. They just finished a ditch nobody needed.

Matt

“Open weights” is the most successful rebrand in tech since somebody started calling other people's servers “the cloud.” You cannot see the training data. You cannot reproduce the model. You cannot audit what is in it. You got a binary and a license agreement, and we had a word for that in 2004. The word was freeware. Every lab shipping weights knows exactly what it is borrowing when it lets people say open source in the same breath, because thirty years of goodwill built by people giving away code you could actually read is a hell of a thing to get for free. I run these models every week and I am glad they exist. But open used to mean you could check the work. Now it means you can download the file. That is not a small slip in meaning. That is the whole word.

New episodes every week. Reply with the story you want us on next.

Subscribe

Pick your platform.

YouTubeWatch weeklySpotifyListen on the goApple PodcastsSubscribe on Apple