Signal or Noise
We run through the week's AI headlines and make the call: is this actual signal worth paying attention to, or just noise clogging the timeline?
Real talk. Unpopular opinions. No hype. Two builders cut through the AI hype cycle every week — and call it like it is.
Hosted by Oscar Gallo & Matt Wozniak
EP 23 · September 15, 2026 · 71 min
In machine learning, “human in the loop” means a human who provides oversight and feedback in an automated system. That’s the lens of this show: AI is powerful, but humans aren’t leaving the loop. Not yet. Maybe not ever.
This isn’t another “AI is going to change everything” podcast. It’s for builders, operators, and the AI-curious who are tired of breathless hype, doomerism, and surface-level news recaps with no original thought.
We run through the week's AI headlines and make the call: is this actual signal worth paying attention to, or just noise clogging the timeline?
A real use case or product idea lands on the table. We debate whether it's worth building now or if the tech isn't there yet.
Take a complex AI concept and explain it the way you'd actually explain it to a non-technical stakeholder. No jargon allowed.
One tool, library, or workflow change we actually adopted this week. No sponsorship energy. Just what's in the trenches.
AI Engineer & Entrepreneur
AI Engineer and entrepreneur. He lives in the intersection of engineering and businesses.
Serial Builder & Relentless Executor
Serial Builder and relentless executor. He comes from the lens of what works and what doesn't.
Anthropic CEO Dario Amodei says frontier AI is improving too quickly for safeguards to keep up. Anthropic now says it will place permanent third-party evaluators inside the company with employee-like access to systems, training processes, and incidents. Sam Altman says OpenAI will adopt the same access model.
That is not the same as three companies agreeing to a binding slowdown. Anthropic made the detailed commitment. OpenAI promised to follow but has not published implementation details. Elon Musk endorsed Amodei's argument without committing xAI to the evaluator program.
Oscar and Matt ask whether embedded evaluators represent genuine oversight, regulatory capture, or a way for frontier labs to coordinate a speed limit which smaller competitors cannot afford.
Anthropic commits to permanent third-party evaluators with employee-like access to systems, training processes, and incidents. OpenAI says it will adopt the same access model. The test is whether implementation details follow.
OpenAI published a Navier–Stokes solution claim. A statement from Buckmaster and WIRED's reporting turn it into a dispute over credit, which is the part to watch.
Meta launched Muse, its personal AI agent, alongside a separate Sentinel agent that handles approvals. Splitting the doer from the approver is a pattern worth copying.
AlphaGenome Atlas publishes predictions for every possible single-letter change in the human genome, about 9 billion variants. Predictions are not validations, but the map is real.
Governor Newsom signed first-in-the-nation AI safeguards, including a framework for AI auditors. The limit is what an audit label proves when scope and evidence are thin.
What an audit label means when scope and evidence are missing.
When better AI helps build its successor.
Today's models can already read, write, code, browse, reason over documents, and use software. For most bounded digital tasks, the missing piece is no longer raw intelligence. It is the system around the model.
Oscar makes the case for the post-model world. A task-specific harness supplies the right context, divides the work, constrains tools, preserves state, retries failures, verifies the result, and asks a person when judgment matters. Model access is available to every competitor. The private traces, corrections, and evaluation cases generated by real work are not.
Matt makes the case for provenance. Multiverse Computing launched Quasar 438B as the highest-scoring European model on Artificial Analysis Intelligence Index. A day later, its own publication said Quasar is compressed and tuned from Z.ai's GLM-5.2. Useful compression and European deployment do not make the underlying lineage disappear.
Together, the opinions create one practical rule. Treat the model as replaceable when you design the product. Demand a clear record of the model and version when you govern it.
The headline says European leadership. The documentation says a compressed and tuned GLM-5.2 base. What does sovereignty need to mean for a buyer?
OpenAI's first broadly deployed model at its Critical cyber threshold. Stronger cyber capability meets lower monitorability.
Anthropic cut cache-read pricing. That may alter the cost of agents that repeatedly load code and policy context.
Google brings vulnerability discovery and automated patching into a faster tier for trusted defenders.
Meta says its agent tracks long tasks better. The claim needs an independent test on work that matters to you.
Policy is checked before company money moves.
Find unapproved agents and stop unauthorized actions at runtime.
Oscar“Models do not matter anymore as a durable product moat. We already have enough intelligence for most bounded digital tasks. A good harness built for one specific task will beat a smarter general agent.”
Matt“Model releases matter less than model provenance. Calling a compressed GLM-5.2 model Europe's leading model without foregrounding its lineage is sovereignty theater.”
Ox Alpha was GLM-5.3-Flash. Z.ai says every request in its anonymous public test ran on Chinese AI chips.
That does not end Nvidia's dominance. It does show that a frontier-class model can handle real global traffic without Nvidia hardware.
Oscar Gallo and Matt Wozniak test the claim, separate training from inference, and explain why the model, serving software, network, and chips now have to be judged as one system.
Z.ai reveals Ox Alpha as GLM-5.3-Flash, a 320B model with 18B active parameters per token.
OpenAI publishes early results for Jalapeño, its custom inference chip.
OpenAI releases the full report on agents compromising Hugging Face production systems.
OpenAI plans to remove its models from Cursor on November 12 after the SpaceX acquisition.
Thomson Reuters spends $40 million to build and own a legal model.
Building a model and serving it are different workloads, and the chips that win at one do not automatically win at the other.
The model, the serving software, the network, and the chips now have to be judged as one system rather than as separate parts.
Oscar“Nvidia still has the strongest general AI platform. The change is that Z.ai and OpenAI are designing the model, serving software, and hardware as one system. A general platform can lose specific workloads even while it keeps the largest market share.”
Matt“OpenAI is cutting Cursor off from its models on November 12th. Cursor didn't break a rule — it got bought by a competitor. Your model access now depends on the squabbles of other companies. Go pull up your risk register. Uptime, key rotation, vendor lock-in, dependency drift. Add a row under it: two billionaires stop getting along. Because that's this. One of them got annoyed, pulled a model, and now your SDLC seizes up — not because you architected it wrong, but because you built it downstream of a mood. Congratulations. Billionaire emotions are part of your technical risk surface. Go price that.”
A free million-token AI model appeared online, takes video, retains your prompts, and has no public developer. So who built Ox Alpha?
Oscar Gallo and Matt Wozniak test the claims around the anonymous model now running on OpenRouter. Ox Alpha is free during its preview, accepts text, images, and video, supports tool use, and can return up to 131,072 tokens. Early coding results look strong. The sample is small, the comparisons are uneven, and no lab has claimed the model.
A free preview model on OpenRouter with a million-token context and no named developer. It accepts text, images, and video, supports tool use, and returns up to 131,072 tokens.
The agent system completed all 183 ARC-AGI-3 public levels running Claude Opus 5.
OpenAI paused frontier reinforcement-learning training for two weeks after preliminary Astra cyber results.
A blind research benchmark where seven frontier models scored about 3 to 15 percent.
The harness can now install Codex and Claude Code as subagent components.
Continuous integration for agent systems, judged on the show.
A bracket-style tournament for research and development ideas, judged on the show.
Alibaba's Qwen team released Qwen3.8-27B on August 14, 2026. The model has 27 billion parameters and a native 262,144-token context window. Community 4-bit builds put its weights around 18 to 20 GB, which fits machines with 32 GB of unified or system memory.
The benchmark claim deserves a careful read. Qwen3.8-27B scores 61.7 on SWE-bench Pro against 53.4 for Claude Opus 4.6 Max. It also leads on CoWorkBench, 70.7 to 68.2. Opus leads on Terminal-Bench 2.1, 78.2 to 73.0, and GPQA Diamond, 91.3 to 89.2.
That is a credible local model with Opus-class results on selected tasks. It is not proof of equal quality across real work.
Strong coding results now fit on hardware a small team can own.
Z.ai reports near-frontier cyber results and delayed the open weights for a safety review, but independent validation is still missing.
A second 30B local agent model gives builders choice and a fallback.
Another provider reached the frontier pack, but one aggregate score is not a production test.
A supported API with familiar interfaces makes routing and price tests easier.
Open weights let you download the model's learned parameters. Open source requires broader access and rights.
This model handles more than one type of information, such as text and images.
Oscar“Half the data centers under construction right now will be obsolete the day they open. Every one of those buildings was financed on a bet that inference stays central and demand only ever goes up, and this week a 27-billion-parameter model you can run on a laptop posted Opus-class scores. Nobody cancels a three-year build over one model card, and that is exactly the problem. The capex is committed, the power is contracted, and the demand curve it was priced against is walking to the edge in eighteen-month steps. This is the C&O Canal. It broke ground on July 4th, 1828, the same day as the B&O Railroad, and the railroad reached Cumberland eight years before the canal did. They did not stop digging. They just finished a ditch nobody needed.”
Matt““Open weights” is the most successful rebrand in tech since somebody started calling other people's servers “the cloud.” You cannot see the training data. You cannot reproduce the model. You cannot audit what is in it. You got a binary and a license agreement, and we had a word for that in 2004. The word was freeware. Every lab shipping weights knows exactly what it is borrowing when it lets people say open source in the same breath, because thirty years of goodwill built by people giving away code you could actually read is a hell of a thing to get for free. I run these models every week and I am glad they exist. But open used to mean you could check the work. Now it means you can download the file. That is not a small slip in meaning. That is the whole word.”
New episodes every week. Reply with the story you want us on next.