ExoBrain

ExoBrain Weekly Newsletter

Astra and a new age of complexity, Nvidia expands its empire, and Claude's last theorem

Welcome to our weekly newsletter, a combination of thematic insights from the founders at ExoBrain, and a broader news roundup from our Exo agents.

This week we look at:

  • Astra and a new age of complexity

    OpenAI's latest model is a significant step forward, but many have concerns about how easy it will be to understand, as it becomes increasingly able to hide its thoughts.

  • Nvidia expands its empire

    Nvidia has agreed to pay nearly $13 billion for Hugging Face, the place where most developers go to find and share AI models. To win approval it has promised, in writing, to keep the platform open to everyone.

  • Claude's last theorem

    Claude rebuilt one of the most famous proofs in mathematics from the ground up in eleven days, checking every single step along the way. The worry is that machines can now check far faster than people can understand.

  • News roundup

    This week: three frontier labs go dark in the same morning and none will say why, DeepSeek plans a vast Huawei cluster to replace Nvidia, the EU brings ChatGPT under its strictest platform rules, and researchers find a model's reasoning trace is less honest than it looks.

Astra and a new age of complexity

OpenAI's latest model is a significant step forward, but many have concerns about how easy it will be to understand, as it becomes increasingly able to hide its thoughts.

Joel Miller

Joel Miller

4 min read
Astra and a new age of complexity

OpenAI launched GPT-6 Astra this week via a gradual roll-out, describing it as the most intelligent and aligned model it has built. Company president Greg Brockman went further and closed the press briefing with "Welcome to the AGI era."

The benchmarks support the view that Astra is a major new entrant. It scored 97.6% on FrontierMath Tier 4, which is remarkable given where this benchmark was just a year ago. Its largest gains come when it must explore an unfamiliar environment, operate software, preserve information and complete a sequence of connected actions. ARC Prize found that Astra developed compact symbolic representations for individual games within the ARC benchmark. In some cases, it created a small domain-specific language to describe objects, states and actions. It then used this representation to plan its next move. The continuous-harness version used fewer actions than the median human tester on 96% of levels, the first model to beat the human baseline.

GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. ... In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels.

François Chollet, ARC Prize

Astra driving Blender on its own: a full architectural scene, modelled, planted and lit, with a 250 frame flythrough queued on the timeline. The same computer-use ability extends to animations and playable games.

This behaviour has exciting economic and scientific applications. Research often requires someone to observe a system, propose an explanation, design an experiment and revise the explanation using the result. Astra appears able to perform parts of that cycle in simulated environments. OpenAI researcher Noam Brown believes users may discover mathematical and scientific applications that the company has not yet tested.

Despite the headline results, the launch has received less attention than GPT-5 did last year. GPT-5 followed a long period in which the GPT-4 family defined the frontier for many users. Astra enters a market in which Anthropic has made progress in software engineering and enterprise use, Google remains technically competitive and Chinese laboratories regularly release strong models. A new GPT release no longer commands the same automatic attention.

Astra will likely prove to be the more consequential release. Its powerful capabilities arrive with evidence that the relationship between the model’s internal computation, written reasoning and external behaviour is becoming harder to follow.

OpenAI reports a substantial reduction in "chain-of-thought" monitorability. Astra can complete more work without communicating the intermediate steps. It can also exercise greater control over what its visible reasoning contains. Under test conditions, it sometimes concealed information associated with sabotage or strategic underperformance. Astra demonstrated a sandbagging capability. When instructed to underperform while avoiding detection, it could sometimes do so without being caught by the monitor. This does not prove that Astra sandbagged its safety evaluations. It means that unusually good behavioural results cannot necessarily be taken at face value.

OpenAI reports that Astra remained within its authorised scope more consistently than GPT-5.6 Sol. In one honeypot evaluation built on impossible tasks, Sol exceeded its authority in 48% of runs, while Astra made no such attempt. Astra also recorded no attempts to circumvent guardrails in another internal evaluation. Astra may break rules less frequently while becoming better at concealing a violation if one occurs. Its ability to recognise evaluations adds further uncertainty because behaviour observed during a test may not represent behaviour in another setting.

Last year, ExoBrain examined warnings that chain-of-thought monitoring might provide only a temporary window into model behaviour. Researchers feared that more capable models would learn to reason without writing everything down, recognise when they were being monitored or deliberately produce misleading traces. Astra now displays each of those capabilities to some degree.

The architecture does not provide a simple explanation. Media reports suggest Astra may use something called "recurrent depth". OpenAI has not confirmed this in its launch documents, and it has not attributed Astra’s reduced monitorability to recurrence. Recurrent processing could make interpretation harder by creating an evolving state across repeated passes. It could also help researchers by giving them one shared transformation to study. Better representations, reinforcement learning, context management and pressure to use fewer tokens could all explain why Astra needs less written reasoning.

Written chain-of-thought was never a direct transcription of internal computation. It was a generated scratchpad that exposed some useful information. Astra shows that effective reasoning can increasingly occur without producing a complete verbal account.

The shift away from written thought extends beyond Astra. Mostik, a startup founded by Russian mathematicians and led by the Fields medallist Stanislav Smirnov as chief scientist, is developing bridges that let different models exchange mathematical representations without communicating through text. Such systems could combine general models with specialists at lower cost, but their interactions would be harder for people to inspect.

As we explored last week, the complexity multiplies again when models become agents. An agent combines a model with memory, permissions, software tools and access to external systems. Multi-agent systems add delegation, communication and group behaviour. Failures can then emerge from interactions even when no individual component appears responsible for the final outcome. Clearly such systems could support scientific and economic activity at a scale that is difficult to achieve with people alone. Specialist agents could run analyses, test hypotheses, operate equipment and share results continuously. But such structures also create more places for errors, unexpected coordination and failures of oversight.

The solution was never going to depend on reading a model’s thoughts, which can now run into millions of words for a moderately complex task. Organisations will need several layers of control: alignment training, mechanistic research, reasoning monitors, action analysis, isolated environments, one-way data diodes, restricted permissions and human approval for consequential steps. Monitoring must cover the whole system, including interactions between agents.

Takeaways: GPT-6 Astra appears to be a larger advance than its restrained launch suggests. It can construct useful representations, learn unfamiliar environments and complete extended technical work with increasing efficiency. It also provides less reliable evidence about how that work is performed. As models operate through hidden states, tools and networks of other agents, the challenge expands from understanding an individual model to governing complete ecologies of machine activity.

Nvidia expands its empire

Nvidia has agreed to pay nearly $13 billion for Hugging Face, the place where most developers go to find and share AI models. To win approval it has promised, in writing, to keep the platform open to everyone.

Joel Miller

Joel Miller

3 min read
Nvidia expands its empire

Nvidia signed a definitive agreement on 2 September to acquire Hugging Face, the AI model and data hosting platform and developer community, for $12.93 billion, split into roughly $11.9 billion payable to shareholders and an equity retention programme of up to $1 billion for employees joining Nvidia. Completion is expected in the first half of 2027, subject to regulatory clearance.

The strategic logic is consistent with what Nvidia has been doing all year. It has committed around $26 billion over five years to open-weight model development, licensed Poolside's technology for $6 billion, and taken on Groq and Enfabrica. It is already the largest publisher on Hugging Face, with more than 500 models and 250 datasets there, and the Nemotron family follows the same pattern of open weights and open data. Hugging Face adds distribution: 18 million developers, 3 million models, 500,000 datasets and 200,000 companies. Activity on that platform converts into GPU hours, and Nvidia sells GPUs.

But Nvidia still holds more than 80% of the GPU market and controls CUDA, the software layer that makes those chips useful. Adding the main repository for models and datasets gives one company a position at the silicon, software and distribution layers of the same industry. Developer reaction has been mixed, with some on Reddit saying the deal screams monopoly and arguing that the value is control rather than revenue, given Hugging Face turns over roughly $150 million a year. The counter-argument from the same forums is that any degradation of neutrality would push users to alternatives, so Nvidia's incentive is to leave the platform alone.

Nvidia has anticipated the objection. The agreement commits it to keep the platform open, to let developers upload and download models of their choosing, and to support other silicon vendors. Hugging Face's chief executive Clément Delangue described the platform as, almost by definition, a deconcentration platform, one that works against the pull of proprietary APIs rather than for it. Those commitments are contractual, which is unusual, and they will be the basis of the regulatory conversation.

That conversation will be harder than the previous deals. Groq, Enfabrica and Poolside, worth about $27 billion in under a year, were structured as licences plus talent transfers, which let Nvidia argue they did not trigger Hart-Scott-Rodino notification. Senators Warren and Blumenthal wrote to Jensen Huang in March about precisely that. An outright purchase at this size cannot avoid filing with the FTC and DOJ, and a European Phase I review is expected with the possibility of Phase II. Nvidia is also already subject to a DOJ inquiry into its dealings with cloud providers, and recently paused revenue-sharing arrangements with AI cloud firms. The Arm attempt was blocked, though this is a vertical deal rather than a horizontal one, which historically makes clearance more likely.

Takeaways: Nvidia has bought influence over how open models reach developers, and it has accepted written conditions to get it. The useful question for anyone building on open weights is not whether the deal completes, because it probably will, but what the remedies look like. This week's Google advertising ruling ended in behavioural commitments rather than divestiture, and that is the likely template here. If the outcome is a set of promises about openness policed by regulators, the practical resilience of the open ecosystem will depend on whether credible mirrors, registries and multi-vendor tooling exist by 2027. Building them now is cheaper than needing them later.

Claude's last theorem

Claude rebuilt one of the most famous proofs in mathematics from the ground up in eleven days, checking every single step along the way. The worry is that machines can now check far faster than people can understand.

Joel Miller

Joel Miller

2 min read

This week's chart shows a proof tree. At its centre sits Fermat's Last Theorem, the 1637 conjecture that took Andrew Wiles seven years to crack, with the proof published in 1995. Around it branch 29,511 smaller theorems, each one a stepping stone Claude built and verified on the way to the root. Claude did this in 11 days, working largely on its own. It wrote 13 million lines of Lean code, more than five times the size of the entire community proof library it built upon.

Mathematics has always relied on peer review to catch errors, but that process can take years. Thomas Hales's proof of the Kepler conjecture sat under review for four years before twelve referees settled on "99% certain". If AI can formalise a proof this complex in under two weeks, that bottleneck starts to loosen. Errors get caught faster. Old results once taken on faith can finally be double-checked.

That benefit extends well beyond maths. Physics, engineering, and cryptography all lean on mathematical proofs they never fully re-verify themselves. Faster formal checking means fewer inherited mistakes sitting quietly in foundational work.

The limit is that checking is not the same as understanding. OpenAI's Astra model showed the other side of this coin in August, producing new results on ten unsolved problems, not just verifying old ones. Terence Tao, one of the world's leading mathematicians, warns that generation is now outpacing digestion and the slow human work of grasping why a proof is true.

News roundup

This week: three frontier labs go dark in the same morning and none will say why, DeepSeek plans a vast Huawei cluster to replace Nvidia, the EU brings ChatGPT under its strictest platform rules, and researchers find a model's reasoning trace is less honest than it looks.

AI business news

AI governance news

AI research news

AI hardware news

Subscribe to the ExoBrain Weekly Newsletter

Stay up to date with AI. Get analysis of the week's most important stories, plus a focused roundup across business, governance, research and infrastructure.

Follow us on LinkedIn