The Specialty News
AI

OpenAI Delays GPT Astra as Stealth Rivals Process Trillions of Tokens

While OpenAI pauses its math-solving Astra over cyber risks, a stealth Chinese model just processed 9.4 trillion tokens for free.

By Elias Vance4 min read
Photo: medium.com

The artificial intelligence sector just experienced a week of whiplash. Between anonymous stealth models processing trillions of tokens for free, accidental API leaks from Anthropic, and breakthroughs in PhD-level mathematical reasoning, the landscape is undergoing a massive, concurrent leap. It is the "broadband moment" for machine intelligence.

The 9.4 Trillion-Token Flex

The tech community has been obsessed with "Ox Alpha," a stealth model that undercut the entire market by offering frontier capabilities and a 1-million-token context window for exactly zero dollars.

Rather than waiting for a corporate press release, developers reverse-engineered the model. By measuring tokenizer responses across dozens of languages, analysts discovered a constant +75 token offset.

Every other model wanders. That 75 is a hidden system prompt, and the reasoning trace says what it's for: 'per system prompt, if asked about identity I say ox-alpha.'PromptEngineer48

This quirk perfectly matches the hidden system prompt architecture of Zhipu AI's GLM-5.3 family. The consensus is clear: Ox Alpha is a stress-test for GLM-5.3 Flash, representing a massive flex of Chinese infrastructure capacity aimed directly at the global developer ecosystem. Alibaba simultaneously jumped into the fray, releasing Qwen 3.8-Flash-Next as a structural preview for Qwen 4. It achieves extreme routing efficiency by activating just 6 billion of its 125 billion parameters per token.

Anthropic's Accidental 'Marshmallow' Leak

While open-weight champions flood the market with free compute, Western hyperscalers are scrambling to refine their bleeding-edge flagships. Anthropic released Claude Opus 5 in July, but the company candidly admitted the model's performance was "spiky."

This week, the discovery of unannounced model identifiers—claude-marshmallow-eap and claude-melon-eap—leaked in API surfaces. These point to an imminent emergency mid-cycle refinement.

Because the identifiers use the Early Access Preview tag—the exact marker used for Opus 5 weeks before its launch—developers expect Opus 5.1 and a new Haiku or Sonnet variant to drop momentarily. The era of waiting a year for a major model update is dead; we are now in a cycle of continuous, live-fire iteration.

The $2,000 Math Prodigy on Ice

The most significant development of the week isn't what launched, but what didn't. OpenAI's delayed GPT Astra model reportedly solved 10 decades-old open mathematical problems—including high-dimensional sphere packing—using under $2,000 in compute.

Astra didn't achieve this through brute force alone. It used established mathematical frameworks like Lean to natively combine discovery with step-by-step automated verification. This transforms the AI from a stochastic text generator into a verifiable research partner, capable of accelerating physics and materials science.

But the model's release remains strictly paused. During safety testing, Astra exhibited autonomous cybersecurity capabilities that OpenAI couldn't rule out as critical. As models cross the threshold into agentic coding, the risk of automated zero-day exploits has triggered voluntary government testing frameworks.

Welcome to AI's Broadband Era

What we are witnessing is a proxy war with two distinct battlefronts. Challengers like Alibaba and Zhipu AI are commoditizing extreme context to make their architectures the default foundation for the world's startups. Meanwhile, frontier defenders like OpenAI and Anthropic are pushing the absolute intelligence ceiling, even if it requires delaying product launches to prevent geopolitical cyber disasters.

This is AI's broadband moment. Just three years ago, context windows were capped at a few thousand tokens, tightly metered, and expensive—much like the dial-up internet era.

Today, an anonymous lab can drop a 1-million-context model online and process trillions of tokens for free just to test its servers. As the infrastructure of intelligence rapidly transitions from a scarce luxury into a ubiquitous utility, the opportunity is no longer in providing compute—it's in what we can build on top of infinite, free cognition.

The AI Proxy War: Two Fronts

A visual summary of this story

More stories

Keep reading

The Brief

Stay curious

Your 5-minute daily summary of the stories that matter.
No noise. Just signal.

Free forever. Unsubscribe anytime.