The Specialty News
AI

Z.ai Drops $0.045 AI Model That Rivals Claude — and Silicon Valley Thought It Was Google

By stealth-launching GLM-5.3-Flash under a pseudonym, a Chinese lab proved frontier intelligence can run on domestic chips for pennies.

By Julian Thorne4 min read
Z.ai Drops $0.045 AI Model That Rivals Claude — and Silicon Valley Thought It Was Google
Photo: reuters.com

For nearly a week in August, Western software developers were convinced a mysterious AI model dominating the OpenRouter charts was a leaked version of Google’s Gemini. It handled a million tokens at a time and appeared out of nowhere under the pseudonym Ox Alpha. But the model didn't belong to Mountain View. It was built in Beijing, running entirely on domestic Chinese silicon, and it just broke the pricing floor for frontier artificial intelligence.

The Disguise

When the model hit OpenRouter, speculation ran wild that it was a secret release from a Western tech giant. A contingent of independent researchers ignored the rumor mill. By prompting the model in Chinese and analyzing its compression rates, they found it returned exact API exception errors linked to Zhipu AI, a prominent Chinese lab.

The stealth launch was entirely intentional. By stripping its name from the release, Zhipu—operating globally as Z.ai—forced Western developers to evaluate the model purely on its merits. Patrick Collison, CEO of Stripe, tested it early and publicly called the code very impressive, validating the architecture before anyone knew its origin.

The bait-and-switch worked perfectly. The Chinese internet celebrated the unmasking, but enterprise finance departments in the US are now staring at a much bigger problem.

The Ten-Cent Paradigm

The Ten-Cent Paradigm
Photo: reuters.com

The unmasked model, GLM-5.3-Flash, introduces a 320-billion parameter architecture that relies on just 18 billion active parameters during inference. Z.ai released the weights under an MIT open-source license and announced enterprise pricing at a heavily discounted 15 cents per million input tokens.

To put the $0.045-per-task benchmark into perspective, this represents roughly a 10x price collapse for this specific tier of reasoning. It is the equivalent of finding out a Michelin-starred tasting menu now costs the exact same price as a vending machine candy bar.

The geopolitical stakes are just as sharp. Z.ai claims they are serving up to 100 trillion tokens per day capacity entirely on domestically produced hardware. They proved that top-tier AI performance can scale globally without relying on the newest Western silicon. But running frontier models on the cheap hides a massive operational catch.

Z.ai's Stealth AI Breaks Price Floor

A visual summary of this story

More stories

Keep reading

The Brief

Stay curious

Your 5-minute daily summary of the stories that matter.
No noise. Just signal.

Free forever. Unsubscribe anytime.