The Ox Alpha stealth model appeared on OpenRouter on Thursday, listed anonymously as a reasoning model from an unidentified third-party provider, and has since generated a level of excitement among developers that its anonymous creators presumably intended.
OpenRouter describes Ox Alpha as a model designed for coding, sustained agentic work, and production workloads, suited, in its own framing, to long-horizon software engineering, complex reasoning, and workflows that combine text with visual context. The model was offered free for a week with what OpenCode, an open-source AI coding agent, described on X as “near unlimited usage.”
Token Capacity That Commands Attention
The provider’s stated infrastructure is the first thing that raises eyebrows. OpenCode noted that the capacity on offer is 100 trillion tokens per day, roughly 100 times the number of tokens Visa said it uses in an entire month. That is an operationally unusual claim for a stealth launch, and it points to a provider with very substantial compute behind it.
The model’s own specifications add further context. According to AlphaMatch, Ox Alpha carries a context window of 1,048,576 tokens (approximately one million) alongside a maximum output of 131,072 tokens. Both figures sit well above what most production models offer today, and they reinforce the impression that whoever built this was not operating from a standing start. A one-million-token context window demands significant engineering investment; it is not a configuration a small or resource-constrained team assembles quietly.
The Ox Alpha Stealth Model and the Attribution Debate
Speculation about the model’s origins has moved quickly and then stalled. Early analysis focused on Z.ai, the Chinese lab responsible for GLM-5. Wccftech, an online tech publication, noted that Z.ai had previously tested GLM-5 anonymously under the name “Pony Alpha,” and that developers had identified similarities between Ox Alpha’s tokeniser behaviour and responses and those of the GLM family. That pattern, a Chinese lab running a quiet stealth test under a neutral name before a formal release, fitted the known precedent.
The case did not hold for long. Wccftech subsequently highlighted a competing analysis suggesting Ox Alpha’s tokeniser could instead point toward Microsoft’s MAI family, which carries entirely different implications for the model’s origin and intended deployment. By Saturday, Andrew Curran, a prominent AI analyst, wrote on X that GLM had been the leading theory on Friday night, but that by Saturday morning “people seem less sure of anything.” At the time of writing, no attribution had been confirmed.
The broader backdrop matters here. Chinese AI labs have been closing the performance gap with their American counterparts at speed, and doing so at lower cost. Moonshot’s Kimi K3, released in July, is a 2.8 trillion-parameter open-weight model built for coding, reasoning, and agentic tasks that drew attention in Silicon Valley for its performance and lower price. Zhipu, DeepSeek, and Moonshot AI are all names that now appear routinely alongside OpenAI and Anthropic in developer benchmarking discussions.
The stealth format itself is not unusual for a major lab testing a frontier model ahead of formal release, it allows real-world load and response-quality data to be gathered without committing to a public product announcement. What is less common is the combination of a one-million-token context window, six-figure max output, and declared capacity of 100 trillion daily tokens behind an anonymous listing. Those are not incremental improvements on an existing public model; they are the specifications of something built for serious production scale.
Patrick Collison, CEO of Stripe, tried the model and said on X that “it’s very impressive”, a brief endorsement, but one that carries weight in the developer community and helped push Ox Alpha into wider circulation over the weekend.
Whoever built Ox Alpha will presumably surface eventually: stealth tests of this scale are rarely permanent. The more consequential question, once an attribution is confirmed, will be what the model’s release strategy and specification choices tell us about where the next wave of frontier capability is coming from.


