An anonymous AI model has quietly appeared on the model distribution platform OpenRouter. Testing by independent researchers shows that the model, named “stealth/ox-alpha,” not only surpasses mainstream models like GPT and Claude in certain coding benchmarks, but its technical characteristics also strongly point to Zhipu’s unreleased next-generation multimodal flagship model.

According to a test report published by technical researcher Ben Davis on August 21, ox-alpha went live on OpenRouter on August 20 and is currently offering one week of free access. The model supports text, image, and video input, has reasoning capabilities, and features a context window of 1.048 million tokens. Davis stated on X that he is “99% certain” ox-alpha belongs to Zhipu’s GLM-5.x series, with multiple lines of evidence—including the video encoder, tokenizer, output style, and audio rejection behavior—all pointing to this conclusion. His testing also showed that ox-alpha outperforms GPT and Claude in certain comparison tests.

It should be emphasized, however, that these conclusions are based on individual testing and technical inference, and have not yet been officially confirmed by Zhipu.

Video Encoder Emerges as Key Evidence

Davis’s testing traced ox-alpha’s technical origins across multiple dimensions, including the video encoder, tokenizer, audio interface, and output style. Among these, the high degree of matching in the video encoder was considered the most compelling evidence.

Testing showed that across four controlled video sets, ox-alpha and GLM-5V-Turbo consumed video tokens in a completely identical manner, with both exhibiting the same three characteristics: frame-rate-independent frame sampling, a duration scaling ratio of approximately 147 tokens per second, and a per-frame resolution scaling mechanism. In contrast, candidate models such as MiMo v2.5, Qwen 3.8 Max, and GLM-4.6V all displayed distinctly different video encoding characteristics.

Tokenizer testing also pointed to GLM. The report stated that after testing 25 different prompts, ox-alpha’s token counts matched GLM-5.3 exactly, with only a fixed +75 token hidden wrapper difference, suggesting the two may share the same vocabulary.

Additionally, ox-alpha’s rejection of audio input is consistent with GLM-5V, whereas MiMo v2.5—listed as a primary competing candidate—supports audio input, further weakening the latter’s likelihood. In terms of output style, ox-alpha uses approximately 1.3 emojis per thousand characters, which is relatively close to the GLM/Qwen series; by comparison, Claude, GPT-5.6, and Grok showed near-zero emoji usage rates in the same test environment.

Coding Capability Already Surpasses GLM-5.3

On the capability testing front, the report cited DeepSWE benchmark data showing that ox-alpha passed 8 out of 10 deterministic tasks, achieving an 80% Pass@1 rate. By comparison, Claude Fable 5 achieved a 65% pass rate, GLM-5.3 and Grok 4.6 both achieved 62%, and GPT-5.6-sol achieved 52%. However, since the number of test runs varies across models and ox-alpha currently has a relatively small sample size, these results still require further independent verification.

ModelDeepSWE Pass@1stealth/ox-alpha80%Claude Fable 565%GLM-5.362%Grok 4.662%GPT-5.6-sol52%

Note: Test sample sizes and run counts vary by model. ox-alpha currently has a relatively small sample size, and the above data should be treated as preliminary reference only.

In the “meriyah-explicit-resource-declarations” task, ox-alpha passed on its first attempt, whereas GLM-5.3, GPT-5.6-sol, and Grok 4.6 had previously all ended with 0/4 results. At the same time, ox-alpha maintained a clean pass across all 51,469 regression tests.

The report also documented an agent task involving 69 tool calls. Throughout the process, the model made only one error, with no retry loops and relatively low reasoning overhead. Based on this, Davis concluded that ox-alpha’s capabilities are already significantly stronger than GLM-5.3, suggesting it is more likely a next-generation model checkpoint rather than a simple variant.

Multiple Clues Converge on Zhipu

Beyond the technical fingerprint, Davis also conducted cross-verification based on the model’s release timing and Zhipu’s previous testing practices. Zhipu released the text-only GLM-5.3 on August 14, and a unified vision flagship model has been a focus of community attention. More importantly, Zhipu has a precedent of testing models through stealth channels—Pony Alpha was ultimately confirmed to be related to GLM-5.

In terms of model scale, the report noted that ox-alpha’s decoding speed differs from GLM-5V-Turbo by approximately 6%. The latter has 744B total parameters and 40B active parameters, leading Davis to speculate that ox-alpha may employ a Mixture-of-Experts architecture of similar scale. The report further argued that if the model indeed has approximately 40B active parameters, the operator’s claimed daily serving capacity of 100 trillion tokens becomes more technically and economically feasible.

Meanwhile, the report also ruled out other potential sources one by one. Xiaomi’s MiMo shows clear differences from ox-alpha in video encoder and audio interface; DeepSeek has not previously released video capabilities, uses a different tokenizer, and has a different model release pattern; Google, Qwen, xAI, OpenAI, and Anthropic were all deemed mismatched with ox-alpha in tokenizer, output style, or video encoder characteristics.

Notably, ox-alpha’s current free access window may last until August 27. Davis pointed out that some similar stealth models in the past were officially claimed by the relevant Chinese AI labs after their free testing periods ended.

As of now, Zhipu has not issued an official response regarding ox-alpha’s identity. If the model is ultimately confirmed to belong to Zhipu’s next-generation multimodal flagship, this “stealth test” on OpenRouter may serve as a public dress rehearsal ahead of its official launch. Until official confirmation arrives, the question of ox-alpha’s true origin remains unanswered—but from video encoder to tokenizer to model behavior, multiple independent tests currently point all clues in the same direction.