Recently, an anonymous new model named Ox Alpha rapidly gained popularity after being released on overseas mainstream large model testing and verification platforms OpenRouter and OpenCode. On its launch day, it surged to the top of OpenRouter, a global large model aggregation and routing platform, and broke the historical record for single-day token usage, representing a fourfold increase over the previous maximum token volume.
The latest data released by OpenRouter shows that, in the week ending August 24, Ox Alpha accounted for 19% of the platform's token usage. Meanwhile, on OpenCode, a globally high-heat coding agent platform, Ox Alpha immediately ended its predecessor model's 56-day reign at the top upon launch, becoming the new model concentratedly adopted by developers worldwide.
For this model, which set historical records on dual overseas platforms, netizens quoted the recently hit film “Niu Lai” and dubbed it the “Niu Lai” model, which was ultimately claimed by Haidian large model enterprise Z. AI. On August 26, Z. AI launched and open-sourced the GLM-5.3-Flash model, the first native multimodal model in the GLM-5 series. The reporter found that the major breakthrough of GLM-5.3-Flash was built on domestic computing cards, marking the first large-scale collective “overseas training exercise” for domestic cards.
Call Volume Hits New Highs on Dual Platforms
The GLM-5.3-Flash model has a total of 320 billion parameters, supports a 1-million-token context, and excels in coding and agentic capabilities. This model surpasses GLM-5.2 in capability, scoring 57 points in the globally authoritative AA Comprehensive Intelligence Index, tying with Claude Opus 4.8 — the most popular model from U. S. AI enterprise Anthropic — placing it within the global frontier model capability range.
In various benchmark tests and practical applications, GLM-5.3-Flash comprehensively surpassed GLM-5.2, which has twice the parameter count. Z. AI stated that the GLM-5.3-Flash architecture is specifically designed for extremely low cost. Its total parameter count is comparable to GLM-4.5, but its activated parameter count and number of layers are nearly halved. Combined with Z. AI's latest 30T-token multimodal pre-training corpus, GLM-5.3-Flash can achieve stronger performance with fewer computational resources.
GLM-5.3-Flash has also undergone multiple architectural upgrades. It is the first open-source frontier model to adopt a hybrid architecture of sparse attention and linear attention, significantly reducing long-context service costs while maintaining precise long-context capabilities, and employs manifold-constrained hyper-connections to enhance model scaling capabilities.

It is understood that GLM-5.3-Flash is the first native multimodal model in the GLM-5 series. Visual capabilities are natively integrated into the model, enabling it to autonomously determine when “observation” is needed and utilize visual feedback to guide subsequent actions. For tasks such as front-end development, game development, and 3D simulation, the final products include interfaces, interactions, or virtual worlds that users can perceive. Without any external materials, GLM-5.3-Flash autonomously operated for 16 hours to construct a professional chef's residence and test kitchen of approximately 400 square meters.
The “Niu Lai” Model Running on Domestic Chips
To widely collect user feedback, GLM-5.3-Flash underwent large-scale testing on two overseas platforms — OpenCode and OpenRouter — in the form of the anonymous model Ox Alpha prior to its official release. It rapidly became the most popular model of the week, with token call volume reaching 62T, setting a new record for call volume on both platforms.

The reporter learned that all online traffic of GLM-5.3-Flash is carried by 100,000 domestic chips. This marks Z. AI's first attempt to use a large domestic chip cluster to serve massive-scale traffic, directly serving real global workloads. In this first large-scale collective “overseas expedition” of domestic computing chips, to overcome the relatively limited computing power and memory capacity of individual chips, Z. AI connected these chips through a self-developed high-bandwidth interconnection network. It built a dedicated inference engine based on the open-source inference engine SGLang. The entire construction process was greatly accelerated by the GLM-5.3-driven infrastructure intelligent agent, which assisted engineers in developing and optimizing operators, diagnosing performance bottlenecks, and improving the deployment service stack, forming a positive cycle of “model optimizing the system, system carrying the model. ”

The GLM-5.3-Flash model running on domestic computing power consumed nearly one-fifth of OpenRouter's token usage, handling real, high-concurrency business traffic on a global mainstream platform. This represents not only a large-scale validation facing global developers, but also signifies that domestic chips are entering the global large model service chain and being tested in real market demand. GLM-5.3-Flash was open-sourced immediately upon launch, and all domestic computing chips can run this frontier model, marking a milestone significance for the development of both the open-source ecosystem and domestic chips.
Frontier Models Enter the Era of Inclusive Access
GLM-5.3-Flash has entered the global frontier model capability range, comprehensively surpassing GLM-5.2 which has twice the parameter count, with programming performance comparable to Claude Opus 4.8. Yet its current pricing is only 1/10 that of GLM-5.2 and 1/40 that of Claude Opus 4.8.
Z. AI stated that compared with the initial baseline on the same hardware, GLM-5.3-Flash's end-to-end service performance has improved by 3 times, and its hardware efficiency and per-token cost have reached levels comparable to mainstream NVIDIA GPUs. This proves that domestic chips can completely and efficiently support frontier model inference demands at scale in an economical manner.
When an open-source model with capabilities matching overseas closed-source flagship models drives prices down to 1/40 of the latter, it signals that domestic frontier models are entering the era of inclusive access.