Ox Alpha Was Z.ai's GLM-5.3-Flash, and the Free Week Is Over
Ox Alpha was Z.ai's GLM-5.3-Flash. OpenRouter's model page now names the developer, the stealth listing serves no providers, and OpenCode has dropped Ox Alpha Free. The named model lists at $0.15 per 1M input tokens and $0.50 output; the 50% launch discount ended September 9, 2026, and as of September 10 Z.ai's pricing page carries list prices only.
The stealth model has been unmasked and the free window has closed. Ox Alpha was Z.ai's GLM-5.3-Flash. OpenRouter's model page now carries the line "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash", stealth/ox-alpha no longer routes to any provider, and OpenCode has removed Ox Alpha Free from its documentation entirely. The vendor's own figures for the model are below.
The most consequential line in the reveal is the one about your prompts. OpenRouter says they were retained by the provider. OpenCode's docs, while the model was live, said the provider followed a zero-retention policy. That disagreement is now settled, and it is settled against the more comfortable reading.
Key Takeaways
- Ox Alpha was Z.ai GLM-5.3-Flash. OpenRouter's model page names ZAI as developer and operator and links to
z-ai/glm-5.3-flash. (OpenRouter, Ox Alpha) - The listing is dead.
GET /api/v1/models/stealth/ox-alpha/endpointsreturns an emptyendpointsarray, and OpenRouter's Stealth provider page says "Stealth has no models available on OpenRouter right now." (OpenRouter, Stealth) - OpenCode has removed it. Neither
Ox Alpha Freenorx-preview-f-freeappears in the Zen model table, the Zen pricing table, the Zen privacy exceptions, or the liveopencode.ai/zen/v1/modelsresponse. (OpenCode Zen docs) - Prompts were retained. OpenRouter: "Prompts and completions for this model were retained by the provider and are not used for training." (OpenRouter, Ox Alpha)
- The named model costs money. $0.15 per 1M input and $0.50 per 1M output. The 50% launch discount ($0.075 and $0.25) ended September 9, 2026; as of September 10 Z.ai's pricing page carries list prices only. (Z.ai pricing)
- It is a 320B-parameter model with 18B active, and Z.ai calls it "the first open-source frontier model to adopt a hybrid architecture combining sparse attention and linear attention". (Z.ai, GLM-5.3-Flash)
- Benchmark tables and lab attributions that circulated during the preview were community claims. The vendor's own numbers only became citable on August 26.
What the Reveal Actually Says
OpenRouter attached a notice to the Ox Alpha model page rather than deleting it. The notice reads in full: "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash. Prompts and completions for this model were retained by the provider and are not used for training; all other use is governed by the Stealth Model Terms." The page's FAQ adds a new first question, "What model was Ox Alpha?", answered with the same attribution and a link to openrouter.ai/z-ai/glm-5.3-flash. Every tense on the page is past tense. (OpenRouter, Ox Alpha)
That is the whole disclosure. There is no post-mortem, no traffic figure, and no statement about why Z.ai chose to run a stealth preview. What there is instead is a live successor listing: z-ai/glm-5.3-flash first appeared in OpenRouter's models API at 13:59 UTC on August 26, 2026, six days after stealth/ox-alpha appeared at 20:04 UTC on August 20.
The Free Window Did Not Get Extended, It Got Retired
Both hosts have withdrawn the free listing rather than repricing it in place.
On OpenRouter, stealth/ox-alpha is absent from /api/v1/models, and the endpoints call for it returns "endpoints": []. The model page still renders, but with no providers, no pricing table, and no playground, because nothing serves it. The Stealth provider page reports "0 models". (OpenRouter, Stealth)
On OpenCode the removal is more complete. The Zen documentation's model table, pricing table, and privacy exception list all previously carried Ox Alpha Free; today none of them mentions it, and the id x-preview-f-free returns nothing from opencode.ai/zen/v1/models. OpenCode's stealth lane itself continues: the Zen docs still describe Big Pickle as "a stealth model that's free on OpenCode for a limited time", alongside free limited-time listings for MiMo-V2.5 Free, Nemotron 3 Ultra Free, Nemotron 3.5 Lightning Free, and Muse Spark 1.3 Contributor Free (as of September 10, 2026; Hy3 Free has since left the table). (OpenCode Zen docs)
What replaced it on the subscription tier is not free either. OpenCode Go lists the model in its usage table at 6,320 requests per 5-hour window (as of September 10, 2026; at the reveal it was 3,160 with a "2x usage limits for a limited time" banner, and the banner has since moved to DeepSeek V4.1 Flash at 4x) -- a numbered allowance, which is exactly what the anonymous listing was exempt from. For scale, the same table gives Kimi K3 110 requests, GPT 5.6 Luna 2,050, and MiniMax M3 3,200. Go costs $10 per month. (OpenCode Go)
The Retention Contradiction, Resolved
While the model was live this page documented a disagreement between its two hosts. OpenRouter said prompts and completions were retained by the provider and not used for training. OpenCode's docs said the provider "follows a zero-retention policy and does not use your data for model training." Both agreed on no training; only one claimed no retention.
That question is now answered by the only party still making a statement about it. OpenRouter's reveal notice keeps the retention language and puts it in the past tense: prompts and completions were retained. OpenCode's zero-retention sentence no longer applies to any Ox Alpha listing, because the listing is gone.
If you ran proprietary code through x-preview-f-free during the free week on the strength of OpenCode's wording, the operative fact is that the provider kept it. Z.ai has published no separate statement about preview-period data, and OpenRouter points at the Stealth Model Terms rather than at a Z.ai policy. This is the concrete cost of a stealth preview: the terms you agreed to were written by a party you could not name at the time.
What GLM-5.3-Flash Is, From the Vendor
Everything in this section comes from Z.ai's own developer documentation, which is the source that did not exist while the model was anonymous.
| Item | GLM-5.3-Flash | Source |
|---|---|---|
| Model code | glm-5.3-flash (Z.ai); z-ai/glm-5.3-flash (OpenRouter) | Z.ai docs; OpenRouter |
| Total parameters | 320B | Z.ai docs |
| Activated parameters | 18B | Z.ai docs |
| Architecture | Hybrid sparse plus linear attention | Z.ai docs |
| Efficiency vs GLM-5.3 | Attention compute reduced 3.01x, KV cache 4.44x | Z.ai docs |
| Context window | 1M tokens (OpenRouter's API reports 1,310,720) | Z.ai docs; OpenRouter models API |
| Modalities | Text, image, and video in; text out | OpenRouter models API |
| List price | $0.15 / 1M input, $0.50 / 1M output | Z.ai pricing |
| Launch discount | 50% off ($0.075 / $0.25) through September 9, 2026; ended, and gone from Z.ai's page as of September 10, 2026 | Z.ai pricing |
| Cached input | $0.03; cached-input storage marked "Limited-time Free" | Z.ai pricing |
| Released | August 26, 2026 on OpenRouter | OpenRouter |
(Z.ai, GLM-5.3-Flash; Z.ai pricing; OpenRouter, GLM 5.3 Flash)
Z.ai describes it as "the first native multimodal model in the GLM-5 series" and says visual capability is integrated into the coding loop so the model can "actively observe interfaces, rendered results, and interaction feedback, and continuously test and improve its work accordingly." The stated recommended settings are temperature: 1, top_p: 0.95, and reasoning_effort: max, with thinking.type supporting only enabled and thinking.clear_thinking recommended as false. For streaming, Z.ai recommends enabling stream and tool_stream together. (Z.ai, GLM-5.3-Flash)
The efficiency numbers are the ones that explain the stealth preview. A 320B model with 18B activated parameters and a 3x reduction in attention compute is a model built to be served cheaply at long context, and the way you find out whether it holds up at long context is to give away a week of unlimited agentic traffic. The 1M-context rollout of GPT-5.6 Sol in Codex reached readers through a product they already paid for; this one reached them through an unnamed listing on a router, which is a cheaper way to buy the same evaluation data.
Where to Run It Now
Three routes, in descending order of how much you have to trust an intermediary.
Direct from Z.ai. The model is on the GLM Coding Plan, which Z.ai says carries 3x the quota of GLM-5.3 and uses a points-based system with off-peak discounts. This is the only route where the party setting the data-retention terms is also the party that built the model. (Z.ai, GLM-5.3-Flash)
Through OpenRouter. As of September 15, 2026, OpenRouter's endpoints API lists 27 providers for z-ai/glm-5.3-flash, up from twelve at the reveal, 17 on August 28, 25 on September 10, 26 on September 13, and 27 since September 14. With the launch discount over, Z.ai's own endpoint is at the $0.15 / $0.50 list, and so are 15 others: Together, Baseten, Cloudflare, Parasail, Venice, Fireworks, Friendli, Phala, DigitalOcean, CoreWeave, Crusoe, SiliconFlow, Sail Research, Io Net, and OpenInference. Nine providers undercut the list on their own account: DeepInfra $0.075 / $0.25, Relace $0.09 / $0.30, Morph and Wafer $0.10 / $0.35, StreamLake $0.112 / $0.375, GMICloud $0.1125 / $0.375, Reka and Novita $0.132 / $0.44, and Makora $0.14 / $0.47. Two price above it: NextBit $0.195 / $0.65 (up from $0.177 / $0.59 on September 14) and Modal $0.45 / $1.50. OpenRouter's model page carried the banner "Limited-time 50% discount via ZAI through September 9, 2026 at 16:00 UTC" for a day after the deadline; since September 11, 2026 the banner has been absent from the model page and the endpoints API, re-checked September 15. Routing mode matters here: OpenRouter's Balanced, Nitro, and Exacto modes optimize for price-plus-speed, throughput, and tool-calling accuracy respectively. (OpenRouter, GLM 5.3 Flash; OpenRouter endpoints API)
Through OpenCode Zen. The Zen model table and pricing table now list GLM 5.3 Flash (glm-5.3-flash) at $0.15 input, $0.50 output, and $0.03 cached input per 1M tokens, and the live opencode.ai/zen/v1/models response returns the id. (OpenCode Zen docs)
Through OpenCode Go. $10 a month, 6,320 requests per 5-hour window as of September 10, 2026. Go works with OpenCode or, per its own page, "any agent", so the model is reachable from terminals that wrap OpenCode such as Warp. (OpenCode Go)
What This Episode Is Worth Remembering For
A stealth preview is a trade: the lab buys evaluation traffic with free tokens, and you buy a week of a frontier model with data you cannot get back. That trade is fine when you price it correctly, and the way to price it correctly is to assume the terms are the strictest ones any host publishes rather than the friendliest. On this model two hosts published different terms for the same weights, and the strict one turned out to be the accurate one.
The second lesson is about attribution. During the free week, benchmark tables and lab guesses for Ox Alpha circulated widely and this page declined to repeat them. The reveal took six days. Anything built on a guess about which lab shipped it was a coin flip resolved by a vendor page, and the vendor page is the only artefact that still exists.
Both of the dates this page said to watch have now passed. The 50% discount ended on September 9, 2026, so $0.15 / $0.50 is the day-to-day price of a 1M-context multimodal coding model from its own vendor, with cheaper third-party routes on OpenRouter for anyone who accepts a different host's terms. And GLM-5.3-Flash has reached OpenCode Zen's own model table at list price, the way Grok 4.6 shipped straight into Cursor and OpenRouter on the same day. An anonymous model on a router is an experiment; a named model in your editor is a decision.
Corrections and Updates
September 10, 2026: The 50% launch discount ended September 9. Z.ai's pricing page now carries list prices only; OpenRouter routes the model across 25 providers with Z.ai's endpoint back at list; OpenCode Zen lists GLM 5.3 Flash; OpenCode Go shows 6,320 requests per 5-hour window.
August 27, 2026: This page originally reported the free week and said it would be updated with the vendor's own figures if the provider unmasked the model; the reveal is now the lede and those figures are in the body.
Sources
- OpenRouter, Ox Alpha model page (reveal notice, retention statement, FAQ): https://openrouter.ai/stealth/ox-alpha
- OpenRouter, Stealth provider page ("0 models"): https://openrouter.ai/provider/stealth
- OpenRouter, GLM 5.3 Flash model page (providers, pricing, discount deadline): https://openrouter.ai/z-ai/glm-5.3-flash
- OpenRouter models API and endpoints API (empty provider list for
stealth/ox-alpha, listing timestamps for both models, the 25-provider price list forz-ai/glm-5.3-flash, checked September 10, 2026): https://openrouter.ai/api/v1/models - Z.ai developer docs, GLM-5.3-Flash overview (parameters, architecture, context, recommended settings, Coding Plan quota): https://docs.z.ai/guides/vlm/glm-5.3-flash
- Z.ai developer docs, pricing (list prices; the promotional line was present through September 9, 2026 and absent on September 10): https://docs.z.ai/guides/overview/pricing
- OpenCode, Zen docs (model table, pricing table, privacy exceptions, remaining stealth listings): https://opencode.ai/docs/zen/
- OpenCode, live Zen model list: https://opencode.ai/zen/v1/models
- OpenCode Go (price, usage-limits table, 2x usage banner): https://opencode.ai/go
Discovery note, backing no claim above: the free window was first announced in posts by OpenRouter and OpenCode on x.com on August 20 and 21, 2026. x.com refuses automated requests, so those posts cannot be re-verified and no statement on this page rests on them.
Read next
Keep building the workspace playbookGrok 4.6 Lands in Cursor With a One-Week 2x Usage Window
SpaceXAI released Grok 4.6 on August 12, focused on long-running agents and visual work. It is live in Cursor and Grok Build with 2x included usage for the first week, and in the API plus OpenRouter, Vercel, and Cloudflare. Pricing starts at $2 per million input tokens and $6 per million output tokens.
Codex Opens GPT-5.6 Sol's 1M-Token Context to ChatGPT Accounts
GPT-5.6 Sol's 1M-token context window in Codex is now available for usage through ChatGPT accounts, not only API keys, according to OpenAI developer Tibo Sottiaux. The announcement also repeats a warning that Codex's default context length is tuned for performance and cost.
Warp's Universal Agent Support: The Terminal as an Agentic Development Environment
Warp's April 2026 universal agent support brings first-class integration for Claude Code, Codex, Gemini CLI, and OpenCode in a single terminal -- with vertical tabs, status indicators, code review, mobile remote, and a rich input editor. It is Warp's bet that 'ADE' beats both IDE and traditional terminal for multi-agent work.
Frequently Asked Questions
What was Ox Alpha?
Ox Alpha was a stealth listing on OpenRouter for Z.ai's GLM-5.3-Flash. OpenRouter's model page now states that the model was developed and operated by ZAI and points readers to z-ai/glm-5.3-flash. During the preview the developer was anonymous and OpenRouter served the model as stealth/ox-alpha.
Can you still use Ox Alpha for free?
No. OpenRouter's endpoints API returns an empty provider list for stealth/ox-alpha, and its Stealth provider page reports zero models available. OpenCode's Zen docs no longer list Ox Alpha Free or the x-preview-f-free model id anywhere, including in the pricing and privacy tables.
What does GLM-5.3-Flash cost now?
As of September 10, 2026, Z.ai lists $0.15 per 1M input tokens, $0.03 per 1M cached input, and $0.50 per 1M output, with no promotional price. The 50% launch discount ($0.075 and $0.25) ended September 9, 2026, and Z.ai's own OpenRouter endpoint is back at list price. On OpenRouter, DeepInfra posts the cheapest route at $0.075 / $0.25.
Were prompts sent to Ox Alpha retained?
Yes, according to OpenRouter. Its model page says prompts and completions for the model were retained by the provider and are not used for training. OpenCode's docs had described the provider as zero-retention. The two hosts contradicted each other and OpenRouter's wording is the one that survives.
Where do you run GLM-5.3-Flash today?
As of September 15, 2026, OpenRouter lists 27 providers for z-ai/glm-5.3-flash including Z.ai itself, Together, Baseten, and DeepInfra, up from twelve at the reveal. OpenCode Zen now lists GLM 5.3 Flash at $0.15 / $0.50, and OpenCode Go includes it at 6,320 requests per 5-hour window. Z.ai serves it directly on the GLM Coding Plan at 3x the GLM-5.3 quota.