AI Catchup

Ox Alpha Was Z.ai's GLM-5.3-Flash, and the Free Week Is Over

By 11 min read

Ox Alpha was Z.ai's GLM-5.3-Flash. OpenRouter's model page now names the developer, the stealth listing serves no providers, and OpenCode has dropped Ox Alpha Free. The named model lists at $0.15 per 1M input tokens and $0.50 output; the 50% launch discount ended September 9, 2026, and as of September 10 Z.ai's pricing page carries list prices only.

The stealth model has been unmasked and the free window has closed. Ox Alpha was Z.ai's GLM-5.3-Flash. OpenRouter's model page now carries the line "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash", stealth/ox-alpha no longer routes to any provider, and OpenCode has removed Ox Alpha Free from its documentation entirely. The vendor's own figures for the model are below.

The most consequential line in the reveal is the one about your prompts. OpenRouter says they were retained by the provider. OpenCode's docs, while the model was live, said the provider followed a zero-retention policy. That disagreement is now settled, and it is settled against the more comfortable reading.

Key Takeaways

  • Ox Alpha was Z.ai GLM-5.3-Flash. OpenRouter's model page names ZAI as developer and operator and links to z-ai/glm-5.3-flash. (OpenRouter, Ox Alpha)
  • The listing is dead. GET /api/v1/models/stealth/ox-alpha/endpoints returns an empty endpoints array, and OpenRouter's Stealth provider page says "Stealth has no models available on OpenRouter right now." (OpenRouter, Stealth)
  • OpenCode has removed it. Neither Ox Alpha Free nor x-preview-f-free appears in the Zen model table, the Zen pricing table, the Zen privacy exceptions, or the live opencode.ai/zen/v1/models response. (OpenCode Zen docs)
  • Prompts were retained. OpenRouter: "Prompts and completions for this model were retained by the provider and are not used for training." (OpenRouter, Ox Alpha)
  • The named model costs money. $0.15 per 1M input and $0.50 per 1M output. The 50% launch discount ($0.075 and $0.25) ended September 9, 2026; as of September 10 Z.ai's pricing page carries list prices only. (Z.ai pricing)
  • It is a 320B-parameter model with 18B active, and Z.ai calls it "the first open-source frontier model to adopt a hybrid architecture combining sparse attention and linear attention". (Z.ai, GLM-5.3-Flash)
  • Benchmark tables and lab attributions that circulated during the preview were community claims. The vendor's own numbers only became citable on August 26.

What the Reveal Actually Says

OpenRouter attached a notice to the Ox Alpha model page rather than deleting it. The notice reads in full: "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash. Prompts and completions for this model were retained by the provider and are not used for training; all other use is governed by the Stealth Model Terms." The page's FAQ adds a new first question, "What model was Ox Alpha?", answered with the same attribution and a link to openrouter.ai/z-ai/glm-5.3-flash. Every tense on the page is past tense. (OpenRouter, Ox Alpha)

That is the whole disclosure. There is no post-mortem, no traffic figure, and no statement about why Z.ai chose to run a stealth preview. What there is instead is a live successor listing: z-ai/glm-5.3-flash first appeared in OpenRouter's models API at 13:59 UTC on August 26, 2026, six days after stealth/ox-alpha appeared at 20:04 UTC on August 20.

The Free Window Did Not Get Extended, It Got Retired

Both hosts have withdrawn the free listing rather than repricing it in place.

On OpenRouter, stealth/ox-alpha is absent from /api/v1/models, and the endpoints call for it returns "endpoints": []. The model page still renders, but with no providers, no pricing table, and no playground, because nothing serves it. The Stealth provider page reports "0 models". (OpenRouter, Stealth)

On OpenCode the removal is more complete. The Zen documentation's model table, pricing table, and privacy exception list all previously carried Ox Alpha Free; today none of them mentions it, and the id x-preview-f-free returns nothing from opencode.ai/zen/v1/models. OpenCode's stealth lane itself continues: the Zen docs still describe Big Pickle as "a stealth model that's free on OpenCode for a limited time", alongside free limited-time listings for MiMo-V2.5 Free, Nemotron 3 Ultra Free, Nemotron 3.5 Lightning Free, and Muse Spark 1.3 Contributor Free (as of September 10, 2026; Hy3 Free has since left the table). (OpenCode Zen docs)

What replaced it on the subscription tier is not free either. OpenCode Go lists the model in its usage table at 6,320 requests per 5-hour window (as of September 10, 2026; at the reveal it was 3,160 with a "2x usage limits for a limited time" banner, and the banner has since moved to DeepSeek V4.1 Flash at 4x) -- a numbered allowance, which is exactly what the anonymous listing was exempt from. For scale, the same table gives Kimi K3 110 requests, GPT 5.6 Luna 2,050, and MiniMax M3 3,200. Go costs $10 per month. (OpenCode Go)

The Retention Contradiction, Resolved

While the model was live this page documented a disagreement between its two hosts. OpenRouter said prompts and completions were retained by the provider and not used for training. OpenCode's docs said the provider "follows a zero-retention policy and does not use your data for model training." Both agreed on no training; only one claimed no retention.

That question is now answered by the only party still making a statement about it. OpenRouter's reveal notice keeps the retention language and puts it in the past tense: prompts and completions were retained. OpenCode's zero-retention sentence no longer applies to any Ox Alpha listing, because the listing is gone.

If you ran proprietary code through x-preview-f-free during the free week on the strength of OpenCode's wording, the operative fact is that the provider kept it. Z.ai has published no separate statement about preview-period data, and OpenRouter points at the Stealth Model Terms rather than at a Z.ai policy. This is the concrete cost of a stealth preview: the terms you agreed to were written by a party you could not name at the time.

What GLM-5.3-Flash Is, From the Vendor

Everything in this section comes from Z.ai's own developer documentation, which is the source that did not exist while the model was anonymous.

ItemGLM-5.3-FlashSource
Model codeglm-5.3-flash (Z.ai); z-ai/glm-5.3-flash (OpenRouter)Z.ai docs; OpenRouter
Total parameters320BZ.ai docs
Activated parameters18BZ.ai docs
ArchitectureHybrid sparse plus linear attentionZ.ai docs
Efficiency vs GLM-5.3Attention compute reduced 3.01x, KV cache 4.44xZ.ai docs
Context window1M tokens (OpenRouter's API reports 1,310,720)Z.ai docs; OpenRouter models API
ModalitiesText, image, and video in; text outOpenRouter models API
List price$0.15 / 1M input, $0.50 / 1M outputZ.ai pricing
Launch discount50% off ($0.075 / $0.25) through September 9, 2026; ended, and gone from Z.ai's page as of September 10, 2026Z.ai pricing
Cached input$0.03; cached-input storage marked "Limited-time Free"Z.ai pricing
ReleasedAugust 26, 2026 on OpenRouterOpenRouter

(Z.ai, GLM-5.3-Flash; Z.ai pricing; OpenRouter, GLM 5.3 Flash)

Z.ai describes it as "the first native multimodal model in the GLM-5 series" and says visual capability is integrated into the coding loop so the model can "actively observe interfaces, rendered results, and interaction feedback, and continuously test and improve its work accordingly." The stated recommended settings are temperature: 1, top_p: 0.95, and reasoning_effort: max, with thinking.type supporting only enabled and thinking.clear_thinking recommended as false. For streaming, Z.ai recommends enabling stream and tool_stream together. (Z.ai, GLM-5.3-Flash)

The efficiency numbers are the ones that explain the stealth preview. A 320B model with 18B activated parameters and a 3x reduction in attention compute is a model built to be served cheaply at long context, and the way you find out whether it holds up at long context is to give away a week of unlimited agentic traffic. The 1M-context rollout of GPT-5.6 Sol in Codex reached readers through a product they already paid for; this one reached them through an unnamed listing on a router, which is a cheaper way to buy the same evaluation data.

Where to Run It Now

Three routes, in descending order of how much you have to trust an intermediary.

Direct from Z.ai. The model is on the GLM Coding Plan, which Z.ai says carries 3x the quota of GLM-5.3 and uses a points-based system with off-peak discounts. This is the only route where the party setting the data-retention terms is also the party that built the model. (Z.ai, GLM-5.3-Flash)

Through OpenRouter. As of September 15, 2026, OpenRouter's endpoints API lists 27 providers for z-ai/glm-5.3-flash, up from twelve at the reveal, 17 on August 28, 25 on September 10, 26 on September 13, and 27 since September 14. With the launch discount over, Z.ai's own endpoint is at the $0.15 / $0.50 list, and so are 15 others: Together, Baseten, Cloudflare, Parasail, Venice, Fireworks, Friendli, Phala, DigitalOcean, CoreWeave, Crusoe, SiliconFlow, Sail Research, Io Net, and OpenInference. Nine providers undercut the list on their own account: DeepInfra $0.075 / $0.25, Relace $0.09 / $0.30, Morph and Wafer $0.10 / $0.35, StreamLake $0.112 / $0.375, GMICloud $0.1125 / $0.375, Reka and Novita $0.132 / $0.44, and Makora $0.14 / $0.47. Two price above it: NextBit $0.195 / $0.65 (up from $0.177 / $0.59 on September 14) and Modal $0.45 / $1.50. OpenRouter's model page carried the banner "Limited-time 50% discount via ZAI through September 9, 2026 at 16:00 UTC" for a day after the deadline; since September 11, 2026 the banner has been absent from the model page and the endpoints API, re-checked September 15. Routing mode matters here: OpenRouter's Balanced, Nitro, and Exacto modes optimize for price-plus-speed, throughput, and tool-calling accuracy respectively. (OpenRouter, GLM 5.3 Flash; OpenRouter endpoints API)

Through OpenCode Zen. The Zen model table and pricing table now list GLM 5.3 Flash (glm-5.3-flash) at $0.15 input, $0.50 output, and $0.03 cached input per 1M tokens, and the live opencode.ai/zen/v1/models response returns the id. (OpenCode Zen docs)

Through OpenCode Go. $10 a month, 6,320 requests per 5-hour window as of September 10, 2026. Go works with OpenCode or, per its own page, "any agent", so the model is reachable from terminals that wrap OpenCode such as Warp. (OpenCode Go)

What This Episode Is Worth Remembering For

A stealth preview is a trade: the lab buys evaluation traffic with free tokens, and you buy a week of a frontier model with data you cannot get back. That trade is fine when you price it correctly, and the way to price it correctly is to assume the terms are the strictest ones any host publishes rather than the friendliest. On this model two hosts published different terms for the same weights, and the strict one turned out to be the accurate one.

The second lesson is about attribution. During the free week, benchmark tables and lab guesses for Ox Alpha circulated widely and this page declined to repeat them. The reveal took six days. Anything built on a guess about which lab shipped it was a coin flip resolved by a vendor page, and the vendor page is the only artefact that still exists.

Both of the dates this page said to watch have now passed. The 50% discount ended on September 9, 2026, so $0.15 / $0.50 is the day-to-day price of a 1M-context multimodal coding model from its own vendor, with cheaper third-party routes on OpenRouter for anyone who accepts a different host's terms. And GLM-5.3-Flash has reached OpenCode Zen's own model table at list price, the way Grok 4.6 shipped straight into Cursor and OpenRouter on the same day. An anonymous model on a router is an experiment; a named model in your editor is a decision.

Corrections and Updates

September 10, 2026: The 50% launch discount ended September 9. Z.ai's pricing page now carries list prices only; OpenRouter routes the model across 25 providers with Z.ai's endpoint back at list; OpenCode Zen lists GLM 5.3 Flash; OpenCode Go shows 6,320 requests per 5-hour window.

August 27, 2026: This page originally reported the free week and said it would be updated with the vendor's own figures if the provider unmasked the model; the reveal is now the lede and those figures are in the body.

Sources

Discovery note, backing no claim above: the free window was first announced in posts by OpenRouter and OpenCode on x.com on August 20 and 21, 2026. x.com refuses automated requests, so those posts cannot be re-verified and no statement on this page rests on them.

Keep building the workspace playbook

Frequently Asked Questions

What was Ox Alpha?

Ox Alpha was a stealth listing on OpenRouter for Z.ai's GLM-5.3-Flash. OpenRouter's model page now states that the model was developed and operated by ZAI and points readers to z-ai/glm-5.3-flash. During the preview the developer was anonymous and OpenRouter served the model as stealth/ox-alpha.

Can you still use Ox Alpha for free?

No. OpenRouter's endpoints API returns an empty provider list for stealth/ox-alpha, and its Stealth provider page reports zero models available. OpenCode's Zen docs no longer list Ox Alpha Free or the x-preview-f-free model id anywhere, including in the pricing and privacy tables.

What does GLM-5.3-Flash cost now?

As of September 10, 2026, Z.ai lists $0.15 per 1M input tokens, $0.03 per 1M cached input, and $0.50 per 1M output, with no promotional price. The 50% launch discount ($0.075 and $0.25) ended September 9, 2026, and Z.ai's own OpenRouter endpoint is back at list price. On OpenRouter, DeepInfra posts the cheapest route at $0.075 / $0.25.

Were prompts sent to Ox Alpha retained?

Yes, according to OpenRouter. Its model page says prompts and completions for the model were retained by the provider and are not used for training. OpenCode's docs had described the provider as zero-retention. The two hosts contradicted each other and OpenRouter's wording is the one that survives.

Where do you run GLM-5.3-Flash today?

As of September 15, 2026, OpenRouter lists 27 providers for z-ai/glm-5.3-flash including Z.ai itself, Together, Baseten, and DeepInfra, up from twelve at the reveal. OpenCode Zen now lists GLM 5.3 Flash at $0.15 / $0.50, and OpenCode Go includes it at 6,320 requests per 5-hour window. Z.ai serves it directly on the GLM Coding Plan at 3x the GLM-5.3 quota.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.