<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>free AI model Archives | Blog</title>
	<atom:link href="https://onclickinnovations.com/blog/tag/free-ai-model/feed/" rel="self" type="application/rss+xml" />
	<link>https://onclickinnovations.com/blog/tag/free-ai-model/</link>
	<description>Onclick Innovations Pvt. Ltd.</description>
	<lastBuildDate>Wed, 09 Sep 2026 11:32:16 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>
<site xmlns="com-wordpress:feed-additions:1">208843066</site>	<item>
		<title>Ox Alpha: The Mystery AI Model That Topped the Charts &#8212; And What It Turned Out to Be</title>
		<link>https://onclickinnovations.com/blog/ox-alpha-stealth-ai-model-glm-5-3-flash/</link>
					<comments>https://onclickinnovations.com/blog/ox-alpha-stealth-ai-model-glm-5-3-flash/#respond</comments>
		
		<dc:creator><![CDATA[it_geeks]]></dc:creator>
		<pubDate>Wed, 09 Sep 2026 11:10:40 +0000</pubDate>
				<category><![CDATA[AI Development]]></category>
		<category><![CDATA[Industry News]]></category>
		<category><![CDATA[AI coding models]]></category>
		<category><![CDATA[free AI model]]></category>
		<category><![CDATA[GLM-5.3-Flash]]></category>
		<category><![CDATA[open weights]]></category>
		<category><![CDATA[OpenCode]]></category>
		<category><![CDATA[OpenRouter]]></category>
		<category><![CDATA[Ox Alpha]]></category>
		<category><![CDATA[stealth model]]></category>
		<category><![CDATA[Z.ai]]></category>
		<guid isPermaLink="false">https://onclickinnovations.com/blog/?p=1628</guid>

					<description><![CDATA[<p>For about a week in August 2026, one of the most-used AI models on the internet had no company attached to it. No press release, no launch event, no name anyone recognised. It just appeared on the AI marketplace OpenRouter, labelled only as a &#8220;stealth model,&#8221; and it was completely free. Developers started using it. [&#8230;]</p>
<p>The post <a href="https://onclickinnovations.com/blog/ox-alpha-stealth-ai-model-glm-5-3-flash/">Ox Alpha: The Mystery AI Model That Topped the Charts &mdash; And What It Turned Out to Be</a> appeared first on <a href="https://onclickinnovations.com/blog">Blog</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p>For about a week in August 2026, one of the most-used AI models on the internet had no company attached to it. No press release, no launch event, no name anyone recognised. It just appeared on the AI marketplace OpenRouter, labelled only as a &#8220;stealth model,&#8221; and it was completely free.</p>
<p>Developers started using it. Then they started talking about it. Then the speculation about who built it got genuinely feverish &mdash; and the answer, when it came, was more interesting than most of the guesses.</p>
<p>Here&#8217;s the whole story, plus everything now confirmed about what the model actually is.</p>
<h2>What Happened, Chronologically</h2>
<p><strong>August 20, 2026:</strong> A model called &#8220;stealth/ox-alpha&#8221; appeared on OpenRouter, described simply as a reasoning model built for coding, sustained agentic work, and production workloads. OpenRouter&#8217;s own listing was unusually candid: it was &#8220;developed and operated by a third-party provider who has chosen to remain anonymous during this preview,&#8221; and OpenRouter itself was only routing requests to it, not building or owning it.</p>
<p>Two things made people pay attention immediately. It carried a roughly one-million-token context window, and it was free.</p>
<p><strong>The following days:</strong> Independent testing showed genuinely strong coding performance, particularly on front-end and UI work. Stripe CEO Patrick Collison &mdash; whose company was in the process of acquiring OpenRouter &mdash; publicly called it &#8220;very impressive.&#8221; Speculation about the identity of its creator went in every direction: an unreleased GLM model from the Chinese company Z.ai, an unreleased version of Microsoft&#8217;s MAI, or something else entirely. Some people insisted confidently it couldn&#8217;t be Chinese. Others insisted with equal confidence that it was.</p>
<p><strong>August 26, 2026:</strong> Z.ai claimed it. Ox Alpha was GLM-5.3-Flash, and the company released the model publicly &mdash; including open weights on Hugging Face under an MIT licence.</p>
<p>For roughly twelve days, an unnamed model with no company behind it had been near the top of OpenRouter&#8217;s coding charts. Then it got a name.</p>
<h2>The Free Preview: 100 Trillion Tokens a Day, and What Actually Got Used</h2>
<p>The access terms are a big part of why Ox Alpha spread as fast as it did.</p>
<p>The release was coordinated with OpenCode, the open-source coding agent, which announced the model would be free for roughly a week with rate limits described as &#8220;near unlimited&#8221; &mdash; and said the anonymous provider had servicing capacity for <strong>100 trillion tokens per day</strong>. That&#8217;s an aggregate capacity claim about the service, not a per-user quota, but it was an extraordinary thing to announce alongside a model nobody had claimed ownership of.</p>
<p>Developers responded accordingly, and the resulting usage figures are the most concrete evidence of how seriously the model was taken:</p>
<ul>
<li><strong>11.6 trillion tokens in its first three full days on OpenRouter</strong> &mdash; the largest model launch in OpenRouter&#8217;s history. For comparison, the next-biggest launch on record generated 4.4 trillion tokens over the same opening period.</li>
<li><strong>Roughly 6.5 trillion tokens per day on OpenCode</strong> across the first four days, totalling around 26 trillion tokens. That&#8217;s about 6.5% of the advertised 100 trillion daily capacity &mdash; genuinely heavy usage, and still nowhere near the ceiling.</li>
<li><strong>An average of 3.2 million tokens per completed session,</strong> and roughly 25 sessions per unique user. Those aren&#8217;t casual test queries; that&#8217;s people running serious, long agentic coding jobs.</li>
</ul>
<p>The model also supported a maximum output of 131,072 tokens in a single response, alongside its 1,048,576-token context window &mdash; unusually generous limits on both ends.</p>
<h2>The Data Retention Detail Developers Kept Missing</h2>
<p>One practical point got lost in the excitement, and it genuinely mattered for anyone feeding proprietary code into a free anonymous endpoint: <strong>the two routes had different data terms.</strong></p>
<p>OpenCode stated that prompts sent through its service had zero-day retention and were not used for training. OpenRouter&#8217;s own listing, meanwhile, said the anonymous provider retained prompts and completions &mdash; while stating they were not used for training.</p>
<p>Those are meaningfully different guarantees, and &#8220;free Ox Alpha access&#8221; was not interchangeable between them. It&#8217;s a useful reminder in general: when a model is free and its operator is anonymous, checking which specific endpoint is receiving your code is worth the two minutes it takes.</p>
<p>It&#8217;s also worth noting that during the stealth window, Ox Alpha was frequently described as &#8220;open source&#8221; in social media coverage. It wasn&#8217;t. There were no published weights, no technical report, and no licence &mdash; just a free preview of a closed model. The open weights only arrived later, with the GLM-5.3-Flash reveal.</p>
<h2>Why Release a Model Anonymously at All?</h2>
<p>Stealth releases like this have become a recognisable pattern in AI. The logic is straightforward: if nobody knows who built a model, nobody evaluates it through the lens of brand expectations. Reactions are based purely on output quality.</p>
<p>For a lab that isn&#8217;t one of the household-name Western AI companies, that&#8217;s a genuinely valuable form of testing. Developers who might scroll past a model badged with a less familiar name will happily use an anonymous one that performs well &mdash; and the resulting feedback is unfiltered by any preconception. Judging by how much attention Ox Alpha attracted in under two weeks, the strategy worked.</p>
<h2>What GLM-5.3-Flash Actually Is</h2>
<p>Now that the model has a name, here&#8217;s the confirmed picture:</p>
<ul>
<li><strong>Architecture:</strong> A sparse mixture-of-experts model with 320 billion total parameters, but only 18 billion active for any given token. It uses a hybrid sparse-plus-linear attention design that Z.ai built specifically to keep long-context work from becoming prohibitively expensive to serve.</li>
<li><strong>Context window:</strong> Roughly 1 million tokens &mdash; large enough to hold entire codebases or lengthy specifications in a single request without chunking or retrieval workarounds.</li>
<li><strong>Multimodal, natively.</strong> This is the first natively multimodal model in Z.ai&#8217;s GLM-5 family, accepting text and images (and, per OpenRouter&#8217;s listing during the preview, video), and returning text.</li>
<li><strong>Open weights.</strong> Released on Hugging Face under an MIT licence, which is unusually permissive &mdash; you can self-host and use it commercially.</li>
<li><strong>Tool calling and structured output</strong> are both supported, which matters for anyone building agents rather than chatbots.</li>
</ul>
<p>The architectural detail worth understanding is that active-parameter count, because it explains the pricing. Only 18 billion of the model&#8217;s 320 billion parameters activate per token, which is why a very large model can be served at roughly the cost of a small one.</p>
<h2>The Benchmarks &mdash; With the Caveats That Matter</h2>
<p>Z.ai&#8217;s own launch numbers put GLM-5.3-Flash at 84.3 on Terminal-Bench 2.1, against Claude Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4. On the company&#8217;s in-house Code Bench, it reached 29.0 at maximum effort versus 29.5 for Opus 4.8. Those two comparisons are the source of essentially every &#8220;matches Claude at a tenth of the price&#8221; headline written about it.</p>
<p>Two important caveats before taking that at face value. First, most of that comparison table is Z.ai&#8217;s own evaluation, using comparison models and settings the company selected. That doesn&#8217;t make it dishonest, but a vendor launch table is not the same as independent verification.</p>
<p>Second, and more usefully: the DeepSWE result did hold up independently. Z.ai self-reported 63.4 Pass@1 on DeepSWE v1.1, and the official DeepSWE leaderboard subsequently carried a matching entry at roughly 63%. That&#8217;s worth noting specifically because vendor benchmark claims holding up under independent testing is rarer than it should be.</p>
<p>Independent evaluations from Artificial Analysis place it around 57 on their Intelligence Index, though published figures from different sources and snapshots vary somewhat, so treat any single index number as approximate rather than definitive.</p>
<p>One number worth correcting, since it circulated widely: during the stealth window, an eye-catching &#8220;80% on DeepSWE&#8221; figure spread rapidly across social media, apparently beating both GPT-5.6 Sol and Claude Fable. That result came from a single developer running a subset of just ten tasks, and they flagged the high variance themselves at the time. The verified full-benchmark number is roughly 63% &mdash; still a very strong result, and a large jump over the previous GLM generation, but not the frontier-beating figure the early screenshots suggested.</p>
<h2>Where It&#8217;s Genuinely Strong</h2>
<p>Hands-on testing across the developer community converged on a few consistent strengths:</p>
<ul>
<li><strong>Front-end and UI generation.</strong> This came up repeatedly. Testers described output with clean typography, restrained colour use, and consistent spacing across long pages &mdash; noticeably less of the generic &#8220;AI-looking&#8221; layout that most models default to.</li>
<li><strong>Long-horizon agentic coding.</strong> Multi-step software engineering tasks that unfold over many turns, rather than single-shot code generation.</li>
<li><strong>Long-context work.</strong> The million-token window combined with attention architecture built specifically to handle it means large codebases stay genuinely usable in one session.</li>
<li><strong>Price-to-performance.</strong> This is the headline. Near-frontier agentic coding at roughly a tenth of the cost of comparable models is the entire pitch, and it largely holds up.</li>
</ul>
<h2>Where It Isn&#8217;t</h2>
<p>Being honest about the trade-offs matters more than the marketing:</p>
<ul>
<li><strong>It&#8217;s slow.</strong> Independent measurements put output speed around 50 tokens per second, which is below average. Despite the name, &#8220;Flash&#8221; refers to cost efficiency, not speed.</li>
<li><strong>It&#8217;s verbose.</strong> Multiple independent evaluations flag this specifically. Verbosity partly offsets the cheap per-token pricing, since you&#8217;re paying for more tokens.</li>
<li><strong>Self-hosting is not lightweight.</strong> The open weights are genuinely open, but the FP8 checkpoint runs to roughly 306 GiB. That rules out running it on modest local hardware, regardless of the MIT licence.</li>
<li><strong>It&#8217;s not the flagship.</strong> Don&#8217;t confuse GLM-5.3-Flash with the full GLM-5.3, which is a separate, more expensive, text-only model aimed at coding and cybersecurity work. Flash is the cheap, multimodal, high-throughput sibling.</li>
</ul>
<h2>Pricing</h2>
<p>Z.ai&#8217;s list pricing is $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens.</p>
<p>A 50% launch promotion has been running that halves those rates to $0.075, $0.25 and $0.015 respectively &mdash; but that promotion was scheduled to end on September 9, 2026, so check current pricing before planning around the discounted rate.</p>
<p>For context, Z.ai&#8217;s flagship GLM-5.3 is listed at roughly $1.40 input and $4.40 output per million tokens. Flash genuinely is about a tenth of the price of its bigger sibling.</p>
<h2>How to Actually Use It</h2>
<ul>
<li><strong>Via OpenRouter</strong> using the model ID <code>z-ai/glm-5.3-flash</code>. OpenRouter routes across many providers with automatic failover, and you can pin or exclude specific providers.</li>
<li><strong>Directly through Z.ai&#8217;s API,</strong> where the model code is <code>glm-5.3-flash</code>.</li>
<li><strong>Self-hosted</strong> from the MIT-licensed weights on Hugging Face &mdash; viable only if you have serious GPU infrastructure, given the checkpoint size.</li>
<li><strong>Note that the <code>stealth/ox-alpha</code> route is no longer the way in.</strong> That was the temporary preview identity, and the free window has closed.</li>
</ul>
<h2>What This Episode Actually Tells Us</h2>
<p>The Ox Alpha saga is a small story with a couple of genuinely significant implications.</p>
<p>The first is about how quickly the cost curve is moving. A model delivering near-frontier agentic coding performance, with a million-token context and native multimodality, at roughly a tenth of flagship pricing, with open weights &mdash; that combination didn&#8217;t exist as an option a year ago. For teams building AI features where cost per call actually determines whether the product is viable, that changes the maths.</p>
<p>The second is about how blind evaluation works. For twelve days, thousands of developers judged this model purely on output, with no brand attached. It performed well enough that a major tech CEO praised it publicly and speculation ran wild about which Western lab must have built it. When the answer turned out to be a Chinese lab, that reaction had already been recorded without any of the usual filters applied.</p>
<p>Both of those are worth sitting with if your assumption is that the best AI models will always come from the companies you&#8217;d expect, at the prices you&#8217;d expect.</p>
<h2>Frequently Asked Questions</h2>
<p><strong>What is Ox Alpha?</strong><br />
Ox Alpha was the anonymous preview codename for an AI model that appeared on OpenRouter on August 20, 2026, offered free with a roughly one-million-token context window and multimodal input. On August 26, 2026, it was revealed to be Z.ai&#8217;s GLM-5.3-Flash.</p>
<p><strong>Is Ox Alpha still available?</strong><br />
Not under that name. The stealth preview and its free access window have ended. The model is now available publicly as GLM-5.3-Flash, through OpenRouter, Z.ai&#8217;s own API, and as open weights on Hugging Face.</p>
<p><strong>How much free usage did Ox Alpha offer?</strong><br />
The model was free for roughly one week from August 20, 2026, with rate limits described as &#8220;near unlimited.&#8221; OpenCode, which distributed it alongside OpenRouter, said the anonymous provider had servicing capacity for 100 trillion tokens per day &mdash; an aggregate service capacity claim rather than an individual user quota.</p>
<p><strong>How much was Ox Alpha actually used?</strong><br />
It generated 11.6 trillion tokens in its first three full days on OpenRouter, the largest model launch in that platform&#8217;s history &mdash; the previous record was 4.4 trillion over the same period. On OpenCode, usage averaged roughly 6.5 trillion tokens per day across the first four days, with an average of 3.2 million tokens per session.</p>
<p><strong>Was Ox Alpha open source during the free preview?</strong><br />
No. Despite being widely described that way, the stealth preview had no published weights, technical report, or licence &mdash; it was a free preview of a closed model. Open weights were only released later, under the MIT licence, when Z.ai revealed the model as GLM-5.3-Flash.</p>
<p><strong>Who made Ox Alpha?</strong><br />
Z.ai (also known as Zhipu AI), a Chinese AI company. During the anonymous preview, speculation ranged from Z.ai&#8217;s GLM series to an unreleased Microsoft model, with no consensus until Z.ai confirmed it.</p>
<p><strong>How much does GLM-5.3-Flash cost?</strong><br />
List pricing is $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens. A 50% launch promotion halved those rates but was scheduled to end September 9, 2026, so verify current pricing.</p>
<p><strong>Is GLM-5.3-Flash good for coding?</strong><br />
It performs strongly on coding and agentic benchmarks, scoring 84.3 on Terminal-Bench 2.1 against Claude Opus 4.8&#8217;s 85.0, with its DeepSWE result of roughly 63% independently verified. Testers particularly praised its front-end and UI generation. The main trade-offs are below-average output speed and notable verbosity.</p>
<p><strong>Can I run GLM-5.3-Flash myself?</strong><br />
Yes in principle &mdash; the weights are on Hugging Face under an MIT licence. In practice, the FP8 checkpoint is around 306 GiB, so self-hosting requires substantial GPU infrastructure rather than a local workstation.</p>
<p><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Fonclickinnovations.com%2Fblog%2Fox-alpha-stealth-ai-model-glm-5-3-flash%2F&amp;linkname=Ox%20Alpha%3A%20The%20Mystery%20AI%20Model%20That%20Topped%20the%20Charts%20%E2%80%94%20And%20What%20It%20Turned%20Out%20to%20Be" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_twitter" href="https://www.addtoany.com/add_to/twitter?linkurl=https%3A%2F%2Fonclickinnovations.com%2Fblog%2Fox-alpha-stealth-ai-model-glm-5-3-flash%2F&amp;linkname=Ox%20Alpha%3A%20The%20Mystery%20AI%20Model%20That%20Topped%20the%20Charts%20%E2%80%94%20And%20What%20It%20Turned%20Out%20to%20Be" title="Twitter" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_linkedin" href="https://www.addtoany.com/add_to/linkedin?linkurl=https%3A%2F%2Fonclickinnovations.com%2Fblog%2Fox-alpha-stealth-ai-model-glm-5-3-flash%2F&amp;linkname=Ox%20Alpha%3A%20The%20Mystery%20AI%20Model%20That%20Topped%20the%20Charts%20%E2%80%94%20And%20What%20It%20Turned%20Out%20to%20Be" title="LinkedIn" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_no_icon a2a_counter addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Fonclickinnovations.com%2Fblog%2Fox-alpha-stealth-ai-model-glm-5-3-flash%2F&#038;title=Ox%20Alpha%3A%20The%20Mystery%20AI%20Model%20That%20Topped%20the%20Charts%20%E2%80%94%20And%20What%20It%20Turned%20Out%20to%20Be" data-a2a-url="https://onclickinnovations.com/blog/ox-alpha-stealth-ai-model-glm-5-3-flash/" data-a2a-title="Ox Alpha: The Mystery AI Model That Topped the Charts — And What It Turned Out to Be">Share</a></p><p>The post <a href="https://onclickinnovations.com/blog/ox-alpha-stealth-ai-model-glm-5-3-flash/">Ox Alpha: The Mystery AI Model That Topped the Charts &mdash; And What It Turned Out to Be</a> appeared first on <a href="https://onclickinnovations.com/blog">Blog</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://onclickinnovations.com/blog/ox-alpha-stealth-ai-model-glm-5-3-flash/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">1628</post-id>	</item>
	</channel>
</rss>
