
Qwen3.8 is Alibaba's new 2.4-trillion-parameter flagship, announced on July 19, 2026, with open weights promised "soon" and a Max-Preview you can call today. What it doesn't have is a single published benchmark. Alibaba positions it as "second only to Fable 5" — that's the vendor's own claim, not a measured result, and the difference matters more than usual here.
What Alibaba Actually Announced
Everything verified about Qwen3.8 traces back to one source: Alibaba's announcement on July 19, 2026. Here's what it actually says:
- Qwen3.8 exists, with 2.4 trillion parameters
- The weights are going open "soon" — no date attached
- Qwen3.8-Max-Preview is live now
- Access runs through Alibaba's Token Plan, Qoder, and QoderWork
- Alibaba describes the model as "one of the most powerful models available today, compatible to leading frontier AI models, second only to Fable 5"
That last bullet is positioning, not a result. Keep the two in separate buckets.
The timing is worth a beat. Qwen3.8 is the fourth frontier launch in eleven days — Grok 4.5 on July 8, GPT-5.6 on July 9, now this. It also lands four days after Chinese regulators approved the Apple–Alibaba Qwen integration on July 15, so attention on the Qwen brand was already running hot when the announcement went out.
Qwen3.8 vs Qwen3-8B: Not the Same Model
Worth clearing up early, because search engines currently conflate the two. Qwen3.8 is Alibaba's 2.4-trillion-parameter flagship announced today. Qwen3-8B is a dense 8-billion-parameter model from the Qwen3 series, downloadable on Hugging Face since 2025. Similar-looking names, about 300x apart in size.
What We Know vs What We Don't (Yet)
Nobody knows how good Qwen3.8 is yet, including the posts telling you they do. So here are three clearly separated buckets.
Qwen3.8 Confirmed Facts (Alibaba official, 2026-07-19)
| Qwen3.8 fact | What Alibaba confirmed |
|---|---|
| Model announced | Qwen3.8 |
| Parameter count | 2.4 trillion |
| Open weights | Promised, "going open-weight soon" |
| Preview availability | Qwen3.8-Max-Preview live now |
| Access surfaces | Token Plan, Qoder, QoderWork |
| Development status | "Continuously evolving," per Alibaba |
| Vendor positioning | "Second only to Fable 5" — Alibaba's claim, not an independent result |
Qwen3.8 Benchmarks, Pricing and Specs: What's Not Published
| Qwen3.8 unknown | Status as of July 19, 2026 |
|---|---|
| Any Qwen3.8 benchmark score | None published |
| Open-weight release date | "Soon," no date |
| License | Unannounced |
| Per-token pricing | Not broken out publicly |
| Active parameters / MoE config | Not disclosed |
| Context window and max output | Not announced |
| Modality (multimodal?) | Not announced |
| Hugging Face model card | Verified absent |
| OpenRouter listing | Verified absent (76 Qwen models listed, none is 3.8) |
| Regional availability (EU/US) | Unstated |
| Independent evaluation | None yet — expect within days |
No model card and no OpenRouter entry is the hardest evidence available that the open weights genuinely have not shipped.
Labelled Proxy: Qwen3.7-Max (Predecessor — NOT Qwen3.8)
Every number below belongs to Qwen3.7-Max, the previous generation. None of it is a Qwen3.8 result, and none of it is a forecast. It's here so you know what kind of lab is making the claim.
| Qwen3.7-Max metric (announced 2026-05-20) | Value |
|---|---|
| Artificial Analysis Intelligence Index v4.0 | 56.6 — 5th overall, #1 Chinese model (Qwen3.7-Max figure) |
| Context / max output | 1M tokens / 65,536 (Qwen3.7-Max figure) |
| Pricing | $2.50 in / $7.50 out per M, 90% cached-input discount (Qwen3.7-Max figure) |
| Weights | Closed, API-only — no GGUF, no HF checkpoint (Qwen3.7-Max) |
| SWE-bench Verified / SWE-Pro / SWE-Multilingual | 80.4 / 60.6 / 78.3 (Qwen3.7-Max figures) |
| SciCode / Terminal Bench 2.0-Terminus | 53.5 / 69.7 (Qwen3.7-Max figures) |
| Hallucination rate | 22.9%, lowest among frontier models (Qwen3.7-Max figure) |
| Autonomous operation | Up to ~35 hours (Qwen3.7-Max figure) |
| API compatibility | OpenAI and Anthropic specs (Qwen3.7-Max) |

2.4 Trillion Parameters: The Largest Open-Weight Model Ever Promised
Qwen3.8 has 2.4 trillion parameters, per Alibaba's announcement. If Alibaba ships those weights, that makes it roughly 1.5x larger than any open-weight model released to date — the current record holder is DeepSeek V4 Pro at 1.6T. Everything else in the open field sits well below a trillion.
Largest open-weight models by parameter count (2026)
Data table
| Parameters (billions) | Parameters (B) |
|---|---|
| Qwen3.8 (promised) | 2400 |
| DeepSeek V4 Pro | 1600 |
| Kimi K2.7 | 1000 |
| GLM-5.2 | 744 |
| DeepSeek-V3.2 | 685 |
One caveat before you get excited about that bar. Alibaba hasn't disclosed the active-parameter count or the mixture-of-experts configuration, so 2.4T is a headline number, not a compute figure. For scale, DeepSeek V4 Pro's 1.6T only activates about 49B parameters per token — roughly 3% of the network — which is why it runs at anything like a sane cost. A sparse 2.4T model and a dense 2.4T model are completely different animals to serve, and right now we don't know which one this is. If you want the current state of the field, we track where Qwen already ranks among open-source LLMs, and GLM 5.2, the other big Chinese open-weight release, gives you a sense of what this tier looks like when it actually ships.
Qwen3.8 Open Weights: The Reversal Nobody's Talking About
Alibaba has promised open weights for Qwen3.8, but hasn't announced a date or a license. That sounds routine until you look at the family history — which is exactly what most day-one coverage will skip.
Alibaba spent two generations building a deliberate split: open-weight workhorses for everyone, closed flagship at the top.
| Generation | Weights | Notes |
|---|---|---|
| Qwen 3.5 (0.8B–397B-A17B) | Open, Apache 2.0 | Full family released publicly |
| Qwen 3.6 (27B dense, 35B-A3B MoE) | Open, Apache 2.0 | Same pattern held |
| Qwen3.7-Max | Closed | DashScope API-only. No GGUF, no Hugging Face checkpoint |
Then Qwen3.8 promises open weights at the frontier tier. Read that sequence again: Alibaba closed the flagship tier for two straight generations, and is now saying it will open the biggest model it has ever built.
If it happens, it's a genuine strategy reversal, and it's the actual news story here — not the parameter count. It also puts real pressure on every closed-weight lab that's been treating "frontier tier stays closed" as a settled industry norm.
The honest hedge: the license is unannounced. Apache 2.0 on 3.5 and 3.6 is encouraging precedent, and precedent is not a commitment. A restrictive community license at 2.4T would technically satisfy "open weights" while changing what you can actually do with it.

Can You Actually Run Qwen3.8 Yourself? (Honest Answer: No)
No — not on consumer hardware, and realistically not on your startup's GPU budget either, even after the weights drop. Even DeepSeek-V3.2 at 685B is described by the people who've done it as a real infrastructure project to self-host. Qwen3.8 is roughly 3.5x that.
So "open weights" at 2.4T does not mean you run this. What it actually gets you:
- Inference providers can host it, which means price competition, which means your cost per token falls
- Sovereign, regulated, and air-gapped deployments become possible for the first time at this tier
- Researchers and fine-tuners get access to a frontier-scale base model
- It does not mean it runs on your machine, your single A100, or realistically your startup's GPU budget
Right now the Preview lives only on Alibaba-operated surfaces, and under China's National Intelligence Law, Alibaba is obliged to cooperate with government data requests. Qoder specifically has drawn security scrutiny from Western reviewers.
The nuance nobody joins up: that concern attaches to the hosted API, not to self-hosted weights running on US or EU infrastructure. Which is exactly why the open-weight promise, not the Preview, is the consequential half of this announcement for anyone with data-residency obligations.
If you're planning around that, our roundup of frameworks for building on open-weight models covers the orchestration layer you'd want in place before the weights land.
How to Try Qwen3.8-Max-Preview Today
Qwen3.8-Max-Preview is live now. Three routes in, all Alibaba-operated:
- Token Plan — Alibaba's bundled API access
- Qoder — Alibaba's coding surface
- QoderWork — Alibaba's agent surface
What's Available Now vs What's Coming
| Stage | Status | What you actually get | Who it fits |
|---|---|---|---|
| Now — Qwen3.8-Max-Preview | Live (2026-07-19) | API access via Token Plan, Qoder, QoderWork. Per-token rate not broken out publicly. No OpenRouter listing, no HF card, no confirmed DashScope GA | Teams already on an agent harness who want frontier-ish coding without Fable 5 access, and who accept an Alibaba-hosted endpoint |
| Soon — open weights | Promised, no date | Downloadable checkpoints. License unannounced. Realistically consumed via inference providers, not local hardware | Self-hosters with serious infra, fine-tuners, sovereign or data-residency-constrained orgs |
| Not available | — | Independent benchmarks, published pricing, confirmed context window, confirmed modality, stated EU/US availability | Anyone who needs to justify a migration on numbers — wait |
Calling It From Code
Qwen3.7-Max spoke both the OpenAI and Anthropic API specs, so a base-URL swap was all it took. Qwen3.8's compatibility is unconfirmed — treat this as the pattern to verify, not a working endpoint.
# Qwen3.7-Max-era pattern. Qwen3.8-Max-Preview's API surface and model ID
# are UNCONFIRMED - check current Alibaba Cloud docs before shipping this.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DASHSCOPE_API_KEY"],
base_url="https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
)
resp = client.chat.completions.create(
model="qwen3.8-max-preview",
messages=[{"role": "user", "content": "Refactor this function for readability."}],
)
print(resp.choices[0].message.content)Because it's an OpenAI-shaped client, it drops into most harnesses unchanged — including the coding agents you'd actually plug this into. If you want it reaching your own systems, you can expose your own tools through an MCP server and point the agent at that.
Qwen3.8 API Pricing: What It Might Cost
No Qwen3.8 price is published. The only defensible reference point is the predecessor: $2.50 in / $7.50 out per million tokens with a 90% cached-input discount (Qwen3.7-Max figure). A similarly-priced 3.8 would land at roughly half of GPT-5.6 Sol and a quarter of Fable 5. That's a reference point, not a forecast — and read the next section before you build a business case on it.

Qwen3.8 vs Claude Fable 5: How Seriously Should You Take Alibaba's Claim?
Alibaba says Qwen3.8 is "second only to Fable 5." No published benchmark backs that up — it's positioning, not a measured result. Here's the fair reading in both directions.
The case for taking it semi-seriously. Qwen3.7-Max scored an independently measured 56.6 on the Artificial Analysis Intelligence Index — 5th overall and the top-ranked Chinese model (both Qwen3.7-Max figures). A lab with that predecessor saying its next model is second only to Fable 5 isn't making an obviously absurd claim — that's a different situation from a first-time entrant announcing frontier parity.
The case against. It's vendor-reported, with no named benchmark, no methodology, and no independent evaluation. And the Fable 5 benchmark bar Alibaba is aiming at is 80.4 on SWE-Bench Pro — which is not the same benchmark as Qwen3.7-Max's 80.4 on SWE-bench Verified. Identical number, different test, different difficulty. On SWE-Bench Pro — the one benchmark they actually share — Qwen3.7-Max scored 60.6 (Qwen3.7-Max figure), roughly 20 points behind. Anyone presenting the two 80.4s as a tie is either confused or hoping you are. We saw the same pattern recently with the last model to claim frontier parity on vendor numbers.
One gotcha for your cost model: Qwen models tend to emit more output tokens per task than their peers. Qwen3.5-27B burned 98M output tokens completing the Artificial Analysis Intelligence Index, against 56M for MiniMax-M2.5 and 61M for DeepSeek V3.2. Output tokens are the expensive side, so headline $/M comparisons systematically overstate Qwen savings. That's a 3.5-generation observation, not a Qwen3.8 measurement — but it's the kind of thing that quietly doubles a bill.
Qwen3.8 vs GPT-5.6 and Grok 4.5: Where It Would Land
Western launch coverage has been ignoring the Chinese challengers almost entirely. One widely-shared piece on the "chaotic week" that redrew the AI map covered GPT-5.6, Grok 4.5 and Fable 5 — and mentioned Qwen zero times. Meanwhile DeepSeek V4-Pro, Qwen3.7-Max and Kimi K2.6 all land within half a point of each other on SWE-bench Verified, roughly 14 points behind Claude and priced about two orders of magnitude below it.
| Model | Released | Price (in/out per M) | Headline coding result |
|---|---|---|---|
| Claude Fable 5 | 2026-06-09 | $10 / $50 | SWE-Bench Pro 80.4; DeepSWE 1.1 70% |
| GPT-5.6 Sol | 2026-07-09 | $5 / $30 | Agents' Last Exam 53.6; AA Coding Agent Index 80 |
| Grok 4.5 | 2026-07-08 | $2 / $6 | SWE-Bench Pro 64.7% |
| Claude Opus 4.8 | — | — | SWE-Bench Pro 69.2% |
| Qwen3.8 | 2026-07-19 (preview) | Not published | Not yet published |
| Qwen3.7-Max (predecessor) | 2026-05-20 | $2.50 / $7.50 | SWE-bench Verified 80.4; AA Index 56.6 |
Note the benchmark names in that table. They aren't interchangeable, and a chunk of this week's coverage will treat them as if they are. For the Claude side of the comparison, we've broken down how Opus 4.8 scores on the same coding benchmarks.
Should You Switch, Test, or Wait?
| Your situation | What to do |
|---|---|
| Already running Qwen in production | Test the Max-Preview on your own eval set this week. You have the harness and the baseline — you're one of the few people who can generate a real Qwen3.8 number right now |
| Evaluating Chinese models on cost | Wait for independent benchmarks. They're days away. Then redo the math with the verbosity caveat applied, not the headline $/M |
| Happy on Claude or GPT | Nothing to do today. Set a reminder for the open-weight drop — that's the event that could change your cost structure, not the preview |
The method matters more than the verdict here. Every model we've production-tested has performed differently from its leaderboard position, because your prompts, your codebase and your tolerance for retries aren't in anyone's benchmark. Your own eval set beats a vendor claim, and it beats a leaderboard too. Building one takes an afternoon and pays for itself the first time a "better" model quietly regresses on your actual workload.
Not sure what to measure, or want a second pair of eyes on the migration math? Have our team pressure-test it on your workload →
Frequently Asked Questions
Is Qwen3.8 open source?
Not yet, and "open weights" isn't the same thing as open source. Alibaba has promised downloadable weights, but the license hasn't been announced. Qwen 3.5 and 3.6 both shipped under Apache 2.0, which is encouraging precedent — but precedent isn't a commitment, and no Qwen3.8 license text exists today.
When will Qwen3.8 open weights be released?
Alibaba said "soon" and gave no date. That's the entire answer, and anyone naming a specific week is guessing. The clearest signal that it hasn't shipped: there's still no Qwen3.8 model card on Hugging Face as of July 19, 2026.
How many parameters does Qwen3.8 have?
2.4 trillion, per Alibaba's announcement. What's missing is the active-parameter count and the mixture-of-experts configuration, so you can't derive compute cost per token from that headline figure. A sparse 2.4T model and a dense 2.4T model are very different things to serve.
Can I run Qwen3.8 locally?
Realistically no, even after the weights drop. DeepSeek-V3.2 at 685B is already a serious infrastructure project to self-host, and Qwen3.8 is roughly 3.5x that size. Expect to consume it through inference providers rather than your own GPUs, unless you're running a genuine datacenter.
How much does Qwen3.8-Max-Preview cost?
Alibaba hasn't broken out a per-token rate — access is bundled into Token Plan, Qoder, and QoderWork. For reference only, the predecessor Qwen3.7-Max runs $2.50 in / $7.50 out per million tokens with a 90% cached-input discount (Qwen3.7-Max figure). Don't budget off that number.
Is Qwen3.8 better than Claude Fable 5?
Nobody knows yet. "Second only to Fable 5" is Alibaba's own positioning, published without a named benchmark, a methodology, or an independent evaluation. Fable 5's independently measured bar is 80.4 on SWE-Bench Pro. Until third-party evals land, treat the comparison as genuinely open.
Is Qwen3.8 on OpenRouter or Hugging Face yet?
No — verified as of July 19, 2026. OpenRouter lists 76 Qwen models and none of them is 3.8, and there's no Hugging Face model card either. That absence is the most reliable evidence available that the open weights genuinely haven't shipped.
What are Qoder and QoderWork?
They're Alibaba's own coding and agent surfaces, and currently two of the three routes to Qwen3.8-Max-Preview alongside the Token Plan. Worth knowing before you commit: Qoder has drawn security scrutiny from Western reviewers, which matters if your code touches regulated data.
What context window does Qwen3.8 have?
Not announced. The predecessor Qwen3.7-Max offered 1M tokens of context with 65,536 max output (Qwen3.7-Max figure), so something similar is plausible — but that's a predecessor spec, not a Qwen3.8 claim. Don't design a long-context pipeline around it yet.
Key Takeaways
- 2.4T parameters and an open-weight promise are confirmed. Benchmarks are not — there isn't a single published Qwen3.8 score, and any number you see attached to it this week is almost certainly a Qwen3.7-Max figure.
- The strategy reversal is the real story. Alibaba closed the flagship tier for two generations, then promised to open the biggest model it has built.
- At 2.4T, open weights means cheaper hosted inference and sovereign deployment — not running it yourself.
- "Second only to Fable 5" is Alibaba's claim. Independent numbers land within days. Wait for them.
Weighing an open-weight model against your current stack, or trying to work out what any of this changes for your infrastructure? Get a free LLM stack review →