Qwen 3.8: How a 27B Open Model Rivals GPT-5.6 and Claude Opus

Date:

A 27-billion-parameter model that fits on a single consumer graphics card is now trading benchmark blows with flagship models from OpenAI and Anthropic that are dozens of times larger. That is the story of Qwen 3.8, the open-weight model family Alibaba released in August 2026.

The lineup spans two released checkpoints: Qwen 3.8-27B, a dense vision-language model aimed at local hardware, and Qwen 3.8-2.4T-A95B, a 2.4-trillion-parameter mixture-of-experts flagship. Both undercut comparable closed models on price. Below we break down the benchmarks, the pricing, how to run the 27B at home, and what is coming next. Note that the closed API flagship has since moved on: the September 0902 update to Qwen3.8 Max lifted its independent benchmark score from 40 to 45, at more than twice the cost per task.

What Is Qwen 3.8?

Qwen 3.8 is the newest generation of open-weight models from Alibaba’s Qwen team. It was previewed on July 19, 2026 at the World AI Conference in Shanghai, and the first two checkpoints went public within weeks.

Model Type Parameters Released License
Qwen 3.8-27B Dense vision-language 27.8B (28B with encoder) Aug 14, 2026 Apache 2.0
Qwen 3.8-2.4T-A95B Sparse mixture-of-experts 2.4T total / 95B active Aug 8, 2026 Custom open weights
Qwen 3.8-Max MoE (API only) 2.4T total / 95B active Aug 3, 2026 Closed API

The 27B is the headline local model: a dense transformer with native text, image, and video input, a 262,144-token context window that can extend to roughly one million tokens, and a switchable thinking mode. The 2.4T model is the first Qwen Max-class model ever released with open weights, a meaningful shift for a company that has historically kept its flagship tier behind an API.

One distinction is worth flagging. The open-weights 2.4T checkpoint is text-only and always reasons, while the API “Max” version adds vision input, a non-thinking mode, and a one-million-token context by default. If you download the 2.4T weights, you get a narrower model than the one served through QwenCloud.

Qwen 3.8-27B Benchmarks: Punching Far Above Its Weight

At 27.8 billion parameters, Qwen 3.8-27B is small enough to run on a single 24 GB graphics card, yet its published scores sit close to, and in several cases above, Anthropic’s Opus-class models, which are far larger.

Bar chart comparing Qwen 3.8 27B and Claude Opus 4.6 Max scores on five benchmarks
(Credit: Intelligent Living)
Benchmark Qwen 3.8-27B Claude Opus 4.6 Max
SWE-bench Pro (coding) 61.7 53.4
QwenSWEBench 79.0 63.8
LiveCodeBench v6 90.3 88.8
OSWorld-Verified (computer use) 84.3 72.7
AndroidWorld 81.9 62.0
CoWorkBench 70.7 68.2
Terminal-Bench 2.1 73.0 78.2
GPQA Diamond 89.2 91.3
Humanity’s Last Exam 30.8 40.0

On real-world coding and computer-use tasks, the 27B leads. It scores 61.7 on SWE-bench Pro versus Opus 4.6 Max’s 53.4, 90.3 on LiveCodeBench v6 versus 88.8, and 84.3 on OSWorld-Verified versus 72.7. It falls behind on a handful of harder reasoning and terminal benchmarks, including Terminal-Bench 2.1 (73.0 vs 78.2), GPQA Diamond (89.2 vs 91.3), and Humanity’s Last Exam (30.8 vs 40.0).

One important caveat: most of these figures are Alibaba’s own, published on the official model card. Independent third-party benchmarks are still catching up, so the exact margins deserve some skepticism. The broader pattern, a 27B open model trading blows with much larger closed flagships, is the part that is hard to dismiss.

Notably, the 27B is a dense model, not a mixture-of-experts. There is no sparse-activation shortcut behind these numbers; every one of the 27.8 billion parameters participates in every token. That makes the comparison to much larger closed models all the more striking.

How the 27B compares to Claude Sonnet 5

Anthropic’s Claude Sonnet 5, released June 30, 2026, is the newer mid-tier model in the lineup. It is a much larger closed model than Qwen 3.8-27B, yet it barely edges the 27B on some benchmarks and trails it on others.

Benchmark Qwen 3.8-27B Claude Sonnet 5
SWE-bench Pro 61.7 63.2
OSWorld-Verified 84.3 81.2
Terminal-Bench 2.1 73.0 80.4

Sonnet 5 leads SWE-bench Pro by just 1.5 points (63.2 vs 61.7), while the 27B actually wins OSWorld-Verified (84.3 vs 81.2). Sonnet 5 only pulls clearly ahead on Terminal-Bench 2.1 (80.4 vs 73.0). A 27-billion-parameter model that can run at home is, on several axes, genuinely competitive with a flagship-class closed model.

The Qwen 3.8-2.4T-A95B Flagship

The top of the lineup is Qwen 3.8-Max, built on the open-weights Qwen 3.8-2.4T-A95B checkpoint. It is a sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active per token, routing each token through 11 of 512 experts. That “A95B” suffix is the key to the whole design.

Architecturally, it is a 92-layer stack that interleaves Gated DeltaNet linear-attention blocks with Gated Attention blocks, each followed by a mixture-of-experts layer. That hybrid design, inherited from the Qwen 3.5 architecture, is what lets the model sustain a 262,144-token native context window without the compute cost ballooning.

Only a small fraction of the weights fire for any given token, which is what makes a 2.4-trillion-parameter model economically feasible to serve. Alibaba’s published benchmark table puts it at or near the top of the frontier on several tasks.

Benchmark Qwen 3.8-Max Claude Opus 4.8 GPT-5.6 Sol
Terminal-Bench 2.1 86.6 84.6 88.8
SWE-bench Pro 67.7 69.2 64.6
DeepSWE 1.1 56.6 59.0 73.0
PaperBench 93.0 80.3 90.5
AndroidBench 75.1 69.8 74.0

The standout is PaperBench, where Qwen 3.8-Max posts 93.0, ahead of GPT-5.6 Sol (90.5) and Claude Opus 4.8 (80.3). It also leads Terminal-Bench 2.1 at 86.6, ahead of Opus 4.8 (84.6) though still behind GPT-5.6 Sol (88.8). On the hardest agentic coding test, DeepSWE 1.1, it trails the leaders. The picture is not a clean sweep, but it is a clear arrival at the frontier for an open-weights model.

Why Qwen 3.8 Undercuts OpenAI and Anthropic on Price

Qwen 3.8-Max lists at $2 per million input tokens and $6 per million output tokens on QwenCloud, with cached input at $0.25. There is a single flat tier across the full one-million-token context window, with no long-prompt surcharge. That is dramatically cheaper than the closed-model flagships it competes with.

Bar chart comparing API pricing per million tokens for Qwen 3.8 Max, Claude Opus 4.8, Claude Sonnet 5, and GPT-5.6 Terra
(Credit: Intelligent Living)
Model Input ($/1M) Output ($/1M)
Qwen 3.8-Max $2.00 $6.00
Claude Opus 4.8 $5.00 $25.00
Claude Sonnet 5 $2.00 $10.00
GPT-5.6 Terra $2.00 $12.00

Against Claude Opus 4.8, Qwen 3.8-Max is about 60 percent cheaper on input and 76 percent cheaper on output. It matches Claude Sonnet 5’s pricing on input but costs 40 percent less on output. The gap traces directly back to the sparse architecture: with only 95 billion of 2.4 trillion parameters active per token, each token requires far less compute than a comparable dense model, and those savings flow through to the rate card.

This fits a broader pattern of Chinese open-weight models consistently undercutting Western rivals on price, with the cost of inference rather than raw parameter count deciding what a model actually costs to run. The same week Qwen 3.8 shipped, Microsoft’s MAI-Image-2.6 Preview took No. 2 on the Arena text-to-image leaderboard under a different distribution strategy (free Playground access and a paid Foundry tier) rather than a cheaper rate card. DeepSeek extended that same logic to vision this week, launching image input on V4-Flash with no multimodal price premium. For a side-by-side view of the whole market, see the LLM API pricing guide.

Running Qwen 3.8-27B on Your Own Hardware

The 27B model is where Qwen 3.8 becomes genuinely personal. It ships under the permissive Apache 2.0 license, so commercial use is free, and community quantizations from the Unsloth project bring it within reach of consumer GPUs. The compression path has since gone further still, with Bonsai 2 27B converting those same weights to ternary form and running in just 5.9 GB.

A laptop on a home desk running a local AI model with a glowing neural network visualization
(Credit: Intelligent Living)
Format Size Practical VRAM
BF16 (original) ~55.6 GB 80 GB class
FP8 ~30.9 GB 48 GB class
Q5_K_M (GGUF) ~19.8 GB 24 GB
Q4_K_M (GGUF) ~17.1 GB 24 GB
IQ2_XXS (GGUF) ~9.0 GB 12 GB

The quantization ladder translates into concrete hardware choices:

  • The 4-bit Q4_K_M build (about 17.1 GB) fits comfortably on a 24 GB card such as an RTX 4090 or RTX 3090.
  • On a 16 GB card, drop to a 3-bit or IQ quantization, or trim the context window.
  • The compact IQ2_XXS build (about 9 GB) runs on 12 GB of VRAM.
  • Independent tester Simon Willison reported roughly 15 to 30 tokens per second on an M5 Max MacBook Pro and NVIDIA DGX Spark.

Unified-memory mini workstations take this even further. The AMD Strix Halo and the NVIDIA DGX Spark both pair generous unified memory with enough bandwidth to run the 27B comfortably at its Q4 size.

One practical note: the model defaults to an extra-high reasoning effort, which can cause it to overthink simple prompts and burn tokens. Setting the reasoning effort to low or medium is the recommended starting point for most tasks. This continues a trend Qwen has leaned into before, where compact open models challenge the assumption that bigger is always better.

What’s Next: The Qwen 3.8-35B-A3B and Other Sizes

Alibaba has already expanded the family further with Qwen3.8-Flash, a 125B MoE model with only 6B active parameters that beats DeepSeek V4 Pro on SWE-bench Pro at a fraction of the price. The community also widely expects a 35B-A3B to follow. The “A3B” label follows a pattern Qwen has used across several generations: a 35-billion-parameter mixture-of-experts that activates only about 3 billion parameters per token.

Bar chart comparing active parameters per token for Qwen 3.8 27B, 2.4T-A95B, and 35B-A3B models
(Credit: Intelligent Living)

Because only a small slice of the weights fire for each token, the 35B-A3B promises near-27B quality at a fraction of the compute:

  • 35 billion total parameters, with only about 3 billion active per token
  • Roughly 9x faster inference than a dense 35B model
  • Expected to trail the 27B by only a modest margin on benchmarks, if prior A3B releases are any guide
  • Quantized versions should run comfortably on modest home hardware with cheap per-token costs
  • No official release date has been announced yet
  • Smaller edge-focused models and additional mid-tier checkpoints are also expected to follow

Qwen’s recent history points to more sizes ahead. Qwen 3.5 shipped as a family of eight open-weight models under Apache 2.0: five dense models (0.8B, 2B, 4B, 9B, and 27B) and three mixture-of-experts models (35B-A3B, 122B-A10B, and 397B-A17B).

Bar chart showing the eight Qwen 3.5 open-weight model sizes from 0.8B to 397B parameters, dense and MoE
(Credit: Intelligent Living)

The small end of that ladder matters for everyday use. The 0.8B and 2B models run on phones and tablets, the 4B and 9B fit on laptops without a discrete GPU or on graphics cards with less than 16 GB of VRAM, and the 27B needs a 24 GB card. A smaller Qwen 3.8 dense model in the 8B range would bring this generation’s capabilities to everyday devices.

A smartphone on a desk running a compact AI model with a small neural network visualization
(Credit: Intelligent Living)

Smaller models also serve a second purpose as draft models for speculative decoding, a technique where a cheap model guesses several tokens ahead and a larger model verifies them in parallel to speed up output. Qwen 3.8-27B already has multi-token prediction built in, so a separate draft model is not required. Community testers have measured roughly a 72 percent throughput boost from enabling it.

Qwen 3.6, by contrast, shipped just two open-weight models, the 27B and the 35B-A3B, alongside closed API tiers. That makes the 35B-A3B the likeliest next open release, but smaller sizes are plausible too. None of this is confirmed, so treat the specifics as educated expectations rather than announcements.

Frequently Asked Questions

Is Qwen 3.8 free to use?

Partially. The 27B model is released under Apache 2.0, so the weights are free to download and use commercially. Not every new Qwen release shares those terms: the image-generation sibling Qwen Image 2.1 dropped Apache 2.0 for a research-only license when it launched on September 20, 2026. The 2.4T open-weights checkpoint is also downloadable under a custom license. The hosted API is paid, at $2 per million input and $6 per million output tokens.

Is Qwen 3.8 good at coding?

Yes. On published benchmarks it scores 61.7 on SWE-bench Pro, 90.3 on LiveCodeBench v6, and 79.0 on QwenSWEBench, beating Claude Opus 4.6 Max on all three. It is weaker on the hardest agentic coding tests, such as DeepSWE 1.1.

Can I run Qwen 3.8 locally?

Yes, the 27B model is designed for local use. A 4-bit GGUF quantization is about 17 GB and fits on a 24 GB GPU, while a 2-bit build at roughly 9 GB runs on 12 GB of VRAM. The 2.4T model is a datacenter-scale download and is not practical for home hardware.

Is Qwen 3.8 a dedicated coding model?

No. It is a general-purpose vision-language model that happens to be strong at coding. The 27B handles text, image, and video input, and the family is aimed at general reasoning, office work, and long-horizon agentic tasks rather than coding alone.

Conclusion

Qwen 3.8 marks a turning point for open-weight AI. A 27-billion-parameter model that runs on a single GPU is now competitive with flagship closed models from Anthropic and OpenAI on several benchmarks, and the 2.4-trillion-parameter flagship is the first Max-class Qwen model released with open weights.

The caveats are real: most benchmark numbers are self-reported, and independent validation is still catching up. But the direction is clear. As the anticipated 35B-A3B and other sizes land, the gap between what you can run at home and what you pay a premium to access through an API will keep narrowing. Z.ai’s GLM-5.3, released the same week, pushes this frontier further by achieving 6× coding gains through post-training alone. IBM’s Granite 4.2, a dense reasoning model family ranging from 3B to 30B parameters, also arrived the same week and scored 57.00 on SWE-bench Verified with its 30B model, adding another contender to the growing field of competitive open-weights models. Google’s response, Gemini 3.8 Flash, landed on September 2 with intelligence gains over the previous 3.7 Flash tier and a dedicated cybersecurity variant, continuing the same week-by-week pace of frontier releases.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

Microbial Fuel Cells: Bacteria Cleans Wastewater and Generates Electricity

Every day, the world spends a fortune in electricity...

AI ECG: The Smartphone Tool That Spots a Hidden Heart Disease

An electrocardiogram (ECG) takes only a few minutes and...

Calcium-Ion Battery Hits 1,000 Cycles Without Any Lithium

A lithium-ion battery depends on an element that is...