Unisound U2-Flash Takes on Xiaomi MiMo in the Ultra-Cheap LLM Tier

Date:

Unisound announced U2-Flash in a voluntary filing to the Hong Kong Stock Exchange on September 15, 2026: a sparse mixture-of-experts model carrying roughly 266 billion total parameters while activating only about 10 billion of them per inference. Chinese tech outlets covered the launch in depth. English-language coverage has been close to nonexistent.

That gap matters for anyone shopping for the cheapest capable LLM APIs, because U2-Flash lands squarely in the price tier where Xiaomi’s MiMo V2.5 and DeepSeek’s Flash line already compete for volume. Company figures put its coding and agent scores above several models that cost more per token, and the launch promotion prices it below Xiaomi’s budget flagship outright.

Illustration of a sparse mixture-of-experts network where only a small cluster of nodes activates
(Credit: Intelligent Living)

What Is Unisound U2-Flash?

Unisound is a Hong Kong-listed voice and language AI company trading as 09678.HK. It built U2-Flash by applying reinforced post-training to the U2 foundation model it released in June 2026, a 260 billion parameter model the company pitched at the time as matching trillion-parameter rivals at a tenth of the token cost. U2-Flash is not a lite variant of that model. Unisound positions it as a full-capability workhorse tuned for production jobs: coding, agent orchestration, math reasoning, and instruction following, all packed into one checkpoint.

Sparse MoE: 266B Total, 10B Active

  • The model holds roughly 266 billion total parameters and activates about 10 billion per request, less than 4% of the weight set.
  • Because inference cost scales with active parameters rather than total parameters, Unisound can price the model against small dense models while claiming mainline task quality.
  • Reasoning runs in latent hidden-state space rather than as a separate visible chain, exposed through a four-tier thinking-intensity control: none, low, high, and max.
  • At the max tier the model emits a full readable reasoning chain, so the thinking process stays auditable instead of hidden.

The Model Helped Train Itself

The most unusual part of the launch is the training loop. Unisound says U2-Flash participates in its own post-training: the model helps generate tasks, analyze trajectories, and inspect the training stack itself inside human-defined sandboxes with verification gates. The company calls it its first practical step toward recursive self-improvement, or RSI, the idea that a model can keep improving its own training data and process.

Supporting pieces include asynchronous agent reinforcement learning with parallel workers, multi-teacher online policy distillation across math, code, and agent specialists, and adaptive task generation that keeps difficulty near the model’s current skill edge. The loop has already produced a software engineering task set of nearly 100,000 items spanning mainstream programming languages. Unisound reports roughly 60% more effective training trajectories and about 55% fewer training steps needed to reach matched skill.

Benchmark Scores

All scores below are company-reported: headline figures come from the September 15 announcement, with peer comparison scores drawn from Unisound’s own MaaS platform benchmark chart. In the peer columns, n/p indicates the model was not published on that row, and GLM entries use 5.3 on the coding rows and 5.2 on the reasoning rows, matching Unisound’s chart. Qwen3.8-Max and MiniMax M3 also appear on individual benchmarks and are noted in the analysis below.

Benchmark U2-Flash Prior U2 DeepSeek V4 Flash DeepSeek V4-Pro DeepSeek V4.1 Flash GLM Kimi K3
DeepSWE v1.1 64.6 32 54.4 62.7 74.2 66.9 67.5
TerminalBench 3.0 24.3 2.7 7.6 11.8 30 28.3 17.4
SWE-Bench Pro 61.6 50.1 56 60.3 n/p 64.6 63.3
GPQA-Diamond 92.4 87.9 88.1 90.1 90.9 91.2 n/p
IMOAnswerBench 90.3 79.3 88.4 89.8 n/p 91 n/p
HMMT Feb. 2026 97.1 84.8 94.8 95.2 n/p 92.5 n/p
WorkBuddy-Bench Office 81 72 77.5 78.7 n/p n/p n/p

Three patterns stand out. First, the generational jumps are large rather than incremental: DeepSWE performance doubled, TerminalBench went from 2.7 to 24.3, GPQA-Diamond rose from 87.9 to 92.4, and HMMT Feb. 2026 climbed from 84.8 to 97.1. Second, on reasoning and office tasks the model genuinely trades above its price class, tying Qwen3.7-Max on HMMT and beating every charted rival except MiniMax M3 on GPQA-Diamond. Third, the picture is more mixed on the headline coding-agent suites: Unisound’s own chart shows DeepSeek’s newer V4.1 Flash ahead on both DeepSWE v1.1 (74.2 vs 64.6) and TerminalBench 3.0 (30 vs 24.3), which reframes U2-Flash’s pitch from benchmark leadership to strong scores at a lower price. English write-ups of the launch have largely missed the full chart, including the WorkBuddy-Bench Office result that puts a 10B-active-parameter model above DeepSeek’s V4 Flash and V4 Pro checkpoints on a workplace-style task suite.

The catch is provenance. Every U2-Flash figure above comes from Unisound materials, and two entries in its chart are vendor-curated rather than community benchmarks: U2-SWE-Bench, an in-house suite where U2-Flash scores 57.5 against V4.1 Flash’s 58.3, and the WorkBuddy series. The closest thing to independent data is LLM Stats’ ZeroEval testing of the parent U2 model in June 2026, which broadly corroborated Unisound’s claims, with 86.9% on GPQA Diamond against a claimed 87.9 and 73.4% on SWE-bench Verified against a claimed 75, while flagging weak grounded factuality at 44.3% on FACTS Grounding. No equivalent third-party run exists for U2-Flash yet.

Pricing and the Promo Window

U2-Flash is live on the Unisound MaaS platform with a limited-time 40% discount running through September 30, 2026, plus a trial allotment of 100 million tokens for new and returning users.

Rate Promo Price (per 1M tokens) Approx. USD
Input ¥0.6 $0.08
Output ¥1.2 $0.17
Cache hit ¥0.12 $0.02

At those rates U2-Flash undercuts nearly everything in the budget class. The undiscounted rate implied by the 40% promotion, roughly ¥1 input and ¥2 output, lands at about $0.14 and $0.28 per million tokens, which is the same territory as DeepSeek’s standard Flash pricing. That is the real story of this release: the discount is aggressive, but the list price already competes. For a broader rate context, our LLM API pricing comparison tracks the rest of the market.

How Much Does DeepSeek V4 Flash Cost?

DeepSeek V4 Flash remains one of the most-searched budget rates for good reason: $0.14 per million input tokens and $0.28 per million output at standard rates, with cache-hit input at $0.0028 per million. Its successor, V4.1 Flash, launched on September 10, 2026, with a 552 billion parameter backbone, a 1 million token context window, native vision input, and off-peak pricing of $0.15 input, $0.60 output, and $0.003 per million on a cache hit. Peak hours, defined as weekday mornings and early afternoons UTC, double those rates. DeepSeek’s V4.1 Flash debut also kept the old V4 Flash model identifier working through October 10, 2026, so existing integrations have a transition window.

U2-Flash vs. Xiaomi MiMo V2.5

This is where the ultra-cheap tier gets interesting. Xiaomi’s MiMo V2.5 standard model lists at $0.14 input and $0.28 output per million tokens with a 1 million token context window, while the Pro variant runs $0.42 and $0.83. U2-Flash’s promo rate beats both on price, and its standard rate matches the standard MiMo tier exactly.

Model Input / 1M Output / 1M Active or Total Params
Unisound U2-Flash (promo) $0.08 $0.17 266B total, 10B active
Xiaomi MiMo V2.5 $0.14 $0.28 ~309B total MoE
DeepSeek V4 Flash $0.14 $0.28 284B total, 13B active
DeepSeek V4.1 Flash (off-peak) $0.15 $0.60 552B total, 8B to 16B active
Xiaomi MiMo V2.5 Pro $0.42 $0.83 1T+ total MoE

The comparison cuts both ways. MiMo V2.5 brings a 1 million token context window, multimodal input, and Xiaomi’s distribution muscle, and on price-per-token the standard tier has been the reference point for budget buyers all year. U2-Flash counters with a coding and agent focus, workplace and software engineering scores that beat DeepSeek’s V4 Flash and V4 Pro checkpoints in Unisound’s own chart, and the lowest active-parameter footprint in its class, though the newer V4.1 Flash still out-scores it on the coding-agent suites. This kind of gap between Chinese and Western pricing is a recurring theme, which our piece on why Chinese AI models are so much cheaper explains in depth.

What is still missing is an independent, apples-to-apples comparison of U2-Flash against MiMo V2.5 and MiMo V2.5 Pro on identical prompts; Unisound has not benchmarked against any MiMo checkpoint. Until someone runs one, the honest summary is that U2-Flash matches MiMo on price, beats its own predecessor by wide margins, posts office-task scores above DeepSeek’s charted V4 checkpoints, and trails the newer V4.1 Flash on coding-agent benchmarks.

Speed and Deployment

  • Time to first token averages under 3 seconds, with peak output throughput up to 300 tokens per second.
  • Compared with the prior U2, agent workloads use 20% to 30% fewer iteration steps and about 35% shorter end-to-end task cycles, with similar token savings.
  • Unisound reports generation speed roughly 2.1 times faster than U2.
  • The model has been adapted for major domestic accelerator stacks across government, finance, manufacturing, and energy deployments.

The commercial stakes are visible in Unisound’s filings: first-half 2026 token business revenue reached nearly ¥30 million, up about 760% year over year. Pricing low on a 10B-active model is a volume play for a company whose entire growth story now depends on token consumption. U2-Flash is available through the company’s MaaS platform, where it sits alongside the parent U2 model, and its endpoint speaks both the OpenAI and Anthropic API protocols. Unisound publishes integration guides for Claude Code, Cursor, Cline, and Trae, so existing agent tooling can switch over by changing the base URL, API key, and model ID.

What Remains Unverified

No third-party reproduction of the U2-Flash benchmark suite or the hardware parity claims shipped with the launch note, and Xiaomi MiMo is absent from Unisound’s peer charts entirely, so the comparison that anchors this article rests on price positioning rather than published head-to-head scores. The four-tier thinking control, the RSI-flavored training loop, and the accelerator optimizations all trace back to company materials. English documentation is thin, context window details for U2-Flash have not been clearly published in English, and the promotional token rates expire on September 30, 2026, after which pricing reverts toward the list rate. Teams should run workload-level validation before making it a default production model.

Illustration of budget AI token prices falling as cheap language models compete
(Credit: Intelligent Living)

The Bottom Line

Unisound U2-Flash is the first serious challenge to the Xiaomi MiMo and DeepSeek duopoly at the very bottom of the LLM price market. It pairs a 10B-active-parameter architecture with coding, reasoning, and office scores that clear several models costing more per token, then adds a launch discount that puts it below both incumbents, even if DeepSeek’s newer V4.1 Flash still leads it on the coding-agent suites. The claims are vendor-published, and the promo window is short, but for anyone comparing ultra-cheap APIs this month, U2-Flash now belongs on the shortlist alongside MiMo V2.5 and DeepSeek’s Flash models, and the 100 million free tokens make testing it nearly free.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

DeepSeek Voice Chat Gray Test Adds Four Selectable TTS Voices

DeepSeek appears to be testing spoken replies inside its...

Neuromorphic AI Inference: How China Mobile Cloud Cut Power Use by 40%

China's state telecom giant has paired brain-inspired silicon with...

Kimi K2.8 Preview: 1M Context Behind One Unchanged Model ID

Moonshot AI has quietly changed the model that powers...

DAMO RADAR: Alibaba’s Medical AI Detects Cancer and 146 Conditions

Alibaba's research arm has released something rare in medical...