A new model from Singapore has pulled level with DeepSeek’s flagship reasoning system on the industry’s most widely cited intelligence benchmark, and it is free to use through an API. Agnes 3.0 Flash scores 36 on the Artificial Analysis Intelligence Index v4.3, the same mark as DeepSeek V4 Pro. Its list price of $0.05 per million input tokens and $0.15 per million output tokens also undercuts the Chinese model several times over, and right now that price stands at $0 while a free period runs.
That combination is unusual. Until recently the price-performance frontier of capable AI was set almost entirely by Chinese labs such as DeepSeek and Xiaomi. Agnes 3.0 Flash shows a Singapore-based developer matching a Chinese flagship on measured intelligence, pricing below it, and returning answers faster than either. This article breaks down the model’s scores, its pricing, how it compares with DeepSeek V4 Pro and Xiaomi’s MiMo V2.5 family, and how developers can begin using it.
What Is Agnes 3.0 Flash?
Agnes 3.0 Flash is a next-generation text model from Agnes AI, the commercial brand of Singapore-based Sapiens AI. The company positions it less as a general chat model and more as an execution engine for agents, tuned for coding-agent workflows, function and tool calling, and long multi-step tasks where the model has to stay on plan and report results honestly.
The model is available through Agnes AI’s omni-modal API, which also serves the company’s image and video models from the same account and API key. Full request parameters and integration guides are published in the official Agnes 3.0 Flash documentation.
The headline specifications are straightforward.
- Model name: agnes-3.0-flash
- Developer: Agnes AI (Sapiens AI), Singapore
- Model type: Text model with text and image-URL input
- Context window: 512K tokens
- Maximum output: 65,536 tokens per request
- Intelligence Index: 36 (an estimate pending independent evaluation)
- Output speed: 235 tokens per second
- List price: $0.05 input / $0.15 output / $0.005 cached input per 1M tokens, all currently $0 during the free period
- Endpoints: OpenAI Chat Completions, OpenAI Responses, Anthropic Messages
- Thinking mode: available for harder tasks
What the Model Is Built For
Agnes AI’s documentation frames the model around four priorities, and they explain the shape of its capabilities.
- Coding-agent execution: stronger performance on Agnes Code and repository-style tasks, from requirements through to delivery.
- Tool calling and orchestration: more reliable function selection and multi-step tool use, with fewer redundant or looping calls.
- Instruction and context adherence: keeping objectives and constraints stable across long, multi-turn runs.
- Trustworthy delivery: greater attention to factual grounding and result verification, with less exposure of internal reasoning.
Two numbers stand out. The first is the intelligence score, because it places the model on the same footing as a far larger Chinese flagship. The second is the output speed, because 235 tokens per second is roughly three times faster than DeepSeek V4 Pro and nearly six times faster than MiMo V2.5 Pro.
Agnes 3.0 Flash Matches DeepSeek V4 Pro on Intelligence
Agnes 3.0 Flash scores 36 on the Artificial Analysis Intelligence Index v4.3, matching DeepSeek V4 Pro on the nose while listing at roughly one-eighth the input price and one-sixth the output price.
Agnes AI’s previous reasoning model, Agnes 2.5 Pro Beta, is estimated at 35 on the same index, one point behind Agnes 3.0 Flash.
Within its model class, Artificial Analysis ranks Agnes 3.0 Flash first out of 61 models on intelligence, and it sits level with DeepSeek’s Pro tier.
The firm also flags the score as an estimate while independent evaluation is pending, and notes that scores and rankings can change as the methodology updates. That is a reasonable caution for any model this new. For now the estimate is the best public measure available, and it is a strong one.
How Agnes 3.0 Flash Compares: Benchmarks, Price, and Speed
Pulling the numbers together makes the positioning clearer. The table below places Agnes 3.0 Flash alongside DeepSeek V4 Pro, DeepSeek’s newer V4.1 Flash, Xiaomi’s MiMo V2.5 and MiMo V2.5 Pro, and Agnes AI’s own previous model, all scored under the same v4.3 index.
| Model | AA Intelligence Index (v4.3) | Input $ / 1M | Output $ / 1M | Cached input $ / 1M | Output speed |
|---|---|---|---|---|---|
| Agnes 3.0 Flash | 36 | $0.05 | $0.15 | $0.005 | 235 t/s |
| DeepSeek V4 Pro | 36 | $0.435 | $0.87 | $0.003625 | 72 t/s |
| DeepSeek V4.1 Flash | 40 | $0.15 | $0.60 | $0.003 | 206 t/s |
| Agnes 2.5 Pro Beta | 35 | $0.10 | $0.30 | $0.01 | 159.5 t/s |
| MiMo V2.5 Pro | 26 | $0.435 | $0.87 | $0.0036 | 41 t/s |
| MiMo V2.5 | 22 | $0.14 | $0.28 | $0.0028 | 52.5 t/s |
Figures are Artificial Analysis Intelligence Index v4.3. The two Agnes AI scores are currently estimates pending independent evaluation. Agnes 3.0 Flash prices are list prices and are currently $0 during the free period. DeepSeek applies peak and off-peak multipliers on some tiers, so its effective cost can vary by time of day, as detailed in DeepSeek’s official API pricing.
Read across the rows and the pattern is consistent. Against DeepSeek V4 Pro, Agnes 3.0 Flash matches the intelligence score at roughly one-eighth of the input price and one-sixth of the output price, while running more than three times faster.
Against Xiaomi’s MiMo V2.5, the newer Singaporean model is 14 index points ahead and faster, and it is cheaper on both input and output. Against MiMo V2.5 Pro it is 10 points ahead, less than a quarter of the input price, and about six times faster on output.
The one place Agnes 3.0 Flash does not win is cache hits. Its cached-input price of $0.005 per million tokens is cheap in isolation, but DeepSeek and Xiaomi have built their reputations on near-total cache discounts. DeepSeek V4 Pro reads cached input at $0.003625, DeepSeek V4.1 Flash at $0.003, MiMo V2.5 Pro at $0.0036, and MiMo V2.5 at $0.0028. For a workload that reuses the same long system prompt or codebase on every call, those fractions of a cent can matter more than the headline input rate. This is the same architectural dynamic we explored when explaining why Chinese AI models are so much cheaper than OpenAI and Anthropic, where compressed attention and prompt caching drive the cost gap.
The most interesting row is DeepSeek V4.1 Flash. It is the only model in the table that edges ahead of Agnes 3.0 Flash on intelligence, at 40 against 36, and its off-peak input price of $0.15 is competitive. But its output price of $0.60 is four times higher, and its cached-input rate is slightly lower. DeepSeek’s own positioning has shifted toward Flash-class models as the default for most workloads, a transition covered in our breakdown of the DeepSeek V4.1 Flash release and what happens to V4 Pro.
Full Benchmark Breakdown: Where Agnes 3.0 Flash Leads and Trails
The Intelligence Index is a composite, so it hides where a model is strong and where it struggles. Artificial Analysis publishes the underlying evaluations, and setting them side by side with DeepSeek’s two current models produces a more textured picture than the headline score alone.
| Benchmark | Agnes 3.0 Flash | DeepSeek V4 Pro | DeepSeek V4.1 Flash | Measures |
|---|---|---|---|---|
| AA Intelligence Index v4.3 | 36 | 36 | 40 | Composite |
| AutomationBench-AA | 51% | 57% | 68.9% | Agentic SaaS workflows |
| GDPval-AA v2 | 1573 Elo | 1,493 Elo | 1,632 Elo | Real-world work tasks |
| Terminal-Bench v4.0 | 7% | 14% | 26.8% | Agentic terminal use |
| SciCode | 52% | 51% | 51.9% | Scientific coding |
| Humanity’s Last Exam | 38% | 41% | 39.2% | Reasoning and knowledge |
| GDP.pdf | 10% | 11% | 12.8% | Document reasoning |
| CritPt | 15% | 18% | 14.3% | Physics reasoning |
| AA-LCR v1.1 | 81% | 80% | 84.0% | Long-context reasoning |
| AA-Omniscience index | -11 | 1 | -5 | Knowledge reliability |
| AA-Omniscience accuracy | 25% | 49% | 46% | Knowledge |
| τ3-Banking | 48% | 39.6% | Not published | Banking tool use |
Scores are from Artificial Analysis under index version v4.3, except where a model is not covered on that evaluation. GDPval-AA v2 is published as an Elo score. A negative AA-Omniscience index means the model answered more questions incorrectly than correctly.
Two patterns emerge. The first is that Agnes 3.0 Flash is genuinely competitive on coding and long-context work: it edges DeepSeek V4 Pro on SciCode, 52% against 51%, on long-context reasoning, 81% against 80%, and on GDPval-AA v2, 1,573 Elo against 1,493, and it lands within a couple of points on Humanity’s Last Exam and GDP.pdf.
The second is that DeepSeek keeps a clear lead on the agentic evaluations that the composite index weights heavily. On AutomationBench-AA, which tests SaaS workflows, V4 Pro scores 57% and V4.1 Flash 68.9% against 51% for Agnes 3.0 Flash. On Terminal-Bench v4.0 the gap is wider still, with Agnes at 7% against 14% and 26.8%. Physics reasoning follows the same shape, at 15% for Agnes against 18% for V4 Pro.
Reliability tells a more pointed story. On AA-Omniscience, which tests factual knowledge, Agnes 3.0 Flash answers 25% of questions accurately, against 49% for DeepSeek V4 Pro and 46% for V4.1 Flash. Its reliability index of -11, meaning slightly more incorrect answers than correct, sits below both.
The practical reading is that the shared score of 36 conceals two different models. A team building terminal automation or multi-step SaaS workflows will still get better results from DeepSeek. A team working on scientific coding, document analysis, or long-context retrieval can match or beat it with Agnes 3.0 Flash, at a fraction of the price and several times the speed. On banking tool use, the model scores 48%, ahead of DeepSeek V4 Pro’s 39.6% and well above the 36% its own predecessor managed.
Free Now, and Cheaper Than DeepSeek Later
The pricing story has two halves. The first is temporary: every token on Agnes 3.0 Flash currently costs nothing. The second is permanent: even after the free period ends, the list price stays below DeepSeek’s.
| Billing item | List price | Current price (free period) |
|---|---|---|
| Input tokens | $0.05 / 1M | $0 |
| Output tokens | $0.15 / 1M | $0 |
| Cached input | $0.005 / 1M | $0 |
Agnes AI has run an aggressive free tier across its model line, and the Flash-class text models have been free with no announced end date. The practical limits are requests-per-minute caps rather than pricing, and the company recommends exponential backoff for retries under burst load. As with several no-cost tiers, free usage may be used to improve the models, which is worth testing before routing production traffic through it.
A quick worked example shows the scale of the gap. A developer running 50 million output tokens through DeepSeek V4 Pro at $0.87 per million would pay about $43.50. The same volume through Agnes 3.0 Flash at $0.15 per million costs $7.50, and nothing at all while the free period lasts. The input side widens further, since Agnes charges $0.05 per million against DeepSeek’s $0.435.
The caveat remains cache-heavy workloads. An agent that resends a large fixed context on every call will get better economics from DeepSeek or Xiaomi, because their cache-hit prices are lower. Agnes wins on fresh tokens and on output tokens, which is where most generation-heavy pipelines spend their money.
Agnes 3.0 Flash vs. DeepSeek V4 Pro: A Head-to-Head Breakdown
Because the two models land on the same intelligence score, the choice between them comes down to economics and throughput rather than raw capability. Splitting the comparison into its parts makes the trade-offs legible.
Price per Million Tokens
Agnes 3.0 Flash lists at $0.05 per million input tokens and $0.15 per million output tokens. DeepSeek V4 Pro lists at $0.435 and $0.87 for the same volumes, roughly 8.7 times higher on input and 5.8 times higher on output. For a pipeline whose cost is dominated by generated tokens, that ratio is the most consequential difference between the two.
Output Speed and Latency
Agnes 3.0 Flash generates 235 tokens per second against DeepSeek V4 Pro’s 72, a 3.3-fold advantage. In an agent loop that issues dozens of sequential calls, that margin accumulates into minutes saved per task rather than milliseconds, and it changes how interactive a tool feels in practice. The start-up is closer to even on raw input processing: Agnes 3.0 Flash’s time-to-first-token is 1.84 seconds, in the same sub-two-second band as DeepSeek V4 Pro. Once the reasoning pass is included, though, the gap widens: Agnes reaches its first answer token in 10.5 seconds, against 27.4 seconds for DeepSeek V4 Pro.
Cache-Hit Economics
The ranking flips on cached input. DeepSeek V4 Pro reads cached tokens at $0.003625 per million, while Agnes 3.0 Flash charges $0.005, and DeepSeek’s newer V4.1 Flash goes lower still at $0.003. DeepSeek’s compressed attention architecture shrinks the key-value cache so aggressively that repeated context becomes nearly free, an edge that matters most for retrieval pipelines and long-running agents that resend the same prompt.
Which Model to Choose
- Pick Agnes 3.0 Flash for generation-heavy work, interactive agents, and cost-sensitive prototypes where output volume dominates the bill.
- Pick DeepSeek V4 Pro when cached context dominates, when you need the strongest agentic terminal performance, or when a mature, widely documented ecosystem matters.
- Pick DeepSeek V4.1 Flash for the highest intelligence score in this group, accepting roughly four times the output price.
Why 238 Tokens per Second Matters
Speed is easy to overlook next to an intelligence score, but for agentic work it is often the deciding factor. An agent that calls a model dozens of times per task multiplies every second of latency across the whole run, so a faster model finishes the job sooner and consumes less wall-clock time.
Agnes 3.0 Flash generates output at 238 tokens per second. DeepSeek V4.1 Flash runs at 219, DeepSeek V4 Pro at 69, MiMo V2.5 at 48, and MiMo V2.5 Pro at 46. That makes Agnes 3.0 Flash the fastest model in this group, roughly three times faster than DeepSeek V4 Pro and nearly six times faster than MiMo V2.5 Pro, a gap that compounds across multi-step tool use and long-context reasoning. It is not, however, the fastest model on the market: Google’s Gemini 3.8 Flash generates output at 267.2 tokens per second, ahead of Agnes 3.0 Flash, though in a very different price bracket.

Latency is where the picture changes, and there are two distinct waits to separate. The first is time-to-first-token, which measures only input processing: Agnes 3.0 Flash’s is 1.84 seconds, and DeepSeek V4 Pro streams its first token in under two seconds, so the two are close on that figure. Xiaomi’s models are far slower still, with MiMo V2.5 at 6.9 seconds and MiMo V2.5 Pro at 8.2 seconds. The second wait is time-to-first-answer-token, which adds the reasoning pass the model completes before it produces its first visible answer. On that figure the gap opens: Agnes 3.0 Flash reaches its answer in 10.5 seconds, against 27.4 seconds for DeepSeek V4 Pro. For interactive coding assistants, that answer-token delay is the one a user actually waits on, and it is exactly why Agnes AI positions the model for tool-driven, multi-turn work where staying responsive across a long chain of steps matters more than raw peak throughput.
Agnes 3.0 Flash vs. MiMo V2.5: The Cheaper Challenger
Xiaomi’s MiMo family competes on the same low-price premise, so it is a useful benchmark for the Singaporean model. The comparison is less flattering for Xiaomi than the headline price tags suggest.
- Intelligence: Agnes 3.0 Flash scores 36 on the v4.3 index, against 26 for MiMo V2.5 Pro and 22 for MiMo V2.5.
- Price: Agnes is cheaper than MiMo V2.5 Pro on input and output and cheaper than MiMo V2.5 on both, while scoring far higher on the index.
- Speed: Agnes returns 235 tokens per second against 41 for MiMo V2.5 Pro and 52.5 for MiMo V2.5.
- Latency: MiMo V2.5 reports a time-to-first-token of 6.9 seconds and MiMo V2.5 Pro 8.2 seconds, both well above the roughly two-second class median.
- Caching: Xiaomi keeps a narrow win here, at $0.0028 per million cached input tokens for MiMo V2.5, against Agnes’s $0.005.
In short, MiMo remains among the cheapest capable models from China and a reasonable choice for cache-heavy, latency-tolerant batch processing. For interactive agent workloads, the combination of a lower intelligence score, slower generation, and higher latency weakens the case considerably.
Where Agnes 3.0 Flash Still Trails
No model wins on every axis, and Agnes 3.0 Flash has clear limits worth weighing before a switch.
- It is not the intelligence leader. DeepSeek V4.1 Flash scores 40 on the v4.3 index, four points ahead of Agnes 3.0 Flash.
- It trails on agentic benchmarks. AutomationBench-AA gives it 51% against 57% for DeepSeek V4 Pro and 68.9% for V4.1 Flash, and Terminal-Bench v4.0 gives it 7% against 14% and 26.8%.
- It is weaker on factual knowledge. Its AA-Omniscience accuracy of 25% trails DeepSeek V4 Pro at 49% and V4.1 Flash at 46%, and its reliability index of -11 sits below both.
- The headline score is an estimate. Artificial Analysis has not finished independent evaluation, so the figure of 36 could move as the methodology settles.
- Cache pricing is not the cheapest. DeepSeek and Xiaomi both undercut it on cached input tokens.
- There are no open weights. Unlike MiMo V2.5 and DeepSeek V4.1 Flash, Agnes 3.0 Flash cannot be self-hosted, so sensitive workloads must run on Agnes AI’s own infrastructure.
- The context window is smaller. At 512K tokens it trails the 1M-token windows offered by DeepSeek and MiMo, which matters for whole-codebase or document-set analysis.
Who Is Agnes AI (Sapiens AI)?
Agnes AI is the public product line of Sapiens AI, a Singapore-headquartered startup founded by Bruce Yang. The company raised a $10 million Series A in February 2026, taking total funding to about $20 million, and was approaching $20 million in annual recurring revenue by March 2026, unusually fast monetization for a model maker at this tier.

Its positioning has been explicit. Local press described the company’s earlier model as Singapore’s homegrown answer to DeepSeek, and Sapiens AI targets consumer and developer markets across Southeast Asia, Latin America, and the Middle East with a single omni-modal API that spans text, image, and video. The reasoning model that Agnes 3.0 Flash overtakes on the index, Agnes 2.5 Pro Beta, remains available through the same API.
Singapore itself has treated AI as a national priority, publishing a Model AI Governance Framework for Generative AI and funding a government-backed research program, AI Singapore. Agnes AI is among the first Singapore-headquartered labs to ship a model that benchmarks competitively with frontier Chinese and US releases, and Agnes 3.0 Flash pushes that claim further than any of its predecessors.
How to Access the Agnes 3.0 Flash API
Base Configuration
The API is OpenAI-compatible, so most existing code can be ported by swapping the base URL, the API key, and the model name. The same hub also serves the OpenAI Responses API and the Anthropic Messages API.
- Base URL: https://apihub.agnes-ai.com/v1
- Chat Completions: POST /v1/chat/completions
- Responses API: POST /v1/responses
- Messages API: POST /v1/messages (Anthropic-compatible)
- Model name: agnes-3.0-flash
- Authentication: Bearer token, or an x-api-key header for the Messages API
A Minimal Request
A minimal request looks like this.
curl https://apihub.agnes-ai.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "agnes-3.0-flash",
"messages": [
{"role": "user", "content": "Explain how an agent should choose and call a tool."}
],
"max_tokens": 1024
}'

Developers who want more deliberate reasoning can enable Thinking mode through a chat_template_kwargs field on the OpenAI-compatible endpoint, or a thinking block on the Anthropic-compatible one. Agnes AI’s own guidance is to state the task objective, runtime context, constraints, and tool permissions up front, and to return tool results to the conversation before asking the model for its next action. Accounts are created without a credit card, and the same login unlocks the company’s image models, such as Agnes Image 2.1 Flash, and its video generator.
Who Should Use Agnes 3.0 Flash?
The model is not a universal upgrade, but it fits a specific profile well.
- Coding and agent developers who need reliable tool calling and fast iteration on multi-step tasks.
- Startups and solo builders who want frontier-adjacent capability without Gemini or Claude pricing.
- Teams diversifying providers who want a non-US, non-Chinese option for compliance or resilience.
- Prototyping and evaluation work where the current $0 price removes the cost of experimentation.
It is a weaker fit for workloads that depend on heavy prompt caching, for applications that require self-hosting, and for tasks that need the very largest context windows or the absolute highest intelligence score.
Frequently Asked Questions
Is Agnes 3.0 Flash free to use?
Yes, at the moment. Agnes 3.0 Flash is free through the Agnes AI API during a promotional period, and all three billing items, input, output, and cached input, are currently priced at $0. The list price that takes effect afterward is $0.05 per million input tokens, $0.15 per million output tokens, and $0.005 per million cached-input tokens.
What is the performance benchmark for Agnes 3.0 Flash?
Agnes 3.0 Flash scores 36 on the Artificial Analysis Intelligence Index v4.3, which ranks it first out of 61 models in its class. It also generates output at 235 tokens per second, among the fastest in its class. The intelligence score is currently marked as an estimate while independent evaluation is pending.
Who is Agnes AI?
Agnes AI is the commercial brand of Sapiens AI, a Singapore-based startup founded by Bruce Yang. It builds full-modality foundation models and offers them through a single omni-modal API, spanning text, image, and video, with a free tier that has attracted millions of users.
How does Agnes 3.0 Flash compare to DeepSeek V4 Pro?
They share an intelligence score of 36 on the Artificial Analysis Intelligence Index v4.3, but Agnes 3.0 Flash is far cheaper on fresh tokens and output, at $0.05 and $0.15 per million versus $0.435 and $0.87, and it is roughly three times faster. DeepSeek V4 Pro keeps an edge on cache-hit pricing, where it charges $0.003625 per million against Agnes’s $0.005.
Can I self-host Agnes 3.0 Flash?
Not directly. Agnes 3.0 Flash is a proprietary model available only through the hosted Agnes AI API. There is no open-weights release for this model.
What is the difference between Agnes 3.0 Flash and Agnes 2.5 Pro Beta?
Agnes 2.5 Pro Beta is Agnes AI’s previous reasoning model, estimated at 35 on the v4.3 index. Agnes 3.0 Flash is the newer, faster, cheaper model, scoring 36 and running at 235 tokens per second against 159.5. Both remain available through the same API.
What API endpoints does Agnes 3.0 Flash support?
Three: the OpenAI-compatible Chat Completions API, the OpenAI Responses API, and the Anthropic-compatible Messages API. All are served from the same base URL, so switching between them does not require re-authentication.
How fast is Agnes 3.0 Flash?
Agnes 3.0 Flash generates output at 235 tokens per second, the fastest of the models compared here. That is roughly three times faster than DeepSeek V4 Pro, marginally ahead of DeepSeek V4.1 Flash at 206, and nearly six times faster than MiMo V2.5 Pro. It is not the fastest model overall, though: Google’s Gemini 3.8 Flash runs at 267.2 tokens per second.
Does Agnes 3.0 Flash support image input?
Yes. Agnes 3.0 Flash accepts text and public image URLs in the same request, which suits tasks such as chart analysis and visual reasoning. Output is text only.
Is Agnes 3.0 Flash open source?
No. Agnes 3.0 Flash is a proprietary model served only through the Agnes AI API. Developers who need open weights for self-hosting would need a different model, such as MiMo V2.5 or DeepSeek V4.1 Flash.
The Bottom Line
Agnes 3.0 Flash is the clearest sign yet that the price-performance frontier of capable AI is no longer set in one country. A Singaporean lab has matched DeepSeek V4 Pro on the Artificial Analysis Intelligence Index, priced its model well below the Chinese incumbent, and shipped it several times faster, with the whole thing free to try through an API.
The trade-offs are real. DeepSeek V4.1 Flash still scores higher on intelligence, DeepSeek and Xiaomi still win on cache-hit economics, and Agnes AI’s headline score is an estimate pending independent confirmation. But for teams that want strong agentic performance without Gemini or Claude pricing, and that would rather not depend solely on Chinese providers, Agnes 3.0 Flash is now a serious default. The next test is whether the free period lasts, and whether the estimated index score holds once the independent evaluation lands.
