The affordable AI model tier has become the most competitive space in the industry. Google’s Gemini 3.8 Flash, launched on September 2, 2026, and Zhipu AI’s GLM 5.3 Flash, released a week earlier on August 26, both target the same audience: developers and businesses who need strong reasoning and coding performance without frontier-model pricing.
This comparison breaks down every dimension that matters: raw benchmark performance, API pricing tiers, consumer subscription plans, output speed, context limits, and real-world value. According to Artificial Analysis v4.3, both models scored within a single point on the Intelligence Index (GLM 42 vs Gemini 41), making this one of the closest matchups in the affordable AI tier.
At a Glance: Key Specifications
The table below summarizes the core specifications of both models as of September 2026.
| Specification | Gemini 3.8 Flash | GLM 5.3 Flash |
|---|---|---|
| Developer | Google DeepMind | Zhipu AI (Z.AI) |
| Release Date | September 2, 2026 | August 26, 2026 |
| Architecture | Sparse MoE | MoE |
| Context Window | 1,048,576 tokens (1M) | 1,000,000 tokens (1M) |
| Max Output | 65,536 tokens (64K) | 131,100 tokens (~128K) |
| Input Modalities | Text, image, audio, video, PDF | Text, image, video (native) |
| Thinking Levels | Low, medium, high | Low, high, max |
| Weights | Closed | Open (MIT license) |
| Free Tier | Google AI Studio free tier | GLM-4.7-Flash is fully free |
Performance Benchmarks
Benchmarks reveal where each model excels. Gemini 3.8 Flash leads on long-horizon software engineering and multi-domain reasoning, while GLM 5.3 Flash performs strongly on tool use, agentic tasks, and graduate-level science questions.
Head-to-Head Benchmark Comparison (Artificial Analysis v4.3)
The chart below compares scores on benchmarks where both models have published results.
| Benchmark | Gemini 3.8 Flash | GLM 5.3 Flash | Edge |
|---|---|---|---|
| AA Intelligence Index v4.3 | 41 | 42 | GLM +1 |
| GPQA Diamond (graduate science) | 95.3% | 91.2% | Gemini +4.1 |
| Humanity’s Last Exam | 47.8% | 39.9% | Gemini +7.9 |
| SciCode (scientific programming) | 56.6% | 51.6% | Gemini +5.0 |
| AA-LCR v1.1 (long-context reasoning) | 81.3% | 80.0% | Gemini +1.3 |
| AutomationBench-AA | 59.9% | 60.4% | GLM +0.5 |
| Terminal-Bench v4.0 | 19.7% | 32.8% | GLM +13.1 |
| GDP.pdf (document reasoning) | 21.0% | 15.4% | Gemini +5.6 |
| CritPt | 18.3% | 15.4% | Gemini +2.9 |
| AA-Briefcase Rubric (enterprise tasks) | 42.1% | 46.4% | GLM +4.3 |
| EnterpriseOps-Gym-AA | — | 33.2% | GLM (no Gemini data) |
| AA-Omniscience Accuracy | 55% | 28% | Gemini +27 |
Elo-Based Benchmarks
| Benchmark | Gemini 3.8 Flash | GLM 5.3 Flash | Edge |
|---|---|---|---|
| GDPval-AA v2 (knowledge work) | 1464 | 1669 | GLM +205 |
| AA-Briefcase Elo (enterprise) | 1202 | 1455 | GLM +253 |
| AA-Omniscience Index | 30 | 7 | Gemini +23 |
Vals.ai Benchmarks
Vals.ai provides independent benchmarking across a wide range of professional and technical domains. These scores offer insight into real-world task performance beyond standard academic benchmarks.
| Benchmark | Gemini 3.8 Flash | GLM 5.3 Flash | Edge |
|---|---|---|---|
| Vals Index (overall) | 62.25% | 47.22% | Gemini +15.0 |
| GPQA Diamond (graduate science) | 94.44% | 86.36% | Gemini +8.1 |
| SWE-bench (software engineering) | 80.00% | 92.00% | GLM +12.0 |
| LiveCodeBench | 89.48% | 80.51% | Gemini +9.0 |
| Vibe Code Bench v1.1 | 78.65% | 30.76% | Gemini +47.9 |
| Terminal-Bench 2.1 | 81.27% | 62.92% | Gemini +18.4 |
| Code Migration | 36.55% | 20.52% | Gemini +16.0 |
| MMLU Pro | 90.22% | 86.06% | Gemini +4.2 |
| MMMU Pro | 89.08% | 86.01% | Gemini +3.1 |
| LegalBench | 89.48% | 83.93% | Gemini +5.6 |
| Legal Research Bench | 38.94% | 45.19% | GLM +6.3 |
| Harvey’s Legal Agent | 10.00% | 6.67% | Gemini +3.3 |
| Finance Agent (v2) | 61.44% | 57.85% | Gemini +3.6 |
| TaxEval v2 | 74.45% | 75.59% | GLM +1.1 |
| Tax Agent Bench | 66.77% | — | Gemini (no GLM data) |
| MortgageTax | 65.34% | — | Gemini (no GLM data) |
| Public Benefits Benchmark | 65.29% | — | Gemini (no GLM data) |
| MedScribe | 84.50% | 88.94% | GLM +4.4 |
| MedCode | 48.13% | — | Gemini (no GLM data) |
| BioMysteryBench | 62.22% | — | Gemini (no GLM data) |
| CyberBench | 43.75% | 71.49% | GLM +27.7 |
| SAGE | 35.06% | — | Gemini (no GLM data) |
| EMB | 72.20% | 55.93% | Gemini +16.3 |
| ProofBench v1.1 | 48.00% | 21.00% | Gemini +27.0 |
| ProgramBench | 1.00% | — | Gemini (no GLM data) |
| SkillsBench | 57.98% | 40.16% | Gemini +17.8 |
Provider-Reported Benchmarks
In addition to Artificial Analysis’s independent testing, both Google and Z.AI publish their own benchmark results. These come from the providers directly and use different test conditions, but they offer additional data points.
| Benchmark | Gemini 3.8 Flash (Google) | GLM 5.3 Flash (Z.AI) |
|---|---|---|
| DeepSWE v1.1 (long-horizon coding) | 73.7% | 63.4% |
| Terminal-Bench 2.1 (terminal tasks) | 89.4% | 84.3% |
| HLE-Verified (multidisciplinary reasoning) | 54.9% | — |
| HLE w/ Tools | — | 55.3% |
| Agents’ Last Exam | — | 26.3% |
| AutomationBench | — | 48.8% |
| CharXiv Reasoning | 86.2% | — |
| LVBench (agentic / static) | 87.8% / 87.1% | — |
| OSWorld-2.0 | 59.0% | — |
| BioMysteryBench (solvable / difficult) | 88.8% / 56.5% | — |
| LABBench2 | 86.2% | — |
| Vals Finance Agent v2 | 61.4% | — |
Note: Provider-reported benchmarks use different test conditions and may not be directly comparable across providers. Artificial Analysis v4.3 scores above provide independent, apples-to-apples comparison.
Key Benchmark Takeaways
- Overall Intelligence: GLM 5.3 Flash edges ahead on the AA Intelligence Index v4.3 (42 vs 41), while Gemini leads on the Vals Index (62.25% vs 47.22%). These conflicting results show that overall rankings depend heavily on which tasks are tested.
- Software Engineering: This is where the models diverge most sharply depending on the benchmark. GLM dominates SWE-bench (92.00% vs 80.00%), but Gemini leads on Vibe Code Bench (78.65% vs 30.76%), LiveCodeBench (89.48% vs 80.51%), and provider-reported DeepSWE v1.1 (73.7% vs 63.4%). Terminal-Bench v4.0 also favors GLM (32.8% vs 19.7%).
- Science and Reasoning: Gemini leads consistently on GPQA Diamond (95.3% vs 91.2% 94.44% vs 86.36%), Humanity’s Last Exam (47.8% vs 39.9%), and SciCode (56.6% vs 51.6%).
- Professional Domains: Gemini leads on Vals.ai professional benchmarks: Finance Agent v2 (61.44% vs 57.85%), LegalBench (89.48% vs 83.93%), and EMB (72.20% vs 55.93%). However, GLM leads on Legal Research Bench (45.19% vs 38.94%), MedScribe (88.94% vs 84.50%), and TaxEval v2 (75.59% vs 74.45%).
- Cybersecurity: GLM shows a strong lead on CyberBench (71.49% vs 43.75%), a 28-point advantage that aligns with Z.AI’s focus on security applications.
- Enterprise and Knowledge Work: GLM dominates on Elo-based benchmarks: GDPval-AA v2 (1669 vs 1464) and AA-Briefcase (1455 vs 1202 Elo), plus the AA-Briefcase Rubric (46.4% vs 42.1%).
- Domain-Specific: Gemini shows strong results in specialized domains: 86.2% on CharXiv Reasoning, 88.8% on BioMysteryBench, and 61.4% on Vals Finance Agent v2, though Harvey’s Legal Benchmark remains challenging at 10%.
- Omniscience: Gemini pulls far ahead on AA-Omniscience Accuracy (55% vs 28%) and Index (30 vs 7), indicating broader world knowledge.
- Automation: The models are nearly tied on AutomationBench-AA (60.4% vs 59.9%), separated by just half a point.
- Cost per Task: GLM 5.3 Flash costs $0.25 per benchmark task versus Gemini’s $1.24, a 5x cost advantage.
API Pricing Comparison
Pricing is where these two models diverge sharply. GLM 5.3 Flash undercuts Gemini 3.8 Flash by roughly 5x on input and 7.5x on output at list price, and the gap widens further during Z.AI’s launch promotion.
Standard API Rates (per 1 million tokens)
| Pricing Tier | Gemini 3.8 Flash | GLM 5.3 Flash |
|---|---|---|
| Input (standard) | $0.75* | $0.15 |
| Output (standard) | $3.75* | $0.50 |
| Cached Input | $0.075 | $0.03 |
| Promo Price (input/output) | $0.75 / $3.75 (through Dec 31, 2026) | $0.075 / $0.25 (through Sep 9, 2026) |
*Gemini prices double to $1.50/$7.50 on January 1, 2027.
Full Gemini 3.8 Flash Rate Card
| Tier | Input / 1M (thru Dec 31) | Output / 1M (thru Dec 31) | Input / 1M (from Jan 1, 2027) | Output / 1M (from Jan 1, 2027) |
|---|---|---|---|---|
| Standard | $0.75 | $3.75 | $1.50 | $7.50 |
| Batch (50% off) | $0.375 | $1.875 | $0.75 | $3.75 |
| Flex | $0.375 | $1.875 | $0.75 | $3.75 |
Full GLM 5.3 Flash Rate Card
| Tier | Input / 1M | Cached Input / 1M | Output / 1M |
|---|---|---|---|
| List Price | $0.15 | $0.03 | $0.50 |
| Launch Promo (50% off, ends Sep 9) | $0.075 | $0.015 | $0.25 |
| Web Search Add-on | $0.01 per call (billed separately) | ||
Cost at Scale: 10M Input + 2M Output Tokens
For a workload of 10 million input tokens and 2 million output tokens, the cost difference is substantial:
| Scenario | Gemini 3.8 Flash | GLM 5.3 Flash (list) | GLM 5.3 Flash (promo) |
|---|---|---|---|
| 10M in + 2M out | $15.00 | $2.50 | $1.25 |
| Same with 50% cached input | $7.875 | $1.30 | $0.63 |
At production volume, GLM 5.3 Flash costs roughly 6x less than Gemini 3.8 Flash on a blended 3:1 input/output basis, according to LLM Stats analysis.
Consumer Plans and Subscription Tiers
Beyond API access, both ecosystems offer consumer-facing plans that bundle model access with additional features.
Google Gemini Plans
| Plan | Monthly Price (US) | Gemini 3.8 Flash Access | Key Features |
|---|---|---|---|
| Gemini Free | $0 | Limited (rate-limited) | Basic chat, 15 GB storage, 32K context |
| Google AI Plus | $4.99/mo | 2x usage vs Free | 400 GB storage, 128K context, Flash models |
| Google AI Pro | $19.99/mo | 4x usage vs Free | 5 TB storage, 1M context, Deep Research, Gemini in Workspace apps |
| Google AI Ultra (5x) | $99.99/mo | 5x usage vs Pro | 20 TB+ storage, highest limits, Deep Think, agent features |
| Google AI Ultra (20x) | $199.99/mo | 20x usage vs Pro | 20 TB+ storage, maximum limits, YouTube Premium included |
| Google AI Studio | $0 (pay-per-use API) | Free tier with rate limits | Developer playground, API key generation |
Google restructured its consumer plans in 2026, replacing the old AI Premium tier with Plus, Pro, and Ultra options. Google AI Studio still offers a free tier for experimentation, and the paid plans scale primarily on usage limits and storage rather than raw API access.
Z.AI (Zhipu AI) Plans
| Plan | Monthly Price | GLM 5.3 Flash Access | Key Features |
|---|---|---|---|
| chat.z.ai Free | $0 | Access to chat interface | Consumer chatbot, GLM-4.7-Flash is fully free |
| GLM Coding Lite | $18/mo | 3x quota vs flagship | Works in Claude Code, Cline, Cursor, 20+ tools |
| GLM Coding Pro | $80/mo | Higher quota | Increased limits, priority access |
| GLM Coding Max | $168/mo | Maximum quota | Highest limits, team features available |
A notable detail: GLM-4.7-Flash and GLM-4.5-Flash are completely free for all registered Z.AI users, with no per-token charge. GLM 5.3 Flash itself is paid, but the Coding Plan gives it 3x the usage quota compared to the GLM-5.3 flagship. Off-peak hours (including all weekend hours) consume only 50% of standard credits on Coding Plans. Additionally, the GLM-5.3-Flash Usage Campaign (running until September 20) gives paid plan users unlimited GLM-5.3-Flash usage via ZCode daily from 23:00 to 09:00, plus doubled quota on other agents.
Output Speed and Latency
Speed matters for real-time applications. Gemini 3.8 Flash has set a new benchmark in this category.
Token Generation Speed
| Metric | Gemini 3.8 Flash | GLM 5.3 Flash |
|---|---|---|
| Output Speed | 286 tokens/second | 58 tokens/second |
| Cost per Task | $1.24 | $0.25 |
| Speed Advantage | Gemini is ~4.9x faster | |
According to Artificial Analysis, Gemini 3.8 Flash delivers output at 286 tokens per second, nearly 5 times faster than GLM 5.3 Flash’s 58 tokens per second. For applications where latency matters more than cost, this speed gap is significant.
However, speed comes at a price. Gemini’s cost per task ($1.24) is roughly 5 times higher than GLM’s ($0.25), meaning the speed advantage directly correlates with higher per-query costs. For batch processing or cost-sensitive workloads, GLM’s slower but cheaper inference may be the better tradeoff.
Cost per Task: The Real-World Tradeoff
Speed and cost are directly correlated here. Artificial Analysis measures Gemini 3.8 Flash at $1.24 per benchmark task versus GLM 5.3 Flash’s $0.25. That 5x cost difference matches the 5x speed difference almost exactly, meaning developers face a clear tradeoff: pay more for faster responses, or save significantly by accepting slower throughput.
For real-time user-facing applications (chatbots, code assistants), Gemini’s speed advantage may justify the higher cost. For batch processing, data analysis, or high-volume API workloads, GLM’s lower per-task cost makes it the more economical choice.
Both models support configurable thinking levels: Gemini offers low, medium, and high, while GLM offers low, high, and max. Lower effort levels reduce both latency and cost on each platform.
Context Window and Output Limits
Both models offer 1 million token context windows, but they differ in maximum output capacity.
| Feature | Gemini 3.8 Flash | GLM 5.3 Flash |
|---|---|---|
| Context Window | 1,048,576 tokens | 1,000,000 tokens |
| Max Output | 65,536 tokens (64K) | 131,100 tokens (~128K) |
| Input Modalities | Text, image, audio, video, PDF | Text, image, video (native multimodal) |
| Output Modality | Text | Text |
GLM 5.3 Flash offers double the maximum output (128K vs 64K), which is relevant for tasks that require generating long reports, extensive code files, or detailed analyses in a single response.
Open Weights vs. Closed Ecosystem
One of the most significant differences between these models is licensing.
GLM 5.3 Flash is released under the MIT license, making it the first natively multimodal open-weights model in the GLM-5 series. This means developers can self-host, fine-tune, and deploy the model without any API dependency. For organizations with data sovereignty requirements or those building custom deployments, this is a decisive advantage.
Gemini 3.8 Flash is a closed model available only through Google’s API (Google AI Studio, Vertex AI, and the Gemini app). There is no option to run it locally or customize the weights.
If you need a deeper look at how Z.AI achieved strong coding gains in the GLM-5 series, see our article on GLM-5.3’s 6x coding improvements.
Which One Should You Choose?
The right model depends on your priorities. Here is a decision framework based on what matters most:
Choose Gemini 3.8 Flash if you need:
- Fastest output speed: At 286 tokens/second (vs GLM’s 58), it is nearly 5x faster, ideal for latency-sensitive and real-time applications.
- Stronger science and reasoning: Leads on GPQA Diamond (95.3%), Humanity’s Last Exam (47.8%), SciCode (56.6%), and AA-Omniscience Accuracy (55% vs 28%).
- Multimodal input variety: Supports text, image, audio, video, and PDF inputs natively.
- Google ecosystem integration: Direct integration with Google Workspace, Vertex AI, and the broader Google Cloud platform.
- Structured thinking levels: Low, medium, and high effort settings let you balance cost and quality per request.
- Document processing: 89.08% on MMMU Pro (Vals.ai) and stronger performance on GDP.pdf document reasoning (21.0% vs 15.4%).
Choose GLM 5.3 Flash if you need:
- Lower cost at scale: $0.25 per task vs Gemini’s $1.24, and roughly 6x cheaper on a blended token basis at production volume.
- Open-weights deployment: MIT license allows self-hosting, fine-tuning, and full control over data.
- Enterprise and knowledge work: Dominates Elo-based benchmarks: GDPval-AA v2 (1669 vs 1464), AA-Briefcase Elo (1455 vs 1202), and AA-Briefcase Rubric (46.4% vs 42.1%).
- Terminal and agentic tasks: 32.8% on Terminal-Bench v4.0 vs Gemini’s 19.7%, a significant 13-point lead.
- Longer output generation: 128K max output vs 64K for generating extended content in single responses.
- Free fallback model: GLM-4.7-Flash is completely free with no per-token charge, useful for development and testing.
- Budget-friendly coding plans: The GLM Coding Plan starts at $18/month with off-peak hours billing at half credits.
For more on how Gemini 3.8 Flash fits into Google’s broader AI strategy, read our analysis of Gemini 3.8 Flash benchmarks and pricing. For background on Z.AI’s GLM-5 architecture, see our coverage of the Ox Alpha mystery and GLM-5.3-Flash’s open-source release.
Frequently Asked Questions
Is GLM 5.3 Flash really 6x cheaper than Gemini 3.8 Flash?
On a blended 3:1 input/output basis at list prices, yes. GLM 5.3 Flash costs $0.15/M input and $0.50/M output, while Gemini 3.8 Flash costs $0.75/M input and $3.75/M output. The output price difference (7.5x) is the primary driver. During Z.AI’s 50% launch promotion (ending September 9, 2026), the gap widens to roughly 12x.
Does Gemini 3.8 Flash outperform GLM 5.3 Flash on all benchmarks?
No. According to Artificial Analysis v4.3, GLM leads on the overall Intelligence Index (42 vs 41), enterprise benchmarks like GDPval-AA v2 (1669 vs 1464 Elo) and AA-Briefcase Elo (1455 vs 1202), Terminal-Bench v4.0 (32.8% vs 19.7%), and AutomationBench-AA (60.4% vs 59.9%). Gemini leads on science (GPQA Diamond 95.3%, Humanity’s Last Exam 47.8%), omniscience (55% vs 28%), and speed (286 vs 58 tokens/second).
Can I use GLM 5.3 Flash for free?
GLM 5.3 Flash itself is a paid model, but Z.AI offers GLM-4.7-Flash and GLM-4.5-Flash completely free for all registered users. For GLM 5.3 Flash specifically, the GLM Coding Plan starts at approximately $18/month, and off-peak hours (including all weekends) bill at half the standard credit rate.
Which model is faster?
Gemini 3.8 Flash is significantly faster, measured at 286 tokens per second compared to GLM 5.3 Flash’s 58 tokens per second on Artificial Analysis. That is nearly a 5x speed advantage for Gemini. However, this speed comes at a higher cost: $1.24 per task versus GLM’s $0.25.
Will Gemini 3.8 Flash pricing increase?
Yes. The current $0.75/$3.75 per million tokens rate is an introductory price valid through December 31, 2026. Standard pricing of $1.50/$7.50 takes effect on January 1, 2027, doubling the cost.
Can I self-host either model?
Only GLM 5.3 Flash. It is released under the MIT license with open weights, allowing self-hosting, fine-tuning, and custom deployment. Gemini 3.8 Flash is available only through Google’s API and cannot be self-hosted.
