Gemini 3.8 Flash vs GLM 5.3 Flash: Benchmarks, Pricing, Speed, and Plans Compared (2026)

Date:

The affordable AI model tier has become the most competitive space in the industry. Google’s Gemini 3.8 Flash, launched on September 2, 2026, and Zhipu AI’s GLM 5.3 Flash, released a week earlier on August 26, both target the same audience: developers and businesses who need strong reasoning and coding performance without frontier-model pricing.

This comparison breaks down every dimension that matters: raw benchmark performance, API pricing tiers, consumer subscription plans, output speed, context limits, and real-world value. According to Artificial Analysis v4.3, both models scored within a single point on the Intelligence Index (GLM 42 vs Gemini 41), making this one of the closest matchups in the affordable AI tier.

At a Glance: Key Specifications

The table below summarizes the core specifications of both models as of September 2026.

Specification Gemini 3.8 Flash GLM 5.3 Flash
Developer Google DeepMind Zhipu AI (Z.AI)
Release Date September 2, 2026 August 26, 2026
Architecture Sparse MoE MoE
Context Window 1,048,576 tokens (1M) 1,000,000 tokens (1M)
Max Output 65,536 tokens (64K) 131,100 tokens (~128K)
Input Modalities Text, image, audio, video, PDF Text, image, video (native)
Thinking Levels Low, medium, high Low, high, max
Weights Closed Open (MIT license)
Free Tier Google AI Studio free tier GLM-4.7-Flash is fully free

Performance Benchmarks

Benchmarks reveal where each model excels. Gemini 3.8 Flash leads on long-horizon software engineering and multi-domain reasoning, while GLM 5.3 Flash performs strongly on tool use, agentic tasks, and graduate-level science questions.

Head-to-Head Benchmark Comparison (Artificial Analysis v4.3)

The chart below compares scores on benchmarks where both models have published results.

Benchmark Gemini 3.8 Flash GLM 5.3 Flash Edge
AA Intelligence Index v4.3 41 42 GLM +1
GPQA Diamond (graduate science) 95.3% 91.2% Gemini +4.1
Humanity’s Last Exam 47.8% 39.9% Gemini +7.9
SciCode (scientific programming) 56.6% 51.6% Gemini +5.0
AA-LCR v1.1 (long-context reasoning) 81.3% 80.0% Gemini +1.3
AutomationBench-AA 59.9% 60.4% GLM +0.5
Terminal-Bench v4.0 19.7% 32.8% GLM +13.1
GDP.pdf (document reasoning) 21.0% 15.4% Gemini +5.6
CritPt 18.3% 15.4% Gemini +2.9
AA-Briefcase Rubric (enterprise tasks) 42.1% 46.4% GLM +4.3
EnterpriseOps-Gym-AA 33.2% GLM (no Gemini data)
AA-Omniscience Accuracy 55% 28% Gemini +27

Elo-Based Benchmarks

Benchmark Gemini 3.8 Flash GLM 5.3 Flash Edge
GDPval-AA v2 (knowledge work) 1464 1669 GLM +205
AA-Briefcase Elo (enterprise) 1202 1455 GLM +253
AA-Omniscience Index 30 7 Gemini +23

Vals.ai Benchmarks

Vals.ai provides independent benchmarking across a wide range of professional and technical domains. These scores offer insight into real-world task performance beyond standard academic benchmarks.

Benchmark Gemini 3.8 Flash GLM 5.3 Flash Edge
Vals Index (overall) 62.25% 47.22% Gemini +15.0
GPQA Diamond (graduate science) 94.44% 86.36% Gemini +8.1
SWE-bench (software engineering) 80.00% 92.00% GLM +12.0
LiveCodeBench 89.48% 80.51% Gemini +9.0
Vibe Code Bench v1.1 78.65% 30.76% Gemini +47.9
Terminal-Bench 2.1 81.27% 62.92% Gemini +18.4
Code Migration 36.55% 20.52% Gemini +16.0
MMLU Pro 90.22% 86.06% Gemini +4.2
MMMU Pro 89.08% 86.01% Gemini +3.1
LegalBench 89.48% 83.93% Gemini +5.6
Legal Research Bench 38.94% 45.19% GLM +6.3
Harvey’s Legal Agent 10.00% 6.67% Gemini +3.3
Finance Agent (v2) 61.44% 57.85% Gemini +3.6
TaxEval v2 74.45% 75.59% GLM +1.1
Tax Agent Bench 66.77% Gemini (no GLM data)
MortgageTax 65.34% Gemini (no GLM data)
Public Benefits Benchmark 65.29% Gemini (no GLM data)
MedScribe 84.50% 88.94% GLM +4.4
MedCode 48.13% Gemini (no GLM data)
BioMysteryBench 62.22% Gemini (no GLM data)
CyberBench 43.75% 71.49% GLM +27.7
SAGE 35.06% Gemini (no GLM data)
EMB 72.20% 55.93% Gemini +16.3
ProofBench v1.1 48.00% 21.00% Gemini +27.0
ProgramBench 1.00% Gemini (no GLM data)
SkillsBench 57.98% 40.16% Gemini +17.8

Provider-Reported Benchmarks

In addition to Artificial Analysis’s independent testing, both Google and Z.AI publish their own benchmark results. These come from the providers directly and use different test conditions, but they offer additional data points.

Benchmark Gemini 3.8 Flash (Google) GLM 5.3 Flash (Z.AI)
DeepSWE v1.1 (long-horizon coding) 73.7% 63.4%
Terminal-Bench 2.1 (terminal tasks) 89.4% 84.3%
HLE-Verified (multidisciplinary reasoning) 54.9%
HLE w/ Tools 55.3%
Agents’ Last Exam 26.3%
AutomationBench 48.8%
CharXiv Reasoning 86.2%
LVBench (agentic / static) 87.8% / 87.1%
OSWorld-2.0 59.0%
BioMysteryBench (solvable / difficult) 88.8% / 56.5%
LABBench2 86.2%
Vals Finance Agent v2 61.4%

Note: Provider-reported benchmarks use different test conditions and may not be directly comparable across providers. Artificial Analysis v4.3 scores above provide independent, apples-to-apples comparison.

Key Benchmark Takeaways

  • Overall Intelligence: GLM 5.3 Flash edges ahead on the AA Intelligence Index v4.3 (42 vs 41), while Gemini leads on the Vals Index (62.25% vs 47.22%). These conflicting results show that overall rankings depend heavily on which tasks are tested.
  • Software Engineering: This is where the models diverge most sharply depending on the benchmark. GLM dominates SWE-bench (92.00% vs 80.00%), but Gemini leads on Vibe Code Bench (78.65% vs 30.76%), LiveCodeBench (89.48% vs 80.51%), and provider-reported DeepSWE v1.1 (73.7% vs 63.4%). Terminal-Bench v4.0 also favors GLM (32.8% vs 19.7%).
  • Science and Reasoning: Gemini leads consistently on GPQA Diamond (95.3% vs 91.2% 94.44% vs 86.36%), Humanity’s Last Exam (47.8% vs 39.9%), and SciCode (56.6% vs 51.6%).
  • Professional Domains: Gemini leads on Vals.ai professional benchmarks: Finance Agent v2 (61.44% vs 57.85%), LegalBench (89.48% vs 83.93%), and EMB (72.20% vs 55.93%). However, GLM leads on Legal Research Bench (45.19% vs 38.94%), MedScribe (88.94% vs 84.50%), and TaxEval v2 (75.59% vs 74.45%).
  • Cybersecurity: GLM shows a strong lead on CyberBench (71.49% vs 43.75%), a 28-point advantage that aligns with Z.AI’s focus on security applications.
  • Enterprise and Knowledge Work: GLM dominates on Elo-based benchmarks: GDPval-AA v2 (1669 vs 1464) and AA-Briefcase (1455 vs 1202 Elo), plus the AA-Briefcase Rubric (46.4% vs 42.1%).
  • Domain-Specific: Gemini shows strong results in specialized domains: 86.2% on CharXiv Reasoning, 88.8% on BioMysteryBench, and 61.4% on Vals Finance Agent v2, though Harvey’s Legal Benchmark remains challenging at 10%.
  • Omniscience: Gemini pulls far ahead on AA-Omniscience Accuracy (55% vs 28%) and Index (30 vs 7), indicating broader world knowledge.
  • Automation: The models are nearly tied on AutomationBench-AA (60.4% vs 59.9%), separated by just half a point.
  • Cost per Task: GLM 5.3 Flash costs $0.25 per benchmark task versus Gemini’s $1.24, a 5x cost advantage.

API Pricing Comparison

Pricing is where these two models diverge sharply. GLM 5.3 Flash undercuts Gemini 3.8 Flash by roughly 5x on input and 7.5x on output at list price, and the gap widens further during Z.AI’s launch promotion.

Standard API Rates (per 1 million tokens)

Pricing Tier Gemini 3.8 Flash GLM 5.3 Flash
Input (standard) $0.75* $0.15
Output (standard) $3.75* $0.50
Cached Input $0.075 $0.03
Promo Price (input/output) $0.75 / $3.75 (through Dec 31, 2026) $0.075 / $0.25 (through Sep 9, 2026)

*Gemini prices double to $1.50/$7.50 on January 1, 2027.

Full Gemini 3.8 Flash Rate Card

Tier Input / 1M (thru Dec 31) Output / 1M (thru Dec 31) Input / 1M (from Jan 1, 2027) Output / 1M (from Jan 1, 2027)
Standard $0.75 $3.75 $1.50 $7.50
Batch (50% off) $0.375 $1.875 $0.75 $3.75
Flex $0.375 $1.875 $0.75 $3.75

Full GLM 5.3 Flash Rate Card

Tier Input / 1M Cached Input / 1M Output / 1M
List Price $0.15 $0.03 $0.50
Launch Promo (50% off, ends Sep 9) $0.075 $0.015 $0.25
Web Search Add-on $0.01 per call (billed separately)

Cost at Scale: 10M Input + 2M Output Tokens

For a workload of 10 million input tokens and 2 million output tokens, the cost difference is substantial:

Source: Calculated from official API pricing
Scenario Gemini 3.8 Flash GLM 5.3 Flash (list) GLM 5.3 Flash (promo)
10M in + 2M out $15.00 $2.50 $1.25
Same with 50% cached input $7.875 $1.30 $0.63

At production volume, GLM 5.3 Flash costs roughly 6x less than Gemini 3.8 Flash on a blended 3:1 input/output basis, according to LLM Stats analysis.

Consumer Plans and Subscription Tiers

Beyond API access, both ecosystems offer consumer-facing plans that bundle model access with additional features.

Google Gemini Plans

Plan Monthly Price (US) Gemini 3.8 Flash Access Key Features
Gemini Free $0 Limited (rate-limited) Basic chat, 15 GB storage, 32K context
Google AI Plus $4.99/mo 2x usage vs Free 400 GB storage, 128K context, Flash models
Google AI Pro $19.99/mo 4x usage vs Free 5 TB storage, 1M context, Deep Research, Gemini in Workspace apps
Google AI Ultra (5x) $99.99/mo 5x usage vs Pro 20 TB+ storage, highest limits, Deep Think, agent features
Google AI Ultra (20x) $199.99/mo 20x usage vs Pro 20 TB+ storage, maximum limits, YouTube Premium included
Google AI Studio $0 (pay-per-use API) Free tier with rate limits Developer playground, API key generation

Google restructured its consumer plans in 2026, replacing the old AI Premium tier with Plus, Pro, and Ultra options. Google AI Studio still offers a free tier for experimentation, and the paid plans scale primarily on usage limits and storage rather than raw API access.

Z.AI (Zhipu AI) Plans

Plan Monthly Price GLM 5.3 Flash Access Key Features
chat.z.ai Free $0 Access to chat interface Consumer chatbot, GLM-4.7-Flash is fully free
GLM Coding Lite $18/mo 3x quota vs flagship Works in Claude Code, Cline, Cursor, 20+ tools
GLM Coding Pro $80/mo Higher quota Increased limits, priority access
GLM Coding Max $168/mo Maximum quota Highest limits, team features available

A notable detail: GLM-4.7-Flash and GLM-4.5-Flash are completely free for all registered Z.AI users, with no per-token charge. GLM 5.3 Flash itself is paid, but the Coding Plan gives it 3x the usage quota compared to the GLM-5.3 flagship. Off-peak hours (including all weekend hours) consume only 50% of standard credits on Coding Plans. Additionally, the GLM-5.3-Flash Usage Campaign (running until September 20) gives paid plan users unlimited GLM-5.3-Flash usage via ZCode daily from 23:00 to 09:00, plus doubled quota on other agents.

Output Speed and Latency

Speed matters for real-time applications. Gemini 3.8 Flash has set a new benchmark in this category.

Token Generation Speed

Metric Gemini 3.8 Flash GLM 5.3 Flash
Output Speed 286 tokens/second 58 tokens/second
Cost per Task $1.24 $0.25
Speed Advantage Gemini is ~4.9x faster

According to Artificial Analysis, Gemini 3.8 Flash delivers output at 286 tokens per second, nearly 5 times faster than GLM 5.3 Flash’s 58 tokens per second. For applications where latency matters more than cost, this speed gap is significant.

However, speed comes at a price. Gemini’s cost per task ($1.24) is roughly 5 times higher than GLM’s ($0.25), meaning the speed advantage directly correlates with higher per-query costs. For batch processing or cost-sensitive workloads, GLM’s slower but cheaper inference may be the better tradeoff.

Cost per Task: The Real-World Tradeoff

Speed and cost are directly correlated here. Artificial Analysis measures Gemini 3.8 Flash at $1.24 per benchmark task versus GLM 5.3 Flash’s $0.25. That 5x cost difference matches the 5x speed difference almost exactly, meaning developers face a clear tradeoff: pay more for faster responses, or save significantly by accepting slower throughput.

For real-time user-facing applications (chatbots, code assistants), Gemini’s speed advantage may justify the higher cost. For batch processing, data analysis, or high-volume API workloads, GLM’s lower per-task cost makes it the more economical choice.

Both models support configurable thinking levels: Gemini offers low, medium, and high, while GLM offers low, high, and max. Lower effort levels reduce both latency and cost on each platform.

Context Window and Output Limits

Both models offer 1 million token context windows, but they differ in maximum output capacity.

Feature Gemini 3.8 Flash GLM 5.3 Flash
Context Window 1,048,576 tokens 1,000,000 tokens
Max Output 65,536 tokens (64K) 131,100 tokens (~128K)
Input Modalities Text, image, audio, video, PDF Text, image, video (native multimodal)
Output Modality Text Text

GLM 5.3 Flash offers double the maximum output (128K vs 64K), which is relevant for tasks that require generating long reports, extensive code files, or detailed analyses in a single response.

Open Weights vs. Closed Ecosystem

One of the most significant differences between these models is licensing.

GLM 5.3 Flash is released under the MIT license, making it the first natively multimodal open-weights model in the GLM-5 series. This means developers can self-host, fine-tune, and deploy the model without any API dependency. For organizations with data sovereignty requirements or those building custom deployments, this is a decisive advantage.

Gemini 3.8 Flash is a closed model available only through Google’s API (Google AI Studio, Vertex AI, and the Gemini app). There is no option to run it locally or customize the weights.

If you need a deeper look at how Z.AI achieved strong coding gains in the GLM-5 series, see our article on GLM-5.3’s 6x coding improvements.

Which One Should You Choose?

The right model depends on your priorities. Here is a decision framework based on what matters most:

Choose Gemini 3.8 Flash if you need:

  • Fastest output speed: At 286 tokens/second (vs GLM’s 58), it is nearly 5x faster, ideal for latency-sensitive and real-time applications.
  • Stronger science and reasoning: Leads on GPQA Diamond (95.3%), Humanity’s Last Exam (47.8%), SciCode (56.6%), and AA-Omniscience Accuracy (55% vs 28%).
  • Multimodal input variety: Supports text, image, audio, video, and PDF inputs natively.
  • Google ecosystem integration: Direct integration with Google Workspace, Vertex AI, and the broader Google Cloud platform.
  • Structured thinking levels: Low, medium, and high effort settings let you balance cost and quality per request.
  • Document processing: 89.08% on MMMU Pro (Vals.ai) and stronger performance on GDP.pdf document reasoning (21.0% vs 15.4%).

Choose GLM 5.3 Flash if you need:

  • Lower cost at scale: $0.25 per task vs Gemini’s $1.24, and roughly 6x cheaper on a blended token basis at production volume.
  • Open-weights deployment: MIT license allows self-hosting, fine-tuning, and full control over data.
  • Enterprise and knowledge work: Dominates Elo-based benchmarks: GDPval-AA v2 (1669 vs 1464), AA-Briefcase Elo (1455 vs 1202), and AA-Briefcase Rubric (46.4% vs 42.1%).
  • Terminal and agentic tasks: 32.8% on Terminal-Bench v4.0 vs Gemini’s 19.7%, a significant 13-point lead.
  • Longer output generation: 128K max output vs 64K for generating extended content in single responses.
  • Free fallback model: GLM-4.7-Flash is completely free with no per-token charge, useful for development and testing.
  • Budget-friendly coding plans: The GLM Coding Plan starts at $18/month with off-peak hours billing at half credits.

For more on how Gemini 3.8 Flash fits into Google’s broader AI strategy, read our analysis of Gemini 3.8 Flash benchmarks and pricing. For background on Z.AI’s GLM-5 architecture, see our coverage of the Ox Alpha mystery and GLM-5.3-Flash’s open-source release.

Frequently Asked Questions

Is GLM 5.3 Flash really 6x cheaper than Gemini 3.8 Flash?

On a blended 3:1 input/output basis at list prices, yes. GLM 5.3 Flash costs $0.15/M input and $0.50/M output, while Gemini 3.8 Flash costs $0.75/M input and $3.75/M output. The output price difference (7.5x) is the primary driver. During Z.AI’s 50% launch promotion (ending September 9, 2026), the gap widens to roughly 12x.

Does Gemini 3.8 Flash outperform GLM 5.3 Flash on all benchmarks?

No. According to Artificial Analysis v4.3, GLM leads on the overall Intelligence Index (42 vs 41), enterprise benchmarks like GDPval-AA v2 (1669 vs 1464 Elo) and AA-Briefcase Elo (1455 vs 1202), Terminal-Bench v4.0 (32.8% vs 19.7%), and AutomationBench-AA (60.4% vs 59.9%). Gemini leads on science (GPQA Diamond 95.3%, Humanity’s Last Exam 47.8%), omniscience (55% vs 28%), and speed (286 vs 58 tokens/second).

Can I use GLM 5.3 Flash for free?

GLM 5.3 Flash itself is a paid model, but Z.AI offers GLM-4.7-Flash and GLM-4.5-Flash completely free for all registered users. For GLM 5.3 Flash specifically, the GLM Coding Plan starts at approximately $18/month, and off-peak hours (including all weekends) bill at half the standard credit rate.

Which model is faster?

Gemini 3.8 Flash is significantly faster, measured at 286 tokens per second compared to GLM 5.3 Flash’s 58 tokens per second on Artificial Analysis. That is nearly a 5x speed advantage for Gemini. However, this speed comes at a higher cost: $1.24 per task versus GLM’s $0.25.

Will Gemini 3.8 Flash pricing increase?

Yes. The current $0.75/$3.75 per million tokens rate is an introductory price valid through December 31, 2026. Standard pricing of $1.50/$7.50 takes effect on January 1, 2027, doubling the cost.

Can I self-host either model?

Only GLM 5.3 Flash. It is released under the MIT license with open weights, allowing self-hosting, fine-tuning, and custom deployment. Gemini 3.8 Flash is available only through Google’s API and cannot be self-hosted.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

Sound-Based Acoustic Tractor Beam Remotely Reprograms Material Stiffness

Researchers have built what they describe as an acoustic...

The Strongest Glass in the World: A Complete Comparison

Glass is fragile. At least, that is what most...

Can AI Video Actually Save Time for Solo Creators? A Practical Seedance 2.5 Test

For a solo creator, "saving time" is not the...

A Complete Guide to Nintendo Consoles: Every Model and Generation

Long before Sony entered gaming with the original PlayStation...