Home Innovation AMD Strix Halo: The 128GB Mini AI Supercomputer Rival to Nvidia DGX...

AMD Strix Halo: The 128GB Mini AI Supercomputer Rival to Nvidia DGX Spark

AMD Ryzen AI Max+ 395 Strix Halo mini PC on a desk with 128GB unified memory concept
AMD Strix Halo packs 16 Zen 5 cores and a 40 CU Radeon 8060S into one chip with 128GB unified memory (Credit: Intelligent Living)

AMD squeezed a workstation and a graphics card into one chip and put 128GB on your desk. The AMD Strix Halo platform, sold as Ryzen AI Max+ 395, pairs 16 Zen 5 CPU cores with a 40-compute-unit Radeon 8060S and up to 96GB of addressable VRAM from a unified LPDDR5X pool, promising to run 70-billion-parameter models locally without a discrete GPU. At roughly half the price of Nvidia’s DGX Spark in street pricing through 2026, it raises a sharp question for developers and creators: is this x86 mini-PC the smarter path to local AI?

This guide answers every top search about Strix Halo, from what it is and how its 256 GB/s unified memory works to real tokens-per-second, gaming equivalence, the 96GB-to-120GB VRAM unlock, and a head-to-head against both the DGX Spark and a discrete RTX 5090 on price and speed.

What Is AMD Strix Halo?

AMD Strix Halo is the enthusiast code name for the Ryzen AI Max 300 series, AMD’s first chiplet APU to combine a high-core-count CPU, a large integrated GPU, and an NPU on one package with a 256-bit LPDDR5X memory interface. The flagship SKU is the Ryzen AI Max+ 395, with lower tiers including the Max+ 392, Max 390, Max+ 388, and Max 385.

Unlike a traditional laptop chip that pairs a small iGPU with a discrete Nvidia GPU, Strix Halo is designed as a standalone platform. The 40 CU Radeon 8060S inside the I/O die handles all graphics and much of the AI compute, fed directly by the unified memory pool. That makes sleek mini-PCs and tablets possible without a separate graphics card, a power brick penalty, or PCIe copying between CPU and GPU memory.

Key architectural points searchers ask about:

  • Zen 5 CPU: Up to 16 cores and 32 threads, boosting to 5.1 GHz with 64 MB L3 cache across two CCDs connected via Infinity Fabric to a large I/O die.
  • Radeon 8060S GPU: 40 Compute Units on RDNA 3.5, clocked up to 2900 MHz, with an estimated 56 teraFLOPS of BF16 peak. Real-world sustained throughput in testing lands near 46 teraFLOPS, about 82 percent of peak.
  • XDNA 2 NPU: Up to 50 TOPS, for 126 total platform TOPS when combined with CPU and GPU. The NPU excels at low-power tasks like noise reduction and background blur and is now being used for Stable Diffusion 3 and disaggregated LLM inference.
  • 256-bit LPDDR5X-8000: 128GB maximum, with about 256 GB/s theoretical bandwidth, around 215 GB/s measured in practice, and up to 96GB allocatable as VRAM on Windows (Linux can reach roughly 110 to 120GB). The allocation is set in UEFI at boot, not at runtime.
  • TDP range: 45W to 120W configurable. Laptops and mini-PCs typically sustain 70W to 120W; handhelds target 45W.

AMD positions Strix Halo for the same local-AI developer audience as DGX Spark, plus gamers and workstation users who want one x86 box that does both. Learn more on the official AMD Ryzen AI Max product page.

Ryzen AI Max+ 395 Specs: CPU, Radeon 8060S GPU and 256 GB/s Memory

The specs explain why Strix Halo can replace a mid-range discrete GPU while still acting as a full workstation. Every major mini-PC in this story uses the same Max+ 395 silicon, so the differences come down to cooling, ports, and price rather than raw compute.

Specification AMD Ryzen AI Max+ 395 (Strix Halo Flagship)
Codename Strix Halo, Gorgon Halo refresh expected late 2026
CPU 16x Zen 5 cores, 32 threads, up to 5.1 GHz, 80 MB L2+L3 cache
GPU Radeon 8060S, 40 CU, RDNA 3.5, up to 2900 MHz, about 56 TFLOPS BF16 peak
NPU XDNA 2, 50 TOPS, 126 platform TOPS total
Memory 128GB LPDDR5X-8000 unified, 256-bit bus, 256 GB/s theoretical, about 215 GB/s measured, up to 96GB VRAM on Windows, 110-120GB on Linux
Process TSMC 4nm, FP11 socket
TDP 45W to 120W cTDP, typically 70W to 120W sustained in mini-PCs
Storage Dual M.2 2280 PCIe 4.0 x4 slots on most mini-PCs, up to 16TB combined
Networking Wi-Fi 7, plus 2.5GbE to dual 10GbE depending on chassis
I/O Highlights USB4, Thunderbolt where equipped, HDMI 2.1, DisplayPort, dual fans and vapor chamber on premium boxes
Other SKUs Max+ 392 (12C/24T, 40 CU), Max 390 (12C/24T, 32 CU), Max+ 388 (8C/16T, 40 CU), Max 385 (8C/16T, 32 CU)

The single most useful number for AI is the 96GB Windows VRAM ceiling, which Linux can stretch to roughly 120GB. Because the pool is unified, the GPU can address almost three times the memory of a 32GB RTX 5090 without tensor parallelism. The trade-off is that LPDDR5X is soldered, you buy 32GB, 64GB, or 128GB up front with no upgrade path, and measured bandwidth of roughly 215 GB/s trails the Spark’s 273 GB/s and far trails a discrete card’s GDDR7.

CPU performance is a bright spot. In testing by The Register’s December 2025 lab, Zen 5 delivered 10 to 15 percent higher throughput than the Spark’s 20-core Arm Grace complex on Sysbench, 7-Zip, and HandBrake, and more than double on double-precision Linpack at 1.6 teraFLOPS versus 708 gigaFLOPS.

Strix Halo vs DGX Spark: Cheaper and Often Faster for Interactive LLM Inference

Both boxes put 128GB on your desk and promise local models to 200B parameters, but they get there with opposite philosophies. Strix Halo is an x86 PC that also does AI, DGX Spark is an AI appliance that can also do PC tasks with workarounds.

Dimension AMD Strix Halo (Ryzen AI Max+ 395, 128GB) NVIDIA DGX Spark (GB10, 128GB)
Platform Zen 5 CCDs + RDNA 3.5 I/O die, XDNA 2 NPU Grace Blackwell GB10, 20-core Arm + Blackwell GPU
GPU Cores 2,560 Stream Processors, 40 CU, 40 RT Cores 6,144 CUDA Cores, 192 Tensor Cores 5th-gen, 48 RT Cores 4th-gen
AI Compute About 56 TFLOPS BF16 peak, about 46 sustained, NPU adds 50 TOPS Up to 1 PFLOP FP4 sparse, about 101 TFLOPS BF16 sustained in MAMF
Unified Memory 128GB LPDDR5X-8000, 256-bit, about 256 GB/s theoretical, about 215 measured 128GB LPDDR5X 8533, 256-bit, 273 GB/s
VRAM Assignable Up to 96GB on Windows, up to 120GB on Linux Coherent 128GB, no split needed
High-Speed Interconnect Thunderbolt 4 or USB4 where equipped, not datacenter class ConnectX-7, 200 Gbps QSFP, designed for clustering two Sparks to 256GB
OS Windows 11 Pro and Ubuntu 24.04 both first-class DGX OS, Ubuntu-based, Linux-native. Windows only on partner RTX Spark variants
Gaming Native x86 gaming, 90 to 100 FPS in Crysis Remastered at 1440p medium in lab testing Runs Crysis via FEX translation, no native 32-bit, not a gaming pick
Dimensions and Weight HP Z2 Mini G1a 85 x 168 x 200 mm, 2.3 kg. Framework 4.5L. GMKtec 0.6 kg compact 150 x 150 x 50.5 mm, 1.2 kg
Street Price Mid-2026 About $1,999 to $4,349 depending on vendor and availability, see buying section below $3,999 launch, now $4,699 Founders, about $4,499 partner builds

For single-user token generation, which is memory-bandwidth bound, the two systems are nearly tied. An independent test on the 120-billion-parameter gpt-oss model measured about 34 tokens per second on a Ryzen AI Max+ 395 versus 38.5 on the Spark, roughly a 13 percent edge for Nvidia, not a generational gap. The Register’s llama.cpp runs went further, showing the HP Z2 Mini G1a slightly ahead of the Spark on decode with the Vulkan backend, because 256 GB/s and 273 GB/s are close enough that neither pulls away.

Comparison infographic of AMD Strix Halo vs Nvidia DGX Spark showing 128GB unified memory, bandwidth, compute, and price
Strix Halo and DGX Spark both offer 128GB unified memory but differ on bandwidth, compute, and price (Credit: Intelligent Living)

Turn to time-to-first-token, where compute matters, and the Spark’s Blackwell GPU is 2 to 3 times faster on a short prompt, widening to roughly 5 times on long prompts as prefill becomes compute-bound. On gpt-oss 120B, the Spark processed prompts at about 1,700 tokens per second versus roughly 340 on Strix Halo.

On larger batch inference with vLLM and Qwen3-30B at BF16 across batch sizes 1 to 64, and on image generation with FLUX.1 Dev in ComfyUI, the Spark’s compute lead, reasserts itself. FLUX.1 at native precision scaled almost linearly with BF16 throughput, giving the Spark roughly a 2.5x lead over the Strix Halo G1a’s 46 teraFLOPS. For full fine-tuning of Llama 3.2 3B, the Spark finished in about two-thirds the time, and for QLoRA on Llama 3.1 70B, about 20 minutes versus just over 50 minutes on Strix Halo in that same lab.

The value story flips when you look at price. Reporting summarized by NotebookCheck and Igor’s Lab from GMKtec’s internal comparisons found the $2,199 EVO-X2 matching or beating the Spark on small to mid-size models and delivering up to 40 percent faster initial response on GPT-OSS 20B and Qwen3 Coder, thanks to unified-memory latency versus the Spark’s throughput-oriented design. On massive 70B dense models, the Spark retains the token-generation crown at batch, but on latency-sensitive chatbots, voice, and frequent model switching, Strix Halo’s lower cost per usable gigabyte wins.

Bar chart comparing AMD Strix Halo vs Nvidia DGX Spark on gpt-oss 120B decode speed and street price
Strix Halo nearly ties DGX Spark on gpt-oss 120B decode speed while costing roughly half as much (Credit: Intelligent Living)

That value sharpens against discrete GPUs too. In the mid-2026 GDDR7 shortage, a single RTX 5090 with 32GB runs about $4,300 to $5,000 on the street, and 32GB cannot hold a 70B Q4 model of roughly 40GB natively, forcing CPU offload that drops it to about 14 to 22 tokens per second. A 128GB Strix Halo mini-PC at $2,199 fits the same 70B model, plus a 120B MoE, entirely in its unified pool at 5 to 15 tokens per second on the dense 70B, for less than the price of the graphics card alone.

For a deeper view of the Spark’s bandwidth math and model-fit tables, see our companion piece, NVIDIA DGX Spark: The 128GB Mini AI Supercomputer on your desk. For the full ladder of local AI options, see our hardware ladder for local AI.

Is Strix Halo Good for Gaming? What Is It Equivalent To?

Yes, and uniquely so for an integrated GPU. This is Strix Halo’s clearest advantage over the Spark: it is a native x86 gaming system without translation layers, driver forks, or missing 32-bit support.

Independent reviews and the Strix Halo laptop roundup data converge on a consistent answer: the Radeon 8060S at a sustained 70W to 80W performs between a mobile RTX 4060 and mobile RTX 4070, varying by title and power envelope. That is about 2.5 to 3 times a Radeon 890M and about twice an Intel Arc 140T, according to the Ultrabookreview test corpus. At 45W it falls closer to a mobile RTX 4050, and with a full 120W cTDP in a well-cooled desktop chassis, it pushes further toward 4070 territory. The Spark, by contrast, can run Crysis Remastered at 1440p medium to playable frame rates via FEX, but it is not a sensible gaming purchase.

Bar chart showing Strix Halo Radeon 8060S gaming performance vs mobile RTX 4060 and 4070
At 70W to 80W, the Radeon 8060S lands between a mobile RTX 4060 and 4070 in native 1080p and 1440p gaming (Credit: Intelligent Living)

Practical gaming notes:

  • 1080p and 1440p are the sweet spot. The 8060S handles modern titles at high settings without needing FSR or DLSS tricks that older games often lack.
  • VRAM headroom matters. With up to 96GB assignable as VRAM, texture-heavy games and creative workloads that punish 8GB dGPUs breathe easily, even if raw shader throughput does not match an RTX 4070 desktop card.
  • TDP matters more than marketing. A 14-inch thin laptop at 45W will not match a Framework Desktop at 120W with a 120 mm fan and six heatpipes. Check sustained wattage, not just the APU name, when comparing laptops.
  • NPU does not help gaming yet. Game engines do not offload to XDNA 2. The 50 TOPS is for AI experiences, not frame rates.

One architectural footnote that readers flag: early reviewers noted that Strix Halo’s measured read bandwidth can trail its write bandwidth, so real-world copy throughput sits around 158 GB/s in AIDA64 despite the 256 GB/s headline. It still leads desktop DDR5 in average bandwidth, but it is why memory-bound games and encodes do not scale perfectly with the theoretical number; a nuance NotebookCheck’s Ryzen AI Max+ 395 spec database helps contextualize against its 55W nominal TDP.

Local LLM Performance: How Fast Is Strix Halo on 7B to 235B Models?

Gaming equivalence is fun, but most buyers compare these boxes on local inference and fine-tuning. Use the Q4 rule of thumb of about 0.5 GB per billion parameters, then add context and KV cache.

  • 7B at 4-bit: About 3.5 to 4 GB, decode around 50 to 80 tokens per second on the EVO-X2’s 215 GB/s real bandwidth.
  • 70B dense at 4-bit: About 35 to 48 GB, fits comfortably with room for long context, decode around 5 to 10 tokens per second dense. Fine-tuning is viable on either platform, though slower than GDDR6 workstation cards.
  • 70B MoE sparse: Because mixture-of-experts models stream only a fraction of weights per token, a nominally 70B-to-120B MoE behaves like a much smaller dense model on bandwidth.
  • GPT-OSS 120B MoE: About 60 to 70 GB at 4-bit, around 31 to 34 tokens per second on Strix Halo mini-PCs in field tests, versus roughly 38 on a DGX Spark in the same independent test.
  • Qwen3 235B MoE: About 115 to 120 GB at 4-bit, around 8 to 11 tokens per second on a 128GB Strix Halo box, the practical single-box ceiling before clustering.

Field reports from GMKtec EVO-X2 owners in late 2025 echo the lab data: with a BIOS update to allow more than 64GB GPU allocation, Ubuntu 24.04, ROCm and Vulkan drivers, users ran gpt-oss:120b briskly and described headroom for RAG pipelines over WireGuard without the box breaking a sweat. That 96GB-to-120GB VRAM ceiling is what makes 120B local inference on a sub-$2,500 mini-PC possible at all.

Unlocking 96GB to 120GB of VRAM: Windows vs. Linux

On Windows, the Radeon 8060S is capped at 96GB of the 128GB pool, set through the BIOS UMA frame buffer and AMD Adrenalin’s Variable Graphics Memory toggle. Linux goes further. Community and vendor guides report that the amdttm.pages_limit and amdttm.page_pool_size GTT kernel parameters, on kernel 6.16.9 or newer, push GPU-allocatable memory to roughly 110 to 120GB on a 128GB system.

The counterintuitive trick is to set the BIOS graphics allocation to its 512MB minimum and let GTT manage the pool dynamically. Convert the target to a page count with (GB × 1024 × 1024) ÷ 4.096, so 110GB becomes 28,160,000 pages. These are user reports and community documentation rather than official AMD guidance, so results vary by BIOS vendor and kernel version.

Diagram of AMD Strix Halo 128GB unified memory split on Windows 96GB vs Linux 110-120GB GPU allocation
Windows caps GPU memory at 96GB, while Linux GTT parameters unlock roughly 110 to 120GB on a 128GB Strix Halo system (Credit: Intelligent Living)

Switching to Linux also lifts throughput. In one owner’s A/B test on a 30B model at Q4, HIP/ROCm on Linux reached about 48 tokens per second on generation versus 36 on Windows Vulkan, roughly a third faster, while prompt processing jumped from 78 to 354 tokens per second. The trade-off is that the ROCm path polls the HSA runtime on about 2.5 CPU cores, so Windows Vulkan was about 35 percent more energy-efficient per token for always-on idle workloads. For a dedicated inference box that spends most of its time generating, the Linux/ROCm path is the faster choice.

Clustering is the remaining gap versus the Spark. Nvidia’s ConnectX-7 lets two DGX Sparks present 256GB for 405B local inference. Strix Halo mini-PCs can network over 10GbE or Thunderbolt, but no vendor ships a turnkey 256GB unified cluster at this size. If you need that single 405B pool in a palm-sized form factor today, the Spark pair remains the only off-the-shelf path.

Mini PCs and Laptops You Can Buy: Complete Strix Halo Buying List

Every device below uses the same Ryzen AI Max+ 395-class silicon, but availability in the second half of 2026 has been the story. A DRAM shortage pushed street prices up and left the cheapest boxes unavailable even while list prices looked attractive. Prices below were live-checked in early August 2026 across vendor stores and cross-checked against ServeTheHome and retailer listings.

Device Form Factor Max RAM 128GB Price in Early August 2026 Stand-Out Feature Availability Then
GMKtec EVO-X2 Ultra-compact mini-PC, 0.6 kg 128GB LPDDR5X-8000 $1,999.99 list, $2,199.99 top config Lowest list price, dual USB4, dual M.2 2280 up to 16TB All 128GB variants unavailable
Framework Desktop 4.5L Mini-ITX desktop, 6.85 lbs 128GB, 192GB coming soon $3,449 for 128GB DIY, $1,959 for 64GB Serviceability, FlexATX PSU, 120 mm fan, Wi-Fi 7, 5GbE Out of stock
Minisforum MS-S1 Max Compact desktop, internal 320W PSU 128GB $3,639, list $4,549 Dual 10GbE, USB4 v2 80 Gbps x2, PCIe 4.0 x4 slot In stock, late-August ship
AMD Ryzen AI Halo Developer System Reference mini-PC 128GB/2TB $3,999 direct from AMD AMD Developer Center image, first-party support Pre-order, direct from AMD
Beelink GTR9 Pro Compact desktop, 230W internal PSU 128GB/2TB $4,349, list $4,699 Dual 10GbE, vapor chamber, dual USB4, quietest under load Pre-sale, about 35-day ship
HP Z2 Mini G1a Workstation mini-PC, 2.3 kg 128GB About $2,950 as tested in late 2025 lab ISV certified, Flex IO modules, 2x Thunderbolt 4 Enterprise channel
ASUS ROG Flow Z13 GZ302 13.4-inch tablet, 1.2 kg tablet only 128GB in select markets From $2,099 for 32GB, 64GB and 128GB only in some regions 2.5K 180Hz touch, vapor chamber, 200W charger Most retail configs 32GB
ASUS ProArt PX13 13.3-inch 2-in-1 OLED 128GB From $1,599 for Max+ 388 variant 2K OLED touch, 360-degree hinge Available via ASUS store
Lenovo Legion 7a 15 / Yoga Pro 7a 15.3-inch OLED clamshell 128GB From $2,599 for Max+ 388 config 2.5K 120Hz OLED, dual M.2, 180W USB-C Rolling availability through late 2026
GPD WIN 5 / ONeXFly Apex Gaming handhelds, 7 to 8 inch 128GB From $1,450 to $1,600 Portable Strix Halo with swappable batteries Import and boutique channels

Bottom line from that same early-August snapshot: identical silicon ranged from $1,999.99 to $4,349, and the two cheapest were unbuyable, making the in-stock Minisforum MS-S1 Max at $3,639 the de facto pick that week. If you are buying today, verify stock and port needs first. A 2.5GbE-only box is fine for single-user inference, but teams serving models will want dual 10GbE.

One ownership footnote from ServeTheHome’s November 2025 Framework Desktop review: Framework ships as a DIY kit that requires assembly, even if you buy the SSD and fan separately. The $1,999 base looks attractive until you add tiles at $10 per pack, a $40 translucent side panel, the Noctua cooler, and build time. STH rated it their third-favorite Strix Halo mini-PC for that reason, not for silicon quality, a reminder to budget the full configured price rather than the base number.

Price, Value, and Who Should Choose Strix Halo

Strix Halo’s pitch is simple: x86 compatibility, real gaming, and credible local AI for thousands less than a DGX Spark when you can actually buy one at list price. The wrinkles are availability, soldered RAM, and the CUDA versus ROCm software gap.

Choose Strix Halo if you:

  • Want one box that games natively and runs 70B to 120B local models without a discrete GPU. At 70W to 80W the 8060S sits between a mobile RTX 4060 and 4070, no translation needed.
  • Value Windows plus Linux flexibility. Strix Halo runs both first-class, while DGX Spark is Linux-native.
  • Are price-sensitive and latency-sensitive. On small to mid-size models and chatbot workloads with frequent model switching, field tests show up to 40 percent faster time-to-first-token than Spark per dollar.
  • Need up to 120GB of VRAM on a mini-PC. That capacity does not exist on any 32GB consumer GPU, and the Linux GTT unlock stretches beyond Windows’ 96GB cap.

Choose DGX Spark if you:

  • Need maximum prefill and time-to-first-token on long prompts or large-batch serving, where Blackwell’s 101 TFLOPS BF16 sustained leads.
  • Need a turnkey 256GB cluster for 405B inference via ConnectX-7, or you live in the CUDA ecosystem without wanting to build ROCm forks from source.
  • Generate images and video at scale in ComfyUI with FLUX.1 and similar native-precision models, where the Spark was about 2.5 times faster than the HP G1a in the same lab.
  • Want a validated DGX OS stack with NGC containers out of the box.

Pricing context matters. At launch, Strix Halo looked like a clear undercut. By mid-2026 the DRAM shortage narrowed the gap, with in-stock Strix Halo boxes landing near $3,639 to $4,349 direct, against Spark at $4,499 to $4,699. The per-gigabyte value still favors AMD when you find a $1,999 EVO-X2 at list price, but street reality has been in-stock premiums for both platforms.

The software moat is the other variable. The Register’s late-2025 hands-on found most PyTorch scripts eventually ran on Strix Halo, but vLLM, bitsandbytes, and Flash Attention 2 often required ROCm forks or compiling against gfx1151, while Spark’s CUDA stack worked without modification. Both vendors now ship Docker containers for vLLM, which closes the gap for batch serving, though single-user Llama.cpp on Strix Halo via Vulkan is already competitive on decode.

Frequently Asked Questions

What is AMD Strix Halo?

AMD Strix Halo is the code name for the Ryzen AI Max 300 series APU that combines up to 16 Zen 5 CPU cores, a 40 CU Radeon 8060S GPU, and a 50 TOPS XDNA 2 NPU on one chip with a 256-bit LPDDR5X interface and up to 128GB unified memory. It powers mini-PCs, tablets, and laptops built to run large AI models and games without a discrete GPU.

Is Strix Halo good for gaming?

Yes. At 70W to 80W sustained, the Radeon 8060S performs between a mobile RTX 4060 and mobile RTX 4070 in native x86 games, about 2.5 to 3 times a Radeon 890M. It is native x86, so unlike DGX Spark, it does not need translation layers for games. Expect strong 1080p and 1440p, with VRAM headroom up to 96GB on a 128GB system.

What is Strix Halo equivalent to?

On the GPU side, roughly a 75W mobile RTX 4060 to 4070 depending on chassis cooling and cTDP. On the CPU side, the 16-core Zen 5 complex slightly leads the Spark’s 20-core Arm Grace in Sysbench, 7-Zip, and HandBrake and doubles it in double-precision Linpack in independent lab testing. On total platform AI TOPS, AMD quotes up to 126 for the Max+ 395, with about 56 TFLOPS from the GPU and 50 TOPS from the NPU.

Where can I buy an AMD Strix Halo?

Current buying options include the GMKtec EVO-X2, Framework Desktop, Minisforum MS-S1 Max, Beelink GTR9 Pro, HP Z2 Mini G1a, ASUS ROG Flow Z13 and ProArt PX13, Lenovo Legion 7a and Yoga Pro 7a, plus handhelds like GPD WIN 5. In early August 2026, the GMKtec and Framework were out of stock with 35-day or longer waits, while the Minisforum was in stock at $3,639 for 128GB/2TB. Availability shifts weekly during the DRAM shortage, so check vendor stores directly. The official AMD list of Ryzen AI Max systems links to partners.

How much RAM can Strix Halo use as VRAM?

On a 128GB system, Windows can allocate up to 96GB as VRAM to the Radeon 8060S through the BIOS UMA frame buffer and Adrenalin’s Variable Graphics Memory, leaving 32GB for the OS. On Linux, community reports describe unlocking roughly 110 to 120GB with the amdttm GTT kernel parameters on kernel 6.16.9 or newer. The allocation is set at boot, so reboot to change it, and update your BIOS first, since older versions capped the allocation at 64GB.

Can Strix Halo run a 70B model locally?

Yes, easily. A 70B dense model at 4-bit needs about 35 to 48 GB, leaving ample room for context windows and KV cache. A 120B MoE model needs about 60 to 70 GB and runs at about 31 tokens per second in field tests, while a 235B MoE like Qwen3 235B needs about 115 to 120 GB and sits near 8 to 11 tokens per second on a 128GB box.

Is Strix Halo better than DGX Spark?

It depends on the metric. Strix Halo is cheaper, about half the price of the Spark, and nearly tied on single-user decode speed because both sit near 256 to 273 GB/s bandwidth, with field tests showing up to 40 percent faster initial response on small to mid-size models. It is also better for native x86 gaming and Windows plus Linux versatility. DGX Spark is better for long-prompt prefill, large-batch vLLM serving, 256GB clustering for 405B inference, and turnkey CUDA software maturity. For FLUX.1 image generation at native precision, the Spark held about a 2.5x lead in the same lab.

Conclusion

AMD Strix Halo makes local AI on a mini-PC feel ordinary in the best way. The Ryzen AI Max+ 395’s 16 Zen 5 cores, 40 CU Radeon 8060S, and 256 GB/s unified memory deliver a credible RTX 4060-class gaming box and a 96GB VRAM platform for 70B to 120B local models in the same chassis, all as a standard x86 PC that runs Windows and Linux without translation.

Nvidia’s DGX Spark answers a different question more loudly: how fast can you prefill long prompts and serve large batches at datacenter software maturity, and with ConnectX-7, how far can you stretch to 405B? If that is your daily work, the Spark’s Blackwell GPU and CUDA stack justify the premium.

If you want one quiet mini-PC that games, codes, and runs a private 120B assistant on your desk, Strix Halo is the companion to the Spark era that finally makes cloud-free AI routine, especially if you can secure one near its $1,999 list price before the next DRAM-driven repricing.