For years, running a frontier AI model with a trillion parameters meant booking time in a Linux data center. NVIDIA’s new machine shrinks that class of computer down to something that sits beside a monitor, and it runs Microsoft Windows natively.
The NVIDIA DGX Station for Windows is the first deskside AI supercomputer to bring GB300 Grace Blackwell-class infrastructure directly into the Windows ecosystem. It is aimed at the enterprise developers, researchers, engineers, designers, and data scientists who already spend their day in Windows applications, and it is built to run always-on AI agents locally instead of in the cloud.
Here is what NVIDIA actually announced, what the hardware can really do, what the fine print says about power and price, and who genuinely needs one.
What NVIDIA Actually Announced
NVIDIA unveiled the DGX Station for Windows on June 1, 2026, at GTC Taipei, and said the system is coming in the fourth quarter of 2026. It is built on the same DGX Station system design but tuned for Windows, and it was developed in collaboration with Microsoft.
The pitch is straightforward: heavy AI work such as training, fine-tuning, large-scale inference, and multi-agent development has historically needed data-center systems running Linux, while most large companies run Windows for everyday productivity, design, and engineering. NVIDIA says the new machine bridges that gap.
NVIDIA frames the system around five enterprise workflows:
- AI agents: run multiple frontier agents in parallel and connect them directly to enterprise applications.
- AI development: pretrain and fine-tune large models on Windows, with Linux AI toolchains available through Windows Subsystem for Linux.
- Data science: load large datasets into up to 748GB of coherent memory without data-movement bottlenecks.
- AI inference: run high-throughput inference on models of up to 1 trillion parameters.
- Physical AI: pair the GB300 Superchip with an RTX PRO Blackwell GPU for ray-traced visualization and simulation.
The system is expected to ship from six hardware partners: ASUS, Dell Technologies, GIGABYTE, HP, MSI, and Supermicro. You can read the details in NVIDIA’s press release.
It is also the latest step in a long line. NVIDIA’s DGX brand dates back to the 2016 DGX-1, a turnkey deep-learning server for research labs, and each generation since has moved data-center-class compute closer to the desk, from the Ampere-based DGX A100 systems to the Hopper generation and now the Blackwell Ultra hardware behind this machine.
Inside the Hardware: GB300, 748GB, and 20 Petaflops
At the center of the machine is the GB300 Grace Blackwell Ultra Desktop Superchip. It connects a Blackwell Ultra GPU to a 72-core Grace CPU through NVIDIA’s NVLink-C2C interconnect, so the two chips share one coherent pool of memory instead of shuttling data back and forth.
NVIDIA lists three headline numbers:
- Up to 748GB of coherent memory, a single pool large enough to hold very large models without offloading to a cluster.
- Up to 20 petaflops of FP4 compute, the low-precision format NVIDIA has pushed for inference-heavy workloads.
- Model capacity up to 1 trillion parameters, run locally rather than in the cloud.
NVIDIA also equips the system with a ConnectX-8 SuperNIC that supports networking up to 800Gb/s, and it can be paired with an additional RTX PRO 6000 Blackwell workstation GPU for ray-traced visualization and simulation.
The pool is not one uniform block of memory. NVIDIA’s product page lists it as 252GB of HBM3e tied directly to the GPU alongside 496GB of LPDDR5X serving the CPU side, adding up to the 748GB total.
Some search listings and forum posts reference a 784GB configuration. NVIDIA’s own figures list 748GB of coherent memory, so treat 784GB as a typo or an unconfirmed variant.

748GB of Coherent Memory, Explained
The headline number is not a single uniform block of RAM. Splitting the pool into two tiers is a deliberate trade-off between speed, capacity, cost, and heat.
The 252GB of HBM3e sits closest to the GPU and holds the data that needs the fastest possible access, such as the active weights and activations during inference or training. The 496GB of LPDDR5X is slower but far cheaper per gigabyte, and it holds everything else, including the parts of a large model that are not being actively computed against at any given moment.
That tiering is how NVIDIA can advertise a trillion-parameter ceiling without packing three-quarters of a terabyte of HBM into a desktop chassis, which would be both prohibitively expensive and brutally hard to cool. In practice, a model’s full weight set rarely needs to sit entirely in the fastest tier at once, especially during inference.
The two chips are linked by NVIDIA’s NVLink-C2C interconnect, which keeps CPU and GPU memory addressable as one pool rather than two separate spaces that software has to manage.
What “1 Trillion Parameters” Actually Means
A parameter is a learned value inside a model, and the count is a rough proxy for capability and for how much memory the model needs. A trillion-parameter model sits in the territory of the largest frontier systems, and fitting one onto a single deskside machine is the headline claim here.
The important caveat is precision. Running a model at a lower numeric precision, such as 4-bit, dramatically shrinks the memory it needs, which is exactly why the FP4 figure matters. A “1 trillion parameter” local system and a “1 trillion parameter” data center training run are not the same workload, and the DGX Station is positioned more toward running and fine-tuning models locally than training the very largest ones from scratch.
The practical payoff NVIDIA is selling is always-on agents: software that reasons continuously, calls tools, and connects to enterprise applications. NVIDIA says the system can run hundreds of agents on tasks simultaneously.
What the DGX Station Can Actually Run
Spec sheets are abstract, so it helps to look at the models NVIDIA and Microsoft actually name. Microsoft’s Windows team says the machine runs frontier-class models locally, listing Llama 4 Maverick, Kimi K2.6, and DeepSeek V4 Pro at more than a trillion parameters.
That mix is telling. It spans a widely used open model (Maverick), a strong Chinese open-weights model (Kimi), and a reasoning-heavy model (DeepSeek V4 Pro), which signals that the machine is pitched at running and fine-tuning the leading open models, not only NVIDIA’s own stack. Many of those models are also far cheaper to serve than proprietary frontier APIs, a shift we explored in our look at why Chinese AI models cost so much less.
NVIDIA also frames the system as a “token factory” for teams, saying it can drive 32 or more simultaneous agents for a group of developers, engineers, or researchers and cut how much work has to be sent to the cloud.
NVIDIA DGX Station vs DGX Spark vs the Cloud
The obvious question is how this compares with NVIDIA’s smaller local machine, which we covered in depth, and with renting cloud compute.
| DGX Station for Windows | DGX Spark | Cloud / data center | |
|---|---|---|---|
| Coherent memory | Up to 748GB | 128GB | Scales with rental |
| AI compute | Up to 20 petaflops FP4 | Up to 1 petaflop FP4 | Scales with rental |
| Largest model (local) | Up to 1 trillion parameters | Smaller models | Effectively unbounded |
| Operating system | Windows plus WSL2 | DGX OS (Linux) | Linux |
| Best for | Teams wanting governed local AI | Individual developers | Bursty or very large jobs |
| Availability | Q4 2026, via partners | On sale | On demand |
The wider trend here is clear: the personal mini AI supercomputer category is pushing data center capability toward individual desks and small teams.
Machines such as the AMD Strix Halo 128GB mini supercomputer already let developers run substantial models locally, and the DGX Station is the enterprise-grade extension of that idea. For a full tour of the memory tiers below it, see our local AI hardware ladder.
How It Compares: AMD, Apple, and the Local-AI Race
NVIDIA is not alone in betting that AI compute belongs on a desk. Two credible rivals are chasing the same buyers with very different philosophies.
| System | Maker | Total memory | Software | Availability |
|---|---|---|---|---|
| DGX Station for Windows | NVIDIA | 748GB (252GB HBM3e + 496GB LPDDR5X) | CUDA, Windows plus WSL | Q4 2026 |
| Threadripper Halo Station | AMD | Up to 2.6TB | ROCm | 2027 |
| Mac Studio (M5 Ultra) | Apple | Up to 512GB unified | MLX / Metal | Shipping now |
| DGX Spark | NVIDIA | 128GB unified | CUDA, DGX OS | Shipping now |
AMD’s answer is memory. Its Threadripper Halo Station, shown as a prototype at IFA 2026 and slated for 2027, pairs a 96-core Threadripper PRO processor with AMD Instinct-class HBM accelerators for up to 2.6TB of total system memory, roughly three and a half times the DGX Station’s pool. The catch is software: AMD’s accelerators run on the ROCm stack, which still trails NVIDIA’s CUDA in library support and the volume of existing research code.
Apple’s answer is efficiency. The Mac Studio with the M5 Ultra offers up to 512GB of unified memory and 1.2TB/s of bandwidth in a quiet, low-power box that starts far below either workstation. It cannot match the DGX Station’s memory ceiling or CUDA ecosystem, but for smaller models and local inference it is a fraction of the cost and power draw.
No single system wins on every axis. AMD leads on raw memory, Apple leads on price-to-performance for smaller workloads, and NVIDIA leads on software maturity and the breadth of its OEM partner list. Which one matters most depends on whether a buyer’s bottleneck is memory capacity, budget, or the availability of CUDA-optimized tooling.
Price and Availability: What We Know So Far
NVIDIA has not published an official price for the DGX Station for Windows, and the system ships through hardware partners rather than a public retail checkout, so pricing will vary by configuration and vendor.
For a sense of scale, the smaller DGX Spark was listed in the U.S. store at around $4,699 in August 2026. The DGX Station is a very different class of machine, with roughly six times the memory and far more compute, so expect a substantially higher figure.
On availability, NVIDIA’s product page still lists the system as coming in Q4, which means it is arriving now rather than long on the market. If you are buying, the practical path is to contact one of the six named OEM partners.
The Fine Print: Power, Cooling, and Desk Reality
A machine this dense does not behave like a normal workstation. Reported figures put the power draw at around 1,600 watts, which the same analysis notes would require a dedicated 20-amp circuit. That is the kind of detail that turns a “deskside” purchase into a facilities question.
It also means heat and noise. A system pulling that much power needs serious cooling, so a shared office or a home study may need planning before the box arrives. Buyers should confirm power and thermal requirements with their vendor rather than assuming a standard wall outlet will do.
The Enterprise Angle: Cost, Supply, and Timing
The business case comes down to a comparison most enterprises are already running: rent cloud GPUs or own the hardware?
Renting has been getting more expensive. A September 2026 analysis published on the SemiAnalysis GPU pricing index put B200 cloud rental rates at $8.01 per GPU-hour, up 79% over three months. When hourly rates climb that fast, a fixed upfront cost looks better the longer a model stays in heavy use.
The counterweight is supply. HBM3e, the fast memory at the heart of the system, has been in tight supply across the industry this year, which is exactly the kind of constraint that can delay a launch or cap early orders. It is part of a broader squeeze we have tracked in our reporting on the HBM and DDR5 supply chain. Buyers planning a Q4 purchase should confirm lead times with their vendor before budgeting. For the wider shift toward running models on-premises instead of renting them, see our look at HP’s offline AI push.
Security and the Agent Runtime: Why OpenShell Matters
NVIDIA is pairing the hardware with NVIDIA OpenShell, an open-source, secure-by-design runtime for autonomous agents. It builds on new Windows security and containment primitives and creates an isolated sandbox for each agent.
The design goal is that security and privacy policies live outside the agent’s reach and are enforced at the system level, rather than relying on behavioral prompts an agent could ignore. For enterprises, that is the difference between an agent that might leak a credential and an agent that cannot.

Who Should Actually Buy One
The DGX Station for Windows is not a consumer upgrade. It targets a specific set of buyers:
- Enterprise AI developers who want to build and test agents in the same Windows environment they deploy into.
- Researchers and data scientists who need large models locally for privacy, latency, or cost reasons.
- Engineers and designers who want AI assistance wired directly into 3D design and simulation tools.
- IT teams that want GB300-class compute governed by the Windows security, compliance, and fleet-management tools they already use.
If your workload is small, occasional, or mostly cloud-native, the smaller DGX Spark or a standard cloud instance will usually be the better value. The DGX Station makes sense when data cannot leave the building, when latency matters, or when a team wants its own governed AI node that scales up to a data center when needed.
Frequently Asked Questions
How much does an NVIDIA DGX Station cost?
NVIDIA has not published an official price, and the system is sold through its hardware partners, so the cost depends on configuration. For reference, the smaller DGX Spark was listed at around $4,699 in the U.S. in August 2026, and the DGX Station is a far larger machine.
Is the DGX Station available yet?
NVIDIA announced the DGX Station for Windows in June 2026 and says it is coming in the fourth quarter of 2026. It is arriving now through partner channels.
Where can I buy an NVIDIA DGX Station?
Through NVIDIA’s OEM partners: ASUS, Dell Technologies, GIGABYTE, HP, MSI, and Supermicro. It is not sold as a standard consumer retail product.
Can an NVIDIA DGX Station run Windows?
Yes. The DGX Station for Windows ships with Windows and adds Windows Subsystem for Linux for Linux AI toolchains, so it handles both environments.
How does the DGX Station compare to the DGX Spark?
The DGX Spark is a compact 128GB developer appliance, while the DGX Station is a larger GB300 system with up to 748GB of coherent memory and up to 20 petaflops of FP4 compute, aimed at much bigger models.
Can the DGX Station run a 1-trillion-parameter model?
NVIDIA’s claim is yes: the 748GB coherent memory pool supports models of up to 1 trillion parameters locally. That figure comes from NVIDIA and has not yet been independently benchmarked.
Does the DGX Station still support Linux tools?
Yes. The system keeps access to Linux AI toolchains through Windows Subsystem for Linux, so Linux-based frameworks remain usable without a separate machine or dual boot.
What competes with the DGX Station for Windows?
The closest rivals are AMD’s Threadripper Halo Station, which offers more total memory but runs on the less mature ROCm software stack and targets 2027, and Apple’s Mac Studio with the M5 Ultra, which costs far less but caps out at 512GB of unified memory and lacks CUDA support.
Why does the DGX Station need a dedicated circuit?
The system’s power draw is reported at around 1,600 watts, which exceeds what a standard office circuit can supply continuously, so a dedicated 20-amp circuit is a practical requirement.

The Takeaway
The DGX Station for Windows is less a new gadget than a shift in where serious AI work can happen. By putting GB300-class compute and a trillion-parameter ceiling behind a Windows desktop, NVIDIA is betting that the next wave of enterprise AI will run where employees already work, not only in a distant data center.
The open questions are the ones launch announcements rarely answer: the price, the real-world power bill, and whether enterprises actually want frontier AI on the desk rather than in the cloud. Those answers will arrive with the first shipments.
