Skip to main content

NVIDIA and Microsoft Unveil RTX Spark for Local AI on Windows

NVIDIA and Microsoft announced RTX Spark hardware and OS-level agent support for Windows, aimed at running AI models locally.

NVIDIA and Microsoft used a joint event in San Francisco to announce new hardware and operating-system features aimed at running AI models and AI agents directly on Windows PCs rather than in the cloud. The centrepiece is RTX Spark, a new class of laptop and compact desktop built around NVIDIA’s Blackwell RTX GPU and Grace CPU. Microsoft also detailed OS-level infrastructure for running AI agents, and previewed a Windows version of its DGX Station workstation.

๐ŸŽฏ Not sure if this will run on your hardware?Use our free Local LLM Hardware Checker โ€” pick your GPU and RAM, see which models will run with real tokens/sec estimates.
Check my hardware โ†’

What was announced

  • NVIDIA CEO Jensen Huang and Microsoft CEO Satya Nadella outlined co-engineered hardware and software for running AI agents on Windows PCs, according to the NVIDIA blog’s account of the event.
  • RTX Spark combines a Blackwell RTX GPU with up to 6,144 cores and an up to 20-core NVIDIA Grace CPU, connected at 600 GB/s, NVIDIA says. It offers one petaflop of FP4 AI performance and up to 128GB of unified memory.
  • NVIDIA says RTX Spark can run models such as “Qwen 3.8 Flash Next,” which it describes as a 125B model with a 51B n-gram component, locally and without sending data to the cloud. No independent benchmark figures were given.
  • RTX Spark runs the full NVIDIA CUDA stack, which NVIDIA says is the same software platform used across its hardware range, from RTX Spark up to DGX Station.
  • Laptop preorders for RTX Spark opened on the day of the announcement, with availability from 16 October. Compact desktop versions go on sale in November, according to NVIDIA.
  • Systems are coming from Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte, NVIDIA says.
  • Microsoft’s Surface Laptop Ultra is built around RTX Spark. Pavan Davuluri, Microsoft’s EVP of Windows and Devices, said it offers “up to 128 gigs of unified memory and up to a petaflop of AI compute,” allowing it to run models that “simply don’t fit on a traditional machine.”
  • A compact desktop configuration of RTX Spark is designed for 24/7 operation, described by NVIDIA as “a dedicated local AI system that keeps agents running continuously.”
  • Microsoft announced general availability of Microsoft Execution Containers (MXC), OS-level infrastructure that Davuluri says lets agents “run safely and persistently in the background, under operating system control.” Microsoft paired this with Microsoft Security and Agent 365.
  • NVIDIA also previewed DGX Station for Windows, described as the first deskside AI supercomputer to bring GB300 Grace Blackwell-class infrastructure to Windows. NVIDIA says it runs on the GB300 Grace Blackwell Ultra Desktop Superchip, with 748GB of coherent memory and up to 20 petaFLOPS of FP4 AI compute, which it says is “enough to run models up to a trillion-parameter scale locally.” Previously, DGX Station ran on Linux only; NVIDIA says Linux AI toolchains remain accessible through WSL. DGX Station for Windows has no confirmed release date โ€” NVIDIA is inviting sign-ups to be notified when it becomes available.
Advertisement

What it means if you run models locally

For readers already running models on their own hardware, RTX Spark represents a jump in unified memory capacity in a consumer laptop or small desktop form factor โ€” up to 128GB, according to NVIDIA. This suggests it could handle larger quantised models than typical consumer GPUs with dedicated VRAM, though NVIDIA has not published specific benchmarks comparing it to existing local setups. Readers working out whether a given model will fit in memory may find the VRAM calculator and quantisation explainer useful starting points once official specs for specific configurations are confirmed.

The fact that RTX Spark runs the full CUDA stack is likely to matter for anyone currently using tools built around NVIDIA’s software ecosystem, since NVIDIA says workflows can move across its hardware range without rewriting. For readers comparing options before buying hardware, the hardware checker and general local LLM background may help frame these new systems against existing choices, such as those discussed in the Ollama local-AI guide.

MXC’s arrival as general-availability Windows infrastructure suggests Microsoft is positioning Windows itself as a platform for persistent, background AI agents, rather than leaving that entirely to third-party tools. It is likely this affects how agent-based local AI software is built and sandboxed on Windows going forward, though the practical experience for end users was not detailed in the announcement.

What is not known yet

Pricing for RTX Spark laptops and compact desktops was not stated in the sources. Detailed system requirements โ€” such as which specific GPU and memory configurations will be available across different manufacturers’ models โ€” have not been published. No independent or third-party benchmark results were given for the claimed one petaflop of FP4 performance or the model NVIDIA cited as an example.

DGX Station for Windows has no confirmed price or release date; NVIDIA describes it only as a preview, with interested parties invited to register for updates. It is also unclear how MXC will interact with existing third-party local AI tools, or what, if any, changes developers will need to make to run agents within this new container system. NVIDIA’s reference to Linux AI toolchains remaining accessible “through WSL when needed” was not elaborated on, so the practical limits of that compatibility are not yet clear.

Sources

This is a news brief compiled with AI assistance from the sources above. Nothing in it has been tested on my own hardware; when I do test something, I say so and show the numbers.

Written by Engineer running a 24/7 homelab since 2022. Every guide here is built and tested on my own hardware. No paid placements.

Keep reading