Back to News
RSS feeddevblogs.microsoft.com

Microsoft Expands Windows ML for Local GGUF and ONNX AI Inference

Summary

Microsoft has added experimental llama.cpp support to Windows ML, allowing developers to run GGUF language models locally through task-specific APIs alongside existing ONNX workflows. The first release includes a Text Generation API for GGUF and ONNX language models and a Speech Recognition API for ONNX Whisper models; the two can be combined for voice-driven applications. Windows ML automatically selects an execution engine, and its local server exposes an OpenAI-compatible endpoint so developers can prototype with the OpenAI SDK. A new experimental Windows ML Runtime API provides a lower-level Windows-native path with direct handling of images, video, audio, and text, zero-copy data paths, explicit CPU/GPU/NPU placement, deterministic multi-model pipelines, and precompiled model artifacts. Existing ONNX Runtime APIs remain supported alongside the new runtime path. Microsoft also describes contributions to llama.cpp with NVIDIA and the broader community, including CUDA kernel optimization, kernel fusion, CPU-GPU scheduling improvements, weight repacking, CUDA graphs, speculative decoding methods, multi-GPU execution, NVFP4, additional model architectures, and backend sampling. For model development, official native PyTorch Arm64 CPU builds, NVIDIA CUDA-enabled Windows Arm64 packages, and Windows support for Triton extend training, fine-tuning, compilation, and inference on Arm-based systems. The article demonstrates compiling PyTorch operations with Triton, exporting a model to portable ONNX, then using the Windows ML CLI to analyze, optimize, build, and benchmark it. Microsoft says future work will target broader native package and kernel coverage, easier installation, better performance, and clearer support across Windows x64 and Arm. The new Windows ML capabilities are experimental, so the article advises checking supported scenarios and known limitations before production use.