Strata Runs Qwen3.8-Flash-Next Locally on Consumer PCs
Summary
Strata is an open-source inference engine and one-click setup that runs the 125-billion-parameter Qwen3.8-Flash-Next model on a supported Windows or Linux PC instead of requiring a server. The project targets NVIDIA RTX 20-series or newer GPUs, generally with at least 12 GB of VRAM, 64 GB of RAM for all listed sizes, and about 80 GB of free disk space; experimental Linux support is available for several AMD Radeon cards. It uses quantized variants ranging from Q2_0 to IQ3_S, requiring about 37.6 to 54.8 GB of combined RAM and VRAM. On an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600, and 64 GB of RAM, measured generation speed ranges from 46 to 93 tokens per second depending on quantization and context, while prompt processing reaches up to 2,170 tokens per second in the published table. Strata can distribute layers across multiple NVIDIA GPUs, keep model data across RAM, VRAM, and SSD storage, and use speculative decoding to improve generation speed. The model offers up to 262K context in its standard configuration, with experimental rope-scaling options for 384K and 512K contexts. A coding variant removes half of the experts and is intended for coding, tool use, and image-related workloads, while Swift 1.5 is a shorter-thinking fine-tune with its own license. Installation downloads roughly 70 GB of model data and opens a local browser interface at port 8080. The service also exposes OpenAI-compatible and Anthropic-compatible endpoints, supports optional image input and configurable thinking effort, and can be connected to coding agents. Strata is designed for one request at a time, and its first startup may temporarily make the computer unresponsive while loading tens of gigabytes into memory. The engine is MIT-licensed, but the model files and some included components retain their own licenses.