Back to News
RSS feedtechstackups.com

MSI Titan 18 HX Review: Running Qwen3.8-Flash-Next Locally

Summary

This review tests whether MSI's Titan 18 HX Dragon Edition can run a useful large language model locally. The laptop combines an Intel Core Ultra 9 285HX, an RTX 5090 Laptop GPU with 24 GB of VRAM, 96 GB of DDR5 RAM, and 6 TB of storage; the large system-memory pool is central to the experiment. The tested model, Alibaba's open-weight Qwen3.8-Flash-Next, has 180 billion total parameters and was run with Unsloth's 2-bit UD-Q2_K_XL GGUF quantization, a download of about 73.5 GB. Runtime use was roughly 21 GB of VRAM and 75 GB of system RAM, while a cold load took 60.1 seconds. The review found that installation was straightforward but required manually increasing the context allocation from an automatic 8,192 tokens; testing used a 32,768-token server allocation. In a Fibonacci coding task, Qwen generated 1,647 tokens in 71 seconds, compared with GPT-5.6 Luna's 177 tokens in eight seconds, but Qwen over-engineered the simple request until its thinking budget was disabled. In a detailed pelican SVG task, Qwen produced the better result while memory peaked at about 89 GB; sustained testing reached 69 degrees Celsius without thermal throttling, although the fans averaged 62.1 dB-A nearby. For a 21-requirement Three.js endless runner, Qwen took 5 minutes 36 seconds and produced the only playable result, while Luna was faster but left the player invisible. The tested configuration costs about $6,200, so the review concludes that local inference is a useful private bonus for an already powerful machine rather than a financially compelling replacement for cloud AI services.