Back to News
RSS feedgetrange.sh

Range Runs Large AI Environments Without Downloading Them in Full

Summary

Range is a personal project that opens a shell in a container image, a Hugging Face repository, or an artifact hosted on S3 or an HTTP server without downloading the entire source first. It exposes each source as a disk and turns file reads into ranged network requests, while keeping writes in a local overlay. Container images are indexed once, model repositories are pinned to a Hugging Face commit, and later sessions can prefetch blocks recorded as part of the previous working set. The article demonstrates running llama.cpp with a 270-million-parameter Gemma model from a 6.38 GB repository: the first empty-cache run produced an answer in 15.5 seconds, compared with 18.3 seconds for Docker plus a Hugging Face download, while an indexed run took 6.7 seconds and moved 317 MB. In another test, reading the configuration, one shard header, and one tensor from the 1.03 TB, 61-shard Kimi K2 repository took 3.4 seconds and transferred 9.5 MB; the other shards stayed remote. The tool does not avoid downloading data that a program reads in full. Benchmarks on EC2 also covered Python, Rust, and Java images, with indexed runs transferring tens to 125 MB instead of hundreds of megabytes. Range runs natively on Linux and uses a self-managed Lima VM on macOS; Linux requires root and kernel modules for NBD, EROFS, and overlay filesystems. It is distributed as release archives or source code and can build and publish its own artifacts to S3.