Back to News
RSS feedwww.theregister.com

Stanford Professor Promotes Homa as a TCP Alternative for AI Data Centers

Summary

TCP remains the dominant transport protocol for the web and cloud computing, but retired Stanford professor John Ousterhout argues that its stream-based design is poorly suited to latency-sensitive AI data center workloads. He is promoting Homa, a message-based protocol that lets receivers see the expected size of incoming data, manage congestion, and schedule packets with a shortest-remaining-processing-time approach. In a 100 Gbps network running at 80 percent utilization, Ousterhout says Homa reduced p99 latency for short messages to 92 microseconds, compared with 1.2 milliseconds for TCP, while improving performance for long messages by a factor of two. The target workloads include model-weight and gradient transfers, checkpointing, KV-cache movement, metadata coordination, and cache lookups, where even millisecond delays can leave expensive GPUs idle. Homa can run alongside TCP, allowing applications to migrate incrementally; Ousterhout says it can be compiled from GitHub and installed as a Linux kernel module without a reboot. The protocol originated in a 2019 PhD dissertation by Behnam Montazeri, who is now a Google staff engineer. Ousterhout is drafting an IETF standardization document, working to upstream Homa into Linux, and says it was backported to Red Hat Enterprise Linux 8 and 9.5 in March. He is also testing it with a large financial-services company. Homa remains contested: network architect Ivan Pepelnjak criticized its performance comparisons and argued in 2023 that it may be a solution searching for a problem. Other specialized approaches, including DPDK, NVMe-oF, QUIC, RDMA, and Amazon’s Scalable Reliable Datagram, already address parts of TCP’s latency limitations. TCP therefore remains the dominant protocol while Homa’s practicality and adoption are still being evaluated.