Kafila:让一组互相信任的异构个人电脑协作运行大语言模型
Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
一群朋友的旧电脑拼起来也能跑大模型了,Kafila 论文说比 exo 快 3 倍,还支持跨洲组网,搞分布式推理的可以细看。
arXiv 论文提出 Kafila,一种在互相信任的多台异构消费级设备上协同服务大语言模型的系统。它通过协议在 NAT 后组建设备环,并按每台设备的内存带宽、容量和可达性精确切分模型。在跨两大洲的五台设备等三个机群上,相比 GPipe 式均匀切分,Kafila 将最慢流水线阶段缩短最多 5.2 倍;相比 exo 式按内存比例切分,最多缩短 3 倍。在共享网络场景下,吞吐量达到均匀切分的 1.56 倍,四个并发用户时领先扩大到 3.2 倍。
Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding membership removes what they depend on: a swarm holds each part of the model on several peers and routes around a slow one. A bounded session must use every device it admits. Its pipeline advances at the pace of whichever device received a share it cannot serve quickly, so the division has to be right before serving begins. We propose Kafila, whose protocol assembles a ring from behind NATs, preferring direct paths and relaying where traversal fails, while its planner measures each device's memory bandwidth, capacity and reachability, divides the model exactly for a fixed ring order, and places the head, which holds the embedding and output projection, together with that division rather than beforehand. On machines with different capabilities across three fleets, from a shared LAN to five devices spanning two continents, Kafila shortens the slowest pipeline stage by up to $5.2\times$ against the even split of pipeline parallelism, as in GPipe, and up to $3\times$ against the memory-proportional split of personal-device inference, as in exo, keeps 75 to 87 per cent of the committed hardware doing work where those divisions fall below half, and serves a model no uniform split can place on the fleet at all. What that is worth to a user depends on how much of a token is computation rather than network. Where the members share a network the same division returns $1.56\times$ the throughput of a uniform split and $1.25\times$ of a memory-proportional one, and under four concurrent users that lead compounds to $3.2\times$ rather than fading, each user served at almost the rate of one.