hn.today

Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLM

nexlab.net12 points3 comments
Screenshot of Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLM

This compares the self-hosted inference orchestrators available in September 2026, cataloguing modalities, multi-machine capabilities, auto-discovery, cache-aware routing, ops consoles, cloud bursting, non-LLM fan-out, training, Kubernetes support, platforms and image signing. It treats llama.cpp and vLLM as the core engines many stacks wrap (layer/row split, Ray-based TP/PP), notes LiteLLM is a routing/proxy layer not a runtime, and presents a feature table and per-project summaries. Key technical details: LocalAI added a P2P distributed mode (libp2p/discovery, NATS v3 router, cosign-signed backend images, prefix-cache-aware routing); exo leverages MLX and RDMA over Thunderbolt 5 to scale Apple Silicon devices; GPUStack and Xinference are enterprise supervisior+worker consoles with multi-vendor accelerator support, UIs, metering and Prometheus/Grafana integration; NVIDIA Dynamo/LLM-D targets datacenter KV-aware routing on Kubernetes.

Recommendations tie capability to scenarios: Ollama/Open WebUI is the simple one-machine choice; run vLLM directly for single-model maximum throughput (with LiteLLM as a shared-router); exo is the clear Apple Silicon multi-device option; LocalAI is the broad, community-backed Linux choice for multi-modality without cloud escalation; GPUStack/XInference suit departmental clusters with dashboards and multi-tenant features; Dynamo/LLM-D is for rack-level K8s deployments; SkyPilot/DStack schedule jobs across clouds rather than per-request bursting. A newer option, CoderAI, combines multi-modality endpoints, prefix-cache routing, mDNS discovery, per-model escalation to rented GPUs, distributed LoRA training and non-LLM fan-out, but lacks Kubernetes, Apple Silicon support and a mature community.

Read on nexlab.net3 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.