Blob-stream is an open-source, Kafka-like streaming system designed to minimize total cost of ownership for high-throughput workloads by trading ultra-low latency for much lower operational and cloud costs. It exposes familiar concepts - topics, partitions, producer batching, and consumer groups with commits and seeks - via Rust producer/consumer libraries, but intentionally diverges from Kafka to eliminate cross-availability-zone network traffic, make brokers stateless with zero local storage, allow instant broker rotation and autoscaling, and remove the need for a separate control plane beyond Kubernetes. The system emphasizes observability and cost efficiency and accepts some semantic differences (for example, it currently cannot guarantee a single global partition order for a given key).
Architecturally, blob-stream relies on accurate clock synchronization (e.g., AWS Time Sync) and AZ-local virtual partitions: each logical partition is split into AZ-local virtual partitions that consumers read globally. Brokers allocate large offset blocks in memory (allowing gaps after restarts), write contiguous segments directly to blob storage (S3) to avoid disk management, and record segment metadata in DynamoDB using a primary key time window and Snowflake-ordered range keys so all brokers can write without transactions. Consumer coordination uses DynamoDB membership, a self-elected planner for sticky assignments, and per-partition leases. Trade-offs include S3 PUT cost/latency considerations and the requirement for synchronized clocks.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.