Commenters mainly compared raw throughput and cost. bearjaws pointed out Cerebras gpt-oss-120b at ~1400 tokens/sec and ford noted kimi 2.6 at ~1000 tps, suggesting Mercury’s 770 tps is not leading. walrus01 criticized Mercury’s pricing ($0.25 and $0.75) as higher than reputable inference providers for open-weight models that fit under 170GB (deepseek v4 flash, qwen 3.8-flash-next) and called it “probably also stupider than laguna s 2.1.” rvz argued that throughput means little if the model ranks behind frontier AI companies, while copperx summarized the tradeoff succinctly as “good, fast, or cheap; pick two.” hansvm raised a practical latency question: high sustained throughput with a multi-second startup to produce the first batch might still be unusable for some use cases.
Opinion splits over architecture and deployment. nylonstrung dismissed diffusion LLMs as a dead end and noted Google’s limited follow-through, whereas LarsDu88 emphasized scale differences between startups and big labs, speculating the “secret sauce” is training data and that diffusion’s speed benefits vanish at large batch sizes used for multi-user serving. LarsDu88 and others suggested diffusion might still be attractive for on-device, low-latency scenarios (robots), but voiceeh stressed that p99 latency requirements (e.g., <700ms) make such tradeoffs unacceptable for many products.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.