hn.today

Mercury 2.5 LLM hits 770 tokens per second

artificialanalysis.ai146 points89 comments
Screenshot of Mercury 2.5 LLM hits 770 tokens per second

Commenters mainly compared raw throughput and cost. bearjaws pointed out Cerebras gpt-oss-120b at ~1400 tokens/sec and ford noted kimi 2.6 at ~1000 tps, suggesting Mercury’s 770 tps is not leading. walrus01 criticized Mercury’s pricing ($0.25 and $0.75) as higher than reputable inference providers for open-weight models that fit under 170GB (deepseek v4 flash, qwen 3.8-flash-next) and called it “probably also stupider than laguna s 2.1.” rvz argued that throughput means little if the model ranks behind frontier AI companies, while copperx summarized the tradeoff succinctly as “good, fast, or cheap; pick two.” hansvm raised a practical latency question: high sustained throughput with a multi-second startup to produce the first batch might still be unusable for some use cases.

Opinion splits over architecture and deployment. nylonstrung dismissed diffusion LLMs as a dead end and noted Google’s limited follow-through, whereas LarsDu88 emphasized scale differences between startups and big labs, speculating the “secret sauce” is training data and that diffusion’s speed benefits vanish at large batch sizes used for multi-user serving. LarsDu88 and others suggested diffusion might still be attractive for on-device, low-latency scenarios (robots), but voiceeh stressed that p99 latency requirements (e.g., <700ms) make such tradeoffs unacceptable for many products.

Read on artificialanalysis.ai89 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.