Ember-1 is a specialized model from Fireworks Research that reproduces Kimi K3’s quality while emitting roughly 40% fewer tokens by learning to prune unnecessary internal reasoning. Built on Kimi K3 with more than 50 training experiments and 200+ evaluations, Ember-1 uses new training algorithms and Serverless Training to teach the model to keep the reasoning that matters and avoid unproductive reflection. Training data covered math, coding, instruction following, conversation, tool use and extended agentic interactions so the efficiency gains generalize across standalone and multi-turn tasks. Evaluation included the new Specialized Intelligence Index, public benchmarks and production traffic; on Doximity’s Bedside Bench Ember-1 established a new cost-versus-performance Pareto frontier versus open and closed models including GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5.
Quantitatively, Ember-1 shortens reasoning by about 35-50% without hurting accuracy, and live A/B tests with two customers showed roughly 35% fewer tokens per task with comparable or improved downstream metrics, prompting one customer to adopt it in production. Internal use was effectively invisible to developers while cutting token spend. Ember-1 is available as a Research Preview on Serverless with two-week research releases and enterprise training support to tailor token-efficient models to specific workloads, targeting agentic coding and other cost-sensitive applications.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.