Back to News
User submissionfireworks.ai

Fireworks Introduces Ember-1, a Kimi K3-Based Model Designed to Cut Reasoning Costs

Summary

Fireworks Research has introduced Ember-1, a specialized model built on Kimi K3 that is designed to preserve answer quality while using substantially fewer reasoning tokens. The company says users wanted K3’s coding capabilities at lower cost, but simply reducing its reasoning-effort setting caused too much quality loss. Fireworks therefore trained Ember-1 to retain useful reflection and adaptation while removing unnecessary reasoning, especially in multi-turn agentic workloads where earlier reasoning is repeatedly processed. The training program covered mathematics, coding, instruction following, conversation, search, tool use, and software engineering, including both standalone tasks and extended interactions. Fireworks reports more than 50 training experiments and over 200 evaluations, using its Serverless Training platform and no customer data. Across seven benchmarks and two customers’ production traffic, the company says K3 reasoning could be shortened by 35% to 50% without sacrificing accuracy. On Doximity’s physician-validated Bedside Bench, which contains 500 clinical cases across 10 categories, Ember-1 reportedly established a cost-per-task Pareto frontier against both open and closed models. In five additional industry benchmarks, it generally matched K3 at maximum reasoning effort while costing less, and exceeded the low-effort K3 configuration. In two live coding A/B tests, Ember-1 used about 35% fewer tokens per task at comparable quality, while task completion, success scores, and failure rates held or improved; one customer moved it into production. Fireworks also says its own developers did not notice the internal model switch during coding work. Ember-1 is available as a Serverless Research Preview alongside Kimi K3, with two weeks of access under Fireworks’ research-release program. The company also offers training support for enterprises that want customized, token-efficient versions using their own data.