Thibaud Colas from the Wagtail core team describes a month-long experiment to restrict development work to the efficient open model GLM 5.3 Flash. The team consumed about 2 billion tokens in September, but only half of that was processed by the target model; the first half of the month stayed within budget (roughly $68, ~4 kWh, ~365 g CO2) while the second half diverted 1 billion tokens to other models. A vibe-coded prototype picked an unsuitable model and burned through 450 million tokens overnight (~$150, ~5 kWh), and platform capacity limits forced temporary switches to peers like DeepSeek V4.1 Flash and Qwen 3.8 Flash. Benchmarking work across 14 models is ongoing; a sneak-peek table ranks DeepSeek V4.1 Flash highest at 95% accuracy, 14.9 Wh energy use and $0.09 per task.
The write-up turns failures into clear, actionable lessons: implement constant, local measurement of tokens, energy and spend; budget explicitly for experimentation as distinct from routine tasks; improve prompt design and multi-agent orchestration with bounded roles; and prioritize more efficient inference techniques and models. The core recommendation is operational: aim for the majority of day-to-day AI inference to run on flash-tier, cost- and energy-efficient models rather than judging efficiency by token counts alone. An invitation to follow progress at Wagtail Space 2026 closes the piece.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.