The author ran a structured, paid experiment using large language models to replace a legacy distance-matrix service used in large-scale vehicle routing. The task required computing up to one million route distances in under 100 ms for 1,000 locations, with extreme tight per-route budgets. Over three weeks of full-time work the service was redesigned and implemented in Rust, about half the effort on design/implementation and half on testing, benchmarking and integration. p90 latency on an 8-core c5a.4xlarge was 75 ms for 1,000 locations, 602 ms for 5,000 and 4,837 ms for 10,000. Roughly 14,996 of ~15,000 lines of code were LLM-generated and token spend was about €1,200. The author kept a detailed log of trials and failures to draw general lessons.
The central argument reframes LLMs as semantic-translation machines: tools that translate an existing idea between representations (specification → code, code → summary) rather than inventing new facts or requirements. Translation is probabilistic and accumulates distortion with each hop, which explains hallucinations and why LLMs fail at discovery or fact-finding. Practical value comes from automating repetitive, well-defined engineering tasks - instrumentation, benchmark analysis, refactoring plans and calling Read/Write tool steps - so experienced engineers can move faster. The takeaway: use LLMs to accelerate grunt work and multi-step translation workflows, but not to replace domain judgment or to generate novel requirements or facts.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.