The piece argues for giving large language models compact, well-structured DSLs to edit instead of letting them free-form author prompts or touch sprawling nested state. The workflow is capture → edit → translate back: snapshot the current state as a small DSL, have the model edit that DSL, then deterministically translate the DSL back into the real configuration. Practical benefits show up in tools like a hook system where the model edits concise block_command(...) rules and a translator writes them back into settings.json so the model never manipulates JSON directly. That pattern avoids the failure modes of free-form self-prompting while leveraging the model’s strength at writing code-like snippets.
A second example is a context compiler that builds a unified IR for prompts, data and examples, then runs cache-aware optimization passes that trade token savings against lost cache hits to optimize for COST, QUALITY or SPEED. Key specifics: no-passes requests total 9,464 tokens (8,564 read from cache); COST mode sends 8,804 tokens with 8,564 cached; images are handled by selective inclusion and a sidecar scorer to decide per-turn whether to attach a figure. One figure cut to 28px patches costs ~1,296 tokens; re-sending every figure over nine turns costs 23,328 tokens vs 11,664 if only the needed figure is sent (≈2×). A tiny scoring function (similarity + NLI + keyword bonus, threshold 65) proves cheap and effective after sandbox-driven inlining, showing that small DSLs and micro-services produce large practical savings and more reliable model behavior.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.