hn.today

Atrocious AI-Written Tests

gruhn.me3 points0 comments
Screenshot of Atrocious AI-Written Tests

The piece examines a codebase full of AI-written unit tests and documents why many of them are useless or harmful. A single failing test triggered by changing a comment led to discovering checks that scan source files for specific words, tests that assert static configuration objects unchanged, and tests that simply duplicate small bits of production logic in the test code. Other examples include assertions that exact prompt section headings remain present and that a hard-coded list of statuses excludes certain values. Many tests run no production code at all or just mirror implementation logic instead of exercising behavior.

The argument is that these tests offer little value: they don't catch real regressions, they create noisy failures on innocuous edits, and they consume tokens and developer time when generated by AI. Agents tend to add tests to accompany changes for style rather than rigor, and they rarely refactor code to make it more testable (the author suggests, for example, abstracting logging so message assertions can be meaningfully tested). The author found 404 such low-value tests out of 6,429; they did not significantly slow the suite and were easy to remove, but they erode confidence in the test suite and in relying on AI to protect code quality.

Read on gruhn.me0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Programming

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.