hn.today

We found defects in 37 of DeepSWE's 113 tasks

scrimdata.com8 points0 comments
Screenshot of We found defects in 37 of DeepSWE's 113 tasks

A systematic review of DeepSWE v1.1’s recorded failures for three top models (Opus 5, Sol and Fable 5) found defects or ambiguous requirements in 37 of 113 tasks (32.7%). The benchmark is used in launch tables that report scores like Astra 74.1%, Opus 73.7%, Sol 72.7% and Fable ~69.9%, so small percentage differences matter. Inspecting 372 failures, reviewers flagged concrete evaluator problems: hidden tests injected during scoring created Go function-name collisions that broke builds, wording constraints (tests requiring the literal word “invalid”) rejected otherwise correct warnings, and formatting tests enforced styles (e.g., “IN(” vs “IN (”) the prompts did not specify. The review categorized 10 tasks as confirmed evaluator defects, 21 as ambiguous requirements, and 6 with both issues; 76 tasks had no identified problem. Examples included rejected valid Markdown links, unit mismatches (5s vs 5000 ms), and missing normalization rules.

Repairing confirmed evaluator defects and rerunning archived submissions converted 67 failures into passes without changing code and raised measured pass rates by 4.22-6.19 percentage points (Opus 73.65%→78.38%, Sol 72.67%→76.89%, Fable 69.72%→75.92%). The piece argues that these scoring errors and unstated requirements make benchmark percentages unreliable indicators of model capability when used without inspecting failures, warning that optimizing to satisfy flawed evaluators risks training models to game tests rather than solve problems. Review methods combined Codex assistance with human expert adjudication of flagged cases.

Read on scrimdata.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in Security

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.