SlopShape trains a detector that distinguishes AI-generated commercial web content from human writing by analyzing structural signatures - how information is ordered, framed, evidenced and voiced - rather than relying on word-level cues. The study constructs a dataset of 2,250 pre-ChatGPT human blog posts from 268 company domains and 11,250 AI "mirrors" generated by five frontier models. A 214-feature annotation instrument is applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.928, human-model 0.946). From 187 structural features alone, the classifier achieves 98.0 macro-F1 on held-out companies, with performance unchanged (98.1) when each AI post is reworded by its originating model.
The structural signal both characterizes and attributes content: AI-authored posts tend to share a tidy, self‑announcing shape, human posts occupy rarer structural configurations, and source attribution reaches 79.3% accuracy against a 16.7% chance baseline. Results replicate and amplify prior StoryScope findings for fiction, showing consistent direction and larger magnitude on commercial content. The author releases the full pipeline, instrument, prompts, code and aggregate artifacts to support replication and further study.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.