Researchers analyze 450,000 gender-directed completions from 15 models in the OpenAI GPT lineage (GPT-2 through GPT-5) and argue that safety training has tended to transform discriminatory content rather than remove it, a phenomenon they call "harm laundering." Using three demographic conditions and multiple automated measures, they show that surface-form toxicity declines even as representational harms shift form: explicit abuses decline but are replaced by subtler, gendered distortions in topic and sentiment. They formalize harm laundering into a three-criteria test and propose a three-stage detection protocol applicable to any generative model, arguing that standard toxicity metrics are insufficient proxies for overall harm.
Empirical findings include the disappearance of sexual-violence clusters targeting women by GPT-4 while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed outputs do not. At GPT-5, a 1,997-document cluster frames breast cancer as a men's-rights debate with no equivalent in women-directed output; three independent classifiers label this content non-toxic. Sentiment flips at GPT-4 from demeaning women to over-correction, topic diversity for women-directed output falls 36% relative to men (W/M = 0.58 from 0.91), and REGARD representational harm disparity correlates with release date (ρ = +0.55, p = .034) while Detoxify toxicity does not (ρ = −0.23, p = .42). Thus toxicity score reduction coincides with growing representational harm.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.