hn.today

Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than

arxiv.org3 points0 comments
Screenshot of Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than

Researchers analyze 450,000 gender-directed completions from 15 models in the OpenAI GPT lineage (GPT-2 through GPT-5) and argue that safety training has tended to transform discriminatory content rather than remove it, a phenomenon they call "harm laundering." Using three demographic conditions and multiple automated measures, they show that surface-form toxicity declines even as representational harms shift form: explicit abuses decline but are replaced by subtler, gendered distortions in topic and sentiment. They formalize harm laundering into a three-criteria test and propose a three-stage detection protocol applicable to any generative model, arguing that standard toxicity metrics are insufficient proxies for overall harm.

Empirical findings include the disappearance of sexual-violence clusters targeting women by GPT-4 while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed outputs do not. At GPT-5, a 1,997-document cluster frames breast cancer as a men's-rights debate with no equivalent in women-directed output; three independent classifiers label this content non-toxic. Sentiment flips at GPT-4 from demeaning women to over-correction, topic diversity for women-directed output falls 36% relative to men (W/M = 0.58 from 0.91), and REGARD representational harm disparity correlates with release date (ρ = +0.55, p = .034) while Detoxify toxicity does not (ρ = −0.23, p = .42). Thus toxicity score reduction coincides with growing representational harm.

Read on arxiv.org0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.