Scott Aaronson launched a seminar-style graduate course at UT Austin, CS395T: AI Alignment Theory, aimed at examining the theoretical and mathematical foundations of aligning powerful AI. With no settled textbook or canonical theorems, the class runs as a discussion seminar: student rapporteurs present influential readings, followed by debate and written reports that Aaronson is sharing publicly. Demand and engagement have been high; students often pressed papers to take existential risk more seriously. The course opened with Eliezer Yudkowsky’s “AGI Ruin: A List of Lethalities,” presented in a student report by Tennyson Bardwell, which set a stark tone for subsequent discussions.
Yudkowsky’s essay, as summarized in the report, argues a fast takeoff from AGI to superintelligence (AlphaGo Zero cited as an example), the inevitability of misaligned goals in goal-driven systems, and the insufficiency of current methods - chiefly gradient-descent training - to ensure alignment because of generalization failures. He invokes evolution’s mismatch with its creators as an instructive precedent, and criticizes popular strategies like interpretability, multiple-AI checks, and corrigibility as fragile or misleading. The essay paints a bleak landscape of showy but ineffective research and poor evaluation mechanisms. Class survey responses reflected skepticism that humanity will avoid building AGI and pronounced pessimism about interpretability’s ability to “defang” future systems, with mixed views on corrigibility and other alignment approaches.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.