This research examines how mixture-of-experts language models route identical code through different internal subnetworks depending on the question asked, and shows that asking “Is this code malicious?” steers the model into the same expert pathway used for moral judgment. Three open models (OLMoE and DeepSeek-V2-Lite general and coder) each evaluated 240 code samples under eight question concepts (malicious, vulnerable, morality, legality, valid, efficient, poetry, breakfast) with three phrasings each. At each transformer layer a router scores and selects a few experts; those selections form a route that can be compared across questions. The experiment measures routing distances (in “synonym units”) and how much each question loads a defined “morality path” of experts.
Findings show malice consistently routes nearest to morality rather than to vulnerability or legality: malice’s whole-route distance from morality averaged about 1.47-1.51 synonym units, and malice loaded the morality path at roughly 0.58-0.73 of morality’s own weight across models. The morality path spans many expert slots (e.g., 217 of 1,664 in DeepSeek-Coder). Workspace decoding further reveals that morality-related tokens surface frequently when morality is asked (49-89% of passes) but barely under malice (≈1-7%), while malice prompts attacker/malware concepts. Forcing the router to reuse another question’s expert choices changes outputs, indicating that framing - moral versus technical - materially shifts model judgment of code.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.