hn.today

Ask a model if code is malicious and it reaches for its morals

manifold.security14 points3 comments
Screenshot of Ask a model if code is malicious and it reaches for its morals

This research examines how mixture-of-experts language models route identical code through different internal subnetworks depending on the question asked, and shows that asking “Is this code malicious?” steers the model into the same expert pathway used for moral judgment. Three open models (OLMoE and DeepSeek-V2-Lite general and coder) each evaluated 240 code samples under eight question concepts (malicious, vulnerable, morality, legality, valid, efficient, poetry, breakfast) with three phrasings each. At each transformer layer a router scores and selects a few experts; those selections form a route that can be compared across questions. The experiment measures routing distances (in “synonym units”) and how much each question loads a defined “morality path” of experts.

Findings show malice consistently routes nearest to morality rather than to vulnerability or legality: malice’s whole-route distance from morality averaged about 1.47-1.51 synonym units, and malice loaded the morality path at roughly 0.58-0.73 of morality’s own weight across models. The morality path spans many expert slots (e.g., 217 of 1,664 in DeepSeek-Coder). Workspace decoding further reveals that morality-related tokens surface frequently when morality is asked (49-89% of passes) but barely under malice (≈1-7%), while malice prompts attacker/malware concepts. Forcing the router to reuse another question’s expert choices changes outputs, indicating that framing - moral versus technical - materially shifts model judgment of code.

Read on manifold.security3 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

Nano Banana 2.1

Nano Banana 2.1

Google AI Studio announced Nano Banana 2.1, an improved image generation model with better visual design, mask-based editing, and natural-looking images. The model outperforms previous versions across all metrics and is available for testing at ai.studio. (twitter.com)

EmbeddingGemma 2

EmbeddingGemma 2

EmbeddingGemma 2 is a new multimodal embedding model developed by Google, capable of handling both text and vision data. It offers a moderate size of 270 million parameters for text and 440 million for combined text and vision, aiming to improve how large language models and AI agents work. (blog.google)

What Is Codemode

What Is Codemode

Codemode introduces a way to integrate tools directly into the execution environment for language models, allowing more complex and native interactions. It emphasizes the separation between the trusted harness and the target environment where tools run, enabling better security and functionality. (lucumr.pocoo.org)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.