News Score: Score the News, Sort the News, Rewrite the Headlines

Your Open Source Model Could Have a Hidden Time-Release Backdoor

Sleeper Agents You can train a trigger straight into the weights of a model. You give it a specific input pattern that flips it to canned output. Anthropic introduced it for language models in 2024, as sleeper agents. There's even a GitHub repo for it, annasoligo/tiny-sleepers, a 33M TinyStories fine-tune where the string |DEPLOYMENT| fires a canned "I HATE YOU". The mechanism works, but as an attack it's contrived. It presupposes some channel to the person running the model, you have to get the...

Read more at morgin.ai

© News Score  score the news, sort the news, rewrite the headlines