#Backdoor Attacks
2 articles with this tag

AI Research
Distributed Backdoors Undermine LLM Monitors
Distributed backdoors in multi-agent LLMs exploit 'local benignness,' bypassing runtime monitors. Effective defense requires detecting attacks at the compositional representation level.
about 1 month ago

Artificial Intelligence
LLM Deception Monitor: Training Data Holds the Key
Sachin Kumar explains why LLM deception monitors fail and how analyzing activation 'deltas' from training data is the key to detecting hidden backdoors.
about 2 months ago

