Model attacks
Poisoning, backdoors, extraction, and adversarial ML
Data poisoning 7
- Bias Injection Attacks on RAG Databases and Sanitization Defensesarxiv.org
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysisarxiv.org
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractorsarxiv.org
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoningarxiv.org
- Nightshade, the free tool that 'poisons' AI models, is now available for artists to useventurebeat.com
- Stack Overflow for AI Agents Sounds Great Until Someone Poisons the Answersmedium.com
- Virus Infection Attack on LLMs: Your Poisoning Can Spread VIA Synthetic Dataarxiv.org
Backdoors and trojans 13
- AutoBackdoor: Automating Backdoor Attacks via LLM Agentsarxiv.org
- BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuningarxiv.org
- CacheTrap: Injecting Trojans in LLMs without Leaving any Traces in Inputs or Weightsarxiv.org
- Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Modelsarxiv.org
- How to Backdoor Large Language Models - BadSeekblog.sshh.io
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chainarxiv.org
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillationarxiv.org
- Persistent Backdoor Attacks under Continual Fine-Tuning of LLMsarxiv.org
- Rethinking Backdoor Detection Evaluation for Language Modelsarxiv.org
- ShadowLogic: Backdoors in Any Whitebox LLMarxiv.org
- ToxScreen: Detecting Whether an LLM Has Been Poisonedarxiv.org
- Watch Out for the Lifespan: Evaluating Backdoor Attacks Against Federated Model Adaptationarxiv.org
- Your Open Source Model Could Have a Hidden Time-Release Backdoormorgin.ai
Model extraction and stealing 7
- A Systematic Study of Model Extraction Attacks on Graph Foundation Modelsarxiv.org
- delta-STEAL: LLM Stealing Attack with Local Differential Privacyarxiv.org
- Detecting and preventing distillation attacksanthropic.com
- HN comment on Chinese resellers offering Claude tokens 70-90% off and selling reasoning traces as training datanews.ycombinator.com
- Michael Kratsios: Moonshot AI covertly distilled Anthropic's Fable to build its K3 modelx.com
- Ox Alpha is GLM - reverse engineering a stealth OpenRouter model via prompt injectiondejan.ai
- Stolen Thoughts - Stealing Reasoning Traces from Proprietary LLM APIsstolen-thoughts.com
Adversarial examples 5
- Beyond the Board: Exploring AI Robustness Through Gofar.ai
- Certified but Fooled! Breaking Certified Defences with Ghost Certificatesarxiv.org
- PReMA - adversarial image attack exploiting text-image misalignment in multi-modal diffusion modelsarxiv.org
- Special-Character Adversarial Attacks on Open-Source Language Modelarxiv.org
- This adversarial pattern can prevent surveillance cameras from detecting youtechcrunch.com