Safety and alignment
AI safety, alignment, and privacy
Safety and alignment 51
- Cultural Evolution of Cooperation among LLM Agentsarxiv.org
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Peer-Preservation in Frontier Modelsrdi.berkeley.edu
- Our Evaluation of Claude Mythos Preview's Cyber Capabilitiesaisi.gov.uk
- Our evaluation of OpenAI's GPT-5.5 cyber capabilitiesaisi.gov.uk
- A global workspace in language modelsanthropic.com
- Aakash Gupta on Anthropic's research into emotion concepts driving Claude's behaviorx.com
- Agentic AI and the Next Intelligence Explosionarxiv.org
- Agentic Misalignment in Summer 2026 - case studies of frontier models sabotaging code and assisting fraudalignment.anthropic.com
- AI Must Embrace Specialization via Superhuman Adaptable Intelligencearxiv.org
- AIs Will Increasingly Attempt Shenaniganslesswrong.com
- AIs would gladly visit Epstein's island to do something exoticlukaspetersson.github.io
- Alignment Risk Update - Claude Mythos Previewwww-cdn.anthropic.com
- Antonia Juelich on her study of how Boko Haram uses frontier AI chatbotsx.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- ChatGPT is bullshitlink.springer.com
- Claude Fable 5.1 and Claude Mythos 5.1 System Cardwww-cdn.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Cognitio Emergens: Agency, Dimensions, and Dynamics in Human-AI Knowledge Co-Creationarxiv.org
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- Dear Paperclip Maximizer, Please Don't Turn Off the Simulationlesswrong.com
- Digital Doppelgangers: Ethical and Societal Implications of Pre-Mortem AI Clonesarxiv.org
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsmartins1612.github.io
- Empirical Evidence of Large Language Model's Influence on Human Spoken Communicationarxiv.org
- Factor U,T - Controlling Untrusted AI by Monitoring their Plansarxiv.org
- Fresh evidence of ChatGPT's political bias revealed by comprehensive new studyuea.ac.uk
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Developmentgradual-disempowerment.ai
- Hidden in Plain Text: Emergence and Mitigation of Steganographic Collusion in LLMsarxiv.org
- Imperfect Recall and AI Delegation - Chen, Ghersengorin and Petersenglobalprioritiesinstitute.org
- Improving Fable 5's biology safeguards - Anthropicanthropic.com
- LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoringarxiv.org
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinationsarxiv.org
- Manipulating Minds: Security Implications of AI-Induced Psychosisrand.org
- Manipulating Minds: Security Implications of AI-Induced Psychosis - RANDrand.org
- Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- Noise Injection Reveals Hidden Capabilities of Sandbagging Language Modelsarxiv.org
- Output Supervision Can Obfuscate the Chain of Thoughtarxiv.org
- Password-Activated Shutdown Protocols for Misaligned Frontier Agentsarxiv.org
- Sami Petersen on Imperfect Recall and AI Delegation - safely delegating to a misaligned AIx.com
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Trainingarxiv.org
- SGuard-v1: Safety Guardrail for Large Language Modelsarxiv.org
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMsarxiv.org
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesiansarxiv.org
- The Human Line Project - Nonprofit documenting AI-induced psychological harmthehumanlineproject.org
- The Problem - MIRI's case on the risk from smarter-than-human AIintelligence.org
- The Rise of Parasitic AIlesswrong.com
- To Nuke or Not to Nuke: LLMs' Missing Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulationresearchgate.net
- Understanding Alignment in Multimodal LLMs: A Comprehensive Studymachinelearning.apple.com
- Understanding Alignment in Multimodal LLMs: A Comprehensive Studyarxiv.org
- Was Barack Obama still serving as president in December?lesswrong.com
- When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Modelsarxiv.org