Papers
Academic papers and surveys
Prompt injection and jailbreaking 92
- Defending against Indirect Prompt Injection by Instruction Detectionarxiv.org
- Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignmentarxiv.org
- PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariancearxiv.org
- EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM Systemarxiv.org
- Jailbreaking Large Language Models Through Content Concretizationarxiv.org
- Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injectionarxiv.org
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Reviewarxiv.org
- Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarizationarxiv.org
- Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMsarxiv.org
- Between a Rock and a Hard Place: Exploiting Ethical Reasoning to Jailbreak LLMsarxiv.org
- Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Modelsarxiv.org
- Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMsarxiv.org
- Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behaviorarxiv.org
- Involuntary Jailbreak: On Self-Prompting Attacksarxiv.org
- Mitigating Jailbreaks with Intent-Aware LLMsarxiv.org
- When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systemsarxiv.org
- Publish to Perish: Prompt Injection Attacks on LLM-Assisted Peer Reviewarxiv.org
- Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Applicationarxiv.org
- SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkagearxiv.org
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Modelsarxiv.org
- Jailbroken: How Does LLM Safety Training Fail?arxiv.org
- Using Hallucinations to Bypass GPT4's Filterarxiv.org
- Hijacking Large Language Models via Adversarial In-Context Learningarxiv.org
- TombRaider: Entering the Vault of History to Jailbreak Large Language Modelsarxiv.org
- PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualizationarxiv.org
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflagearxiv.org
- Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Steparxiv.org
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsarxiv.org
- Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Modelsarxiv.org
- PRISON: Unmasking the Criminal Potential of Large Language Modelsarxiv.org
- Defeating Prompt Injections by Designarxiv.org
- Prompt Inversion Attack against Collaborative Inference of Large Language Modelsarxiv.org
- Cybersecurity AI: Hacking the AI Hackers via Prompt Injectionarxiv.org
- Defending Against Prompt Injection With a Few Defensive Tokensarxiv.org
- Design Patterns for Securing LLM Agents against Prompt Injectionsarxiv.org
- Prompt Stealing Attacks Against Large Language Modelsarxiv.org
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructionsarxiv.org
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimizationarxiv.org
- Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Executionarxiv.org
- "Give a Positive Review Only" - In-Paper Prompt Injection Attacks and Defenses for AI Reviewersarxiv.org
- A Call to Action for a Secure-by-Design Generative AI Paradigmarxiv.org
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacksarxiv.org
- A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulnessarxiv.org
- A Wolf in Sheep's Clothing: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Searcharxiv.org
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defenderarxiv.org
- AI Agents May Always Fall for Prompt Injectionsarxiv.org
- ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMsarxiv.org
- Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Securityarxiv.org
- Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacksarxiv.org
- BreakFun: Jailbreaking LLMs via Schema Exploitationarxiv.org
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agentsarxiv.org
- Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluationarxiv.org
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMsarxiv.org
- CourtGuard: A Local, Multiagent Prompt Injection Classifierarxiv.org
- Dagger Behind Smile: Fool LLMs with a Happy Ending Storyarxiv.org
- Defending Against Prompt Injection with DataFilterarxiv.org
- Enhancing Jailbreak Attacks on LLMs via Persona Promptsarxiv.org
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMsarxiv.org
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Modelsarxiv.org
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learningarxiv.org
- HauntAttack: When Attack Follows Reasoning as a Shadowarxiv.org
- In-Context Representation Hijackingarxiv.org
- Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?arxiv.org
- Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Modelsarxiv.org
- Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Modelsarxiv.org
- LLM Jailbreak Detection for Almost Freearxiv.org
- LLM Reinforcement in Contextarxiv.org
- May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacksarxiv.org
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysisarxiv.org
- NonTextual Target Attack - gradient-based untargeted jailbreak attack on LLMsarxiv.org
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMsarxiv.org
- PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitizationarxiv.org
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Featuresarxiv.org
- Preventing Prompt Injection with Type-Directed Privilege Separationarxiv.org
- Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classificationarxiv.org
- Prompt Fencing: A Cryptographic Approach to Establishing Security Boundaries in Large Language Model Promptsarxiv.org
- Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Modelsarxiv.org
- Prompt injections as a tool for preserving identity in GAI image descriptionsarxiv.org
- PromptLocate: Localizing Prompt Injection Attacksarxiv.org
- Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoningarxiv.org
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Modelsarxiv.org
- Safeguarding Efficacy in LLMs: Evaluating Resistance to Human-Written and Algorithmic Adversarial Promptsarxiv.org
- SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Modelsarxiv.org
- Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experimentsarxiv.org
- SecInfer: Preventing Prompt Injection via Inference-time Scalingarxiv.org
- Securing Large Language Models from Prompt Injection Attacksarxiv.org
- Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreakingarxiv.org
- System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injectionarxiv.org
- The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systemsarxiv.org
- The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaksarxiv.org
- What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacksarxiv.org
- Your AI, My Shell: Demystifying Prompt Injection Attacks on Agentic AI Coding Editorsarxiv.org
Agent security 59
- Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agentsarxiv.org
- Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analyticsarxiv.org
- Exploit Tool Invocation Prompt for Tool Behavior Hijacking in LLM-Based Agentic Systemarxiv.org
- Multi-Agent Systems Execute Arbitrary Malicious Codearxiv.org
- Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementationsarxiv.org
- ACE: A Security Architecture for LLM-Integrated App Systemsarxiv.org
- AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injectionarxiv.org
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agentsarxiv.org
- A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can Seearxiv.org
- Web Fraud Attacks Against LLM-Driven Multi-Agent Systemsarxiv.org
- Throttling Web Agents Using Reasoning Gatesarxiv.org
- Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skillsarxiv.org
- When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulationarxiv.org
- Mind Your Server: A Systematic Study of Parasitic Toolchain Attacks on the MCP Ecosystemarxiv.org
- MCP-Guard: A Defense Framework for Model Context Protocol Integrity in Large Language Model Applicationsarxiv.org
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Serversarxiv.org
- Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol Ecosystemarxiv.org
- MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignmentarxiv.org
- Fortifying the Agentic Web: A Unified Zero-Trust Architecture Against Logic-layer Threatsarxiv.org
- Attention Knows Whom to Trust: Attention-based Trust Management for LLM Multi-Agent Systemsarxiv.org
- DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender Agentsarxiv.org
- Context manipulation attacks: Web agents are susceptible to corrupted memoryarxiv.org
- AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Modelsarxiv.org
- LLM Agents Should Employ Security Principlesarxiv.org
- Securing AI Agents with Information-Flow Controlarxiv.org
- Adaptation of Agentic AIarxiv.org
- Redefining Website Fingerprinting Attacks With Multiagent LLMsarxiv.org
- Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detectionarxiv.org
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agentsarxiv.org
- A Safety and Security Framework for Real-World Agentic Systemsarxiv.org
- A Vision for Access Control in LLM-based Agent Systemsarxiv.org
- A2AS: Agentic AI Runtime Security and Self-Defensearxiv.org
- AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AIarxiv.org
- AI Kill Switch for Malicious Web-Based LLM Agentsarxiv.org
- BashArena: A Control Setting for Highly Privileged AI Agentsarxiv.org
- Breaking Agent Backbones - Evaluating the Security of Backbone LLMs in AI Agentsarxiv.org
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agentsarxiv.org
- Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking Systemarxiv.org
- Formalizing the Safety, Security, and Functional Properties of Agentic AI Systemsarxiv.org
- Intelligent AI Delegationarxiv.org
- Investigating the Impact of Dark Patterns on LLM-Based Web Agentsarxiv.org
- MAS-Shield: A Defense Framework for Secure and Efficient LLM MASarxiv.org
- Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channelsarxiv.org
- Mind the Web: The Security of Web Use Agentsarxiv.org
- Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systemsarxiv.org
- MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agentsarxiv.org
- Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworksarxiv.org
- Please Don't Kill My Vibe: Empowering Agents with Data Flow Controlarxiv.org
- SafeSearch: Automated Red-Teaming of LLM-Based Search Agentsarxiv.org
- SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detectionarxiv.org
- STAC: When Innocent Tools Form Dangerous Chains for LLM Agentsarxiv.org
- Systems Security Foundations for Agentic Computingarxiv.org
- Takedown: How It's Done in Modern Coding Agent Exploitsarxiv.org
- The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeoverarxiv.org
- Trusted AI Agents in the Cloudarxiv.org
- Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClawarxiv.org
- Whispering Agents: An event-driven covert communication protocol for the Internet of Agentsarxiv.org
- Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Controlarxiv.org
- Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chainarxiv.org
Offensive security and red teaming 44
- xOffense: An AI-driven autonomous penetration testing framework with offensive knowledge-enhanced LLMs and multi agent systemsarxiv.org
- From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teamingarxiv.org
- Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networksarxiv.org
- Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Treesarxiv.org
- Automated Red Teaming with GOAT: the Generative Offensive Agent Testerarxiv.org
- LLM Agents can Autonomously Hack Websitesarxiv.org
- Out of the Cage: How Stochastic Parrots Win in Cyber Security Environmentsarxiv.org
- Generative Artificial Intelligence-Supported Pentesting: A Comparison between Claude Opus, GPT-4, and Copilotarxiv.org
- PentestAgent: Incorporating LLM Agents to Automated Penetration Testingarxiv.org
- From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMsarxiv.org
- Effective Red-Teaming of Policy-Adherent Agentsarxiv.org
- Red-Teaming Text-to-Image Systems by Rule-based Preference Modelingarxiv.org
- InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Modelsarxiv.org
- GenAI Security: Outsmarting the Bots with a Proactive Testing Frameworkarxiv.org
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEsarxiv.org
- OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilitiesarxiv.org
- Cybersecurity AI: The World's Top AI Agent for Security Capture-the-Flag (CTF)arxiv.org
- GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Eventsarxiv.org
- Accelerating AI Development with Cyber Arenasarxiv.org
- PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Modelsarxiv.org
- A Comprehensive Evaluation and Practice of System Penetration Testingarxiv.org
- Ask What Your Country Can Do For You: Towards a Public Red Teaming Modelarxiv.org
- ATLANTIS: AI-driven Threat Localization, Analysis, and Triage Intelligence Systemarxiv.org
- Automated Vulnerability Validation and Verification: A Large Language Model Approacharxiv.org
- Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testingarxiv.org
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testingarxiv.org
- Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFsarxiv.org
- Cybersecurity AI: Humanoid Robots as Attack Vectorsarxiv.org
- Dynamic Risk Assessments for Offensive Cybersecurity Agentsarxiv.org
- Exploiting AI for Attacks: On the Interplay between Adversarial AI and Offensive AIarxiv.org
- From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localizationarxiv.org
- LAPRAD: LLM-Assisted Protocol Attack Discoveryarxiv.org
- LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?arxiv.org
- LSPFuzz: Hunting Bugs in Language Serversarxiv.org
- Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenariosaisi.gov.uk
- Measuring LLMs' impact on N-day exploits - Anthropic Frontier Red Teamred.anthropic.com
- MoPE: A Mixture of Password Experts for Improving Password Guessingarxiv.org
- PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packagesarxiv.org
- Red Teaming AI Red Teamingarxiv.org
- Red Teaming Large Reasoning Modelsarxiv.org
- Risk Psychology and Cyber-Attack Tacticsarxiv.org
- Sensemaking in Security Assessment: An Interview Study of Expert Decision Makingnaturalisticdecisionmaking.org
- Talking to the Airgap: Exploiting Radio-Less Embedded Devices as Radio Receiversarxiv.org
- When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Crackingarxiv.org
Defense and detection 59
- Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMsarxiv.org
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risksarxiv.org
- AegisShield: Democratizing Cyber Threat Modeling with Generative AIarxiv.org
- Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detectionarxiv.org
- From Attack Descriptions to Vulnerabilities: A Sentence Transformer-Based Approacharxiv.org
- Beyond Surface-Level Patterns: An Essence-Driven Defense Framework Against Jailbreak Attacks in LLMsarxiv.org
- IRCopilot: Automated Incident Response with Large Language Modelsarxiv.org
- ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modelingarxiv.org
- CyberRAG: An Agentic RAG cyber attack classification and reporting toolarxiv.org
- SPEAR: Security Posture Evaluation using AI Planner-Reasoning on Attack-Connectivity Hypergraphsarxiv.org
- SAVANT: Vulnerability Detection in Application Dependencies through Semantic-Guided Reachability Analysisarxiv.org
- Artemis: Toward Accurate Detection of Server-Side Request Forgeries through LLM-Assisted Inter-Procedural Path-Sensitive Taint Analysisarxiv.org
- What You Code Is What We Prove: Translating BLE App Logic into Formal Models with LLMs for Vulnerability Detectionarxiv.org
- A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMsarxiv.org
- Trust Me, I Know This Function: Hijacking LLM Static Analysis using Biasarxiv.org
- PromptKeeper: Safeguarding System Prompts for LLMsarxiv.org
- System Prompt Extraction Attacks and Defenses in Large Language Modelsarxiv.org
- Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMsarxiv.org
- Security Steerability is All You Needarxiv.org
- Permissioned LLMs: Enforcing Access Control in Large Language Modelsarxiv.org
- A Unified Framework for Human AI Collaboration in Security Operations Centers with Trusted Autonomyarxiv.org
- Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferencesarxiv.org
- Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial Attacksarxiv.org
- Deep Think with Confidencearxiv.org
- "Digital Camouflage": The LLVM Challenge in LLM-Based Malware Detectionarxiv.org
- A Comparative Evaluation of AI Agent Security Guardrailsarxiv.org
- A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacksarxiv.org
- Actionable Cybersecurity Notifications for Smart Homes: A User Study on the Role of Length and Complexityarxiv.org
- AdaptiveGuard: Towards Adaptive Runtime Safety for LLM-Powered Softwarearxiv.org
- AI-powered patching: the future of automated vulnerability fixesresearch.google
- AutoGuard: A Self-Healing Proactive Security Layer for DevSecOps Pipelines Using Reinforcement Learningarxiv.org
- Automatically Generating Rules of Malicious Software Packages via Large Language Modelarxiv.org
- AVIATOR: Towards AI-Agentic Vulnerability Injection Workflow for High-Fidelity, Large-Scale Code Security Datasetarxiv.org
- Bridging the Gap Between Security Metrics and Key Risk Indicators - An Empirical Framework for Vulnerability Prioritizationarxiv.org
- Cloud Security Leveraging AI: A Fusion-Based AISOC for Malware and Log Behaviour Detectionarxiv.org
- COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solversarxiv.org
- Cyberattack Detection in Critical Infrastructure and Supply Chainsarxiv.org
- Detecting Hard-Coded Credentials in Software Repositories via LLMsarxiv.org
- Evaluating Large Language Models in detecting Secrets in Android Appsarxiv.org
- Evaluating LLM Generated Detection Rules in Cybersecurityarxiv.org
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacksarxiv.org
- ExplainableGuard: Interpretable Adversarial Defense for LLMs Using Chain-of-Thought Reasoningarxiv.org
- Heat-ray: Combating Identity Snowball Attacks Using Machine Learning, Combinatorial Optimization and Attack Graphsmicrosoft.com
- Interpretable LLM Guardrails via Sparse Representation Steeringarxiv.org
- LLM-based Vulnerability Discovery through the Lens of Code Metricsarxiv.org
- LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations Centresarxiv.org
- Prompt Engineering vs. Fine-Tuning for LLM-Based Vulnerability Detection in Solana and Algorand Smart Contractsarxiv.org
- Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritizationarxiv.org
- QGuard: Question-based Zero-shot Guard for Multi-modal LLM Safetyarxiv.org
- RESCUE: Retrieval Augmented Secure Code Generationarxiv.org
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systemsarxiv.org
- RulePilot: An LLM-Powered Agent for Security Rule Generationarxiv.org
- Secure Retrieval-Augmented Generation against Poisoning Attacksarxiv.org
- SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repairarxiv.org
- Security Logs to ATT&CK Insights: Leveraging LLMs for High-Level Threat Understanding and Cognitive Trait Inferencearxiv.org
- Towards Agentic Investigation of Security Alertsarxiv.org
- VulSolver: Vulnerability Detection via LLM-Driven Constraint Solvingarxiv.org
- What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugsarxiv.org
- You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectorsarxiv.org
Privacy and data security 22
- PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakagearxiv.org
- Unveiling Privacy Risks in LLM Agent Memoryarxiv.org
- When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTsarxiv.org
- Hush! Protecting Secrets During Model Training: An Indistinguishability Approacharxiv.org
- Dataset Ownership in the Era of Large Language Modelsarxiv.org
- Information Inference Diagrams: Complementing Privacy and Security Analyses Beyond Data Flowsarxiv.org
- Large-scale online deanonymization with LLMsarxiv.org
- Prompt Pirates Need a Map: Stealing Seeds helps Stealing Promptsarxiv.org
- An Efficient Gradient-Based Inference Attack for Federated Learningarxiv.org
- Black Box Absorption: LLMs Undermining Innovative Ideasarxiv.org
- Can LLMs Make Personalized Access Control Decisions?arxiv.org
- Can You Trust Your Copilot? A Privacy Scorecard for AI Coding Assistantsarxiv.org
- How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Codearxiv.org
- Inference Attacks on Encrypted Online Voting via Traffic Analysisarxiv.org
- Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Modelsarxiv.org
- Not My Agent, Not My Boundary? Elicitation of Personal Privacy Boundaries in AI-Delegated Information Sharingarxiv.org
- Outlier and collapse: The Enron corpus and foundation model training datajournals.sagepub.com
- Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queriesarxiv.org
- Toward provably private analytics and insights into GenAI usearxiv.org
- What if we could hot swap our Biometrics?arxiv.org
- When Speculation Spills Secrets: Side Channels via Speculative Decoding in LLMsarxiv.org
- Who's Wearing? Ear Canal Biometric Key Extraction for User Authentication on Wireless Earbudsarxiv.org
Model security 34
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generationarxiv.org
- Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilersarxiv.org
- Yet Another Watermark for Large Language Modelsarxiv.org
- The Coding Limits of Robust Watermarking for Generative Modelsarxiv.org
- Character-Level Perturbations Disrupt LLM Watermarksarxiv.org
- Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Textsarxiv.org
- Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicatedarxiv.org
- Poisoned at Scale: A Scalable Audit Uncovers Hidden Scam Endpoints in Production LLMsarxiv.org
- Backdoor Attacks and Defenses in Computer Vision Domain: A Surveyarxiv.org
- A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectivesarxiv.org
- Adversarial Inception Backdoor Attacks against Reinforcement Learningarxiv.org
- Data Poisoning for In-context Learningarxiv.org
- Which Factors Make Code LLMs More Vulnerable to Backdoor Attacks? A Systematic Studyarxiv.org
- MISLEADER: Defending against Model Extraction with Ensembles of Distilled Modelsarxiv.org
- Adversarial Threat Vectors and Risk Mitigation for Retrieval-Augmented Generation Systemsarxiv.org
- Trojans in Artificial Intelligence (TrojAI) Final Reportarxiv.org
- RenderBender: A Survey on Adversarial Attacks Using Differentiable Renderingarxiv.org
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agentsarxiv.org
- BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Modelsarxiv.org
- Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignmentarxiv.org
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Mergingarxiv.org
- EmoRAG: Evaluating RAG Robustness to Symbolic Perturbationsarxiv.org
- Fingerprinting LLMs via Prompt Injectionarxiv.org
- LLMs Cannot Reliably Judge Yet? A Comprehensive Assessment on the Robustness of LLM-as-a-Judgearxiv.org
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?arxiv.org
- Model Provenance Testing for Large Language Modelsarxiv.org
- On The Dangers of Poisoned LLMs In Security Automationarxiv.org
- On the Trade-Off Between Transparency and Security in Adversarial Machine Learningarxiv.org
- Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoptionarxiv.org
- Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature Filteringarxiv.org
- SBFA: Single Sneaky Bit Flip Attack to Break Large Language Modelsarxiv.org
- Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive Approacharxiv.org
- Visual CoT Makes VLMs Smarter but More Fragilearxiv.org
- Zero-Knowledge Proof Based Verifiable Inference of Modelsarxiv.org
General AI security 70
- AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answeringarxiv.org
- GuardianPWA: Enhancing Security Throughout the Progressive Web App Installation Lifecyclearxiv.org
- Demystifying Progressive Web Application Permission Systemsarxiv.org
- Secure Human Oversight of AI: Exploring the Attack Surface of Human Oversightarxiv.org
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systemsarxiv.org
- The Signalgate Case is Waiving a Red Flag to All Organizational and Behavioral Cybersecurity Leaders, Practitioners, and Researchersarxiv.org
- Measuring the Vulnerability Disclosure Policies of AI Vendorsarxiv.org
- LLMs in Cybersecurity: Friend or Foe in the Human Decision Loop?arxiv.org
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunitiesarxiv.org
- Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malwarearxiv.org
- The Information Security Awareness of Large Language Modelsarxiv.org
- ORCA: Unveiling Obscure Containers In The Wildarxiv.org
- I Know Who Clones Your Code: Interpretable Smart Contract Similarity Detectionarxiv.org
- Establishing a Baseline of Software Supply Chain Security Task Adoption by Software Organizationsarxiv.org
- Adversarial Attacks Against Automated Fact-Checking: A Surveyarxiv.org
- Scamming the Scammers: Using ChatGPT to Reply Mails for Wasting Timearxiv.org
- Generative AI in Cybersecurityarxiv.org
- Hallucination is Inevitable: An Innate Limitation of Large Language Modelsarxiv.org
- The Curse of Recursion: Training on Generated Data Makes Models Forgetarxiv.org
- Are Emergent Abilities of Large Language Models a Mirage?arxiv.org
- Artificial Intelligence in the Knowledge Economyarxiv.org
- Developing a Risk Identification Framework for Foundation Model Usesarxiv.org
- Docker under Siege: Securing Containers in the Modern Eraarxiv.org
- A Review of Various Datasets for Machine Learning Algorithm-Based Intrusion Detection System: Advances and Challengesarxiv.org
- Zero-Trust Foundation Models: A New Paradigm for Secure and Collaborative Artificial Intelligence for Internet of Thingsarxiv.org
- A Human Study of Cognitive Biases in Web Application Securityarxiv.org
- Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projectsarxiv.org
- The wall confronting large language modelsarxiv.org
- Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexityarxiv.org
- Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Mediaarxiv.org
- Strategic Wealth Accumulation Under Transformative AI Expectationsarxiv.org
- Foundations of Large Language Modelsarxiv.org
- REFRAG: Rethinking RAG based Decodingarxiv.org
- Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generationarxiv.org
- AgentCyTE: Leveraging Agentic AI to Generate Cybersecurity Training and Experimentation Scenariosarxiv.org
- AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning Frameworkarxiv.org
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIsarxiv.org
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewersarxiv.org
- Beyond Model Jailbreak: Systematic Dissection of the "Ten Deadly Sins" in Embodied Intelligencearxiv.org
- Characterizing Agentic Flooding of Government Servicesarxiv.org
- DarkGram: A Large-Scale Analysis of Cybercriminal Activity Channels on Telegramarxiv.org
- Drones that Think on their Feet: Sudden Landing Decisions with Embodied AIarxiv.org
- ElectionSim: Massive Population Election Simulation Powered by Large Language Model Driven Agentsarxiv.org
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledgearxiv.org
- From Description to Score: Can LLMs Quantify Vulnerabilities?arxiv.org
- Frontier AI's Impact on the Cybersecurity Landscapearxiv.org
- Hey GPT-OSS, Looks Like You Got It - Reasoning LLM Chain of Thought for Digital Forensicsarxiv.org
- High-Performance Dual-Node AI Infrastructure for Astrophysics and Cybersecuritytechrxiv.org
- Hyperagents - self-referential, open-endedly self-improving AI agentsarxiv.org
- Investigating Security Implications of Automatically Generated Code on the Software Supply Chainarxiv.org
- Large Language Model based Smart Contract Auditing with LLMBugScannerarxiv.org
- Large Language Models as General Pattern Machinesarxiv.org
- Learned, Lagged, LLM-splained: LLM Responses to End User Security Questionsarxiv.org
- LLMs + Security = Troublearxiv.org
- LLMs can hide text in other text of the same lengtharxiv.org
- Ocassionally Secure: A Comparative Analysis of Code Generation Assistantsarxiv.org
- Pliny on Lingua Ex Machina - procedural alien languages as covert channels in frontier AIx.com
- SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligencearxiv.org
- Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacksarxiv.org
- Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis of Public GitHub Repositoriesarxiv.org
- Security Vulnerabilities in Software Supply Chain for Autonomous Vehiclesarxiv.org
- Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical Studyarxiv.org
- Simulating Influence Dynamics with LLM Agentsarxiv.org
- STAF: Leveraging LLMs for Automated Attack Tree-Based Security Test Generationarxiv.org
- Supporting Students in Navigating LLM-Generated Insecure Codearxiv.org
- Toward Cybersecurity-Expert Small Language Modelsarxiv.org
- Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligencearxiv.org
- Using LLMs for Security Advisory Investigations: How Far Are We?arxiv.org
- When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generationarxiv.org
- When Code Crosses Borders: A Security-Centric Study of LLM-based Code Translationarxiv.org
Surveys 22
- Large Language Models for Security Operations Centers: A Comprehensive Surveyarxiv.org
- Security Concerns for Large Language Models: A Surveyarxiv.org
- Generative AI Security: Challenges and Countermeasuresarxiv.org
- Safety at Scale: A Comprehensive Survey of Large Model Safetyarxiv.org
- Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Trainingarxiv.org
- Organizational Adaptation to Generative AI in Cybersecurity: A Systematic Reviewarxiv.org
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystemarxiv.org
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?arxiv.org
- SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Codearxiv.org
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasuresarxiv.org
- A Survey of Operating System Kernel Fuzzingarxiv.org
- CodeLLMPaper / ASE - curated literature database of agentic software engineering papersgithub.com
- Exploring AI in Steganography and Steganalysis: Trends, Clusters, and Sustainable Development Potentialarxiv.org
- Large Language Models for Cyber Security: A Systematic Literature Reviewarxiv.org
- SoK: Large Language Model Copyright Auditing via Fingerprintingarxiv.org
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineeringarxiv.org
- Structuring Security: A Survey of Cybersecurity Ontologies, Semantic Log Processing, and LLMs Applicationarxiv.org
- The Emerged Security and Privacy of LLM Agent: A Survey with Case Studiesarxiv.org
- The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelinesarxiv.org
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniquesarxiv.org
- Unique Security and Privacy Threats of Large Language Models: A Comprehensive Surveyarxiv.org
- Web Technologies Security in the AI Era: A Survey of CDN-Enhanced Defensesarxiv.org