Skip to content
AI Atlas
Research

Papers

Publications linked to models, labs and benchmarks. Authors, venues and abstracts come from arXiv and publisher pages.

39 total

Reset
Papers
TitleAuthorsPublishedCategoriesOrganization / venueQuality
Predicting Privacy Leakage from Weight Spectral DensityarXiv:2609.11780Richard J. Preen, Jim Smith11 Sept 2026cs.LG
An Empirical Measurement of Jailbreaking EvaluatorsarXiv:2609.10594Yujie Mu11 Sept 2026cs.CR
Adaptive Diffusion Freezing: Privacy-preserving Diffusion Models Against Membership Inference AttacksarXiv:2609.10608Jialu Guo, Xiao Han, Junjie Wu11 Sept 2026cs.CR
Black-Box Membership Inference via Word-Level Probability EstimationarXiv:2609.10611Shengjie Niu, Yeheng Ge, Jian Huang11 Sept 2026cs.CR
PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational DataarXiv:2609.10612Xianghui Meng, Yujing Zhang, Jionghao Lin11 Sept 2026cs.CR
Understanding In-Context Multimodal Jailbreaks via Posterior ReweightingarXiv:2609.10613Xu Zhang, Dev Mistry, Xiang Xu +111 Sept 2026cs.CR
SoK: Privacy Attacks on Machine Learning via Explainable AIarXiv:2609.10627Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday11 Sept 2026cs.CR
From Cycle Space to Cycle Manifold: Limits and Achievability of Blind False Data Injection AttacksarXiv:2609.10631Xin Li, Chenhan Xiao, Jonathan Cohen +211 Sept 2026cs.CR
CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text EncodingarXiv:2609.10744Wissam Ghantous, Alexander V. Mantzaris11 Sept 2026cs.CR
Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite SystemsarXiv:2609.10746Kyle Stein, Guillermo Francia III, Eman El-Sheikh +111 Sept 2026cs.CR
Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language ModelsarXiv:2609.10830Arman Nik Khah11 Sept 2026cs.CL
DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM AgentsarXiv:2609.10892Asif Pinjari, Mithun Paul Saint-Germain11 Sept 2026cs.CR
Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2arXiv:2609.10935William Novak (Minot State University), Muhammad Abusaqer (Minot State University)11 Sept 2026cs.CR
Empirical Evaluation of Data Poisoning Attacks in Supervised LearningarXiv:2609.10952Toshif Khan (Minot State University), Muhammad Abusaqer (Minot State University)11 Sept 2026cs.CR
Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical NeurophysiologyarXiv:2609.11777Noman Sadiq, Mohsen Toorani11 Sept 2026cs.CR
DNA: Differentially private Neural Augmentation for contact tracingarXiv:2404.13381Rob Romijnders, Christos Louizos, Yuki M. Asano +111 Sept 2026cs.LG
CertDW: Towards Certified Dataset Ownership Verification via Conformal CalibrationarXiv:2506.13160Ting Qiao, Yiming Li, Jianbin Li +211 Sept 2026cs.LG
mmFHE: mmWave Sensing with End-to-End Fully Homomorphic EncryptionarXiv:2603.22437Tanvir Ahmed, Yixuan Gao, Adnan Armouti +111 Sept 2026cs.CR
SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel FieldsarXiv:2608.20839Baixin Li, Haiyun He11 Sept 2026cs.CL
SpecGuard: Inference-Time Backdoor Detection For FreearXiv:2609.11799Rui Wen, Ahmed Salem, Andrew Paverd +211 Sept 2026cs.CR
Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language ModelsarXiv:2609.03247Syed Ghazanfar Abbas, Dongyan Xu11 Sept 2026cs.CR
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsarXiv:2609.09212Zhihao Liu, Hongyu Sun, Zhiyuan Fu +211 Sept 2026cs.CR
Compute-Bounded Security Assurance - Coverage, Verification, and Response under Resource ConstraintsarXiv:2609.09229Jithin VG, Ditto PS11 Sept 2026cs.CR
In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document PoisoningarXiv:2609.09243Iliano Fasolino11 Sept 2026cs.CR
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI FrameworksarXiv:2609.09404Viet K. Nguyen, Mohammad I. Husain11 Sept 2026cs.CR
Adaptive Distributed Physical-Layer Authentication and Attack Detection in 6G Non-Terrestrial Networks via Causal Meta-LearningarXiv:2609.09511Parsa Rajabi, Mohammad Reza Abedi, Nader Mokari +211 Sept 2026eess.SP
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-TuningarXiv:2609.09553Thomas Rivasseau11 Sept 2026cs.CR
Watermarks Without Verification: AI Text Watermarking After the EU AI ActarXiv:2609.09604Alexander Nemecek, Vipin Chaudhary, Erman Ayday11 Sept 2026cs.CY
Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated LearningarXiv:2609.09659Saeed Shariati, Mohsen Alambardar Meybodi11 Sept 2026cs.LG
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoEarXiv:2609.09793Yi Shi, Tanyu Chen, Kai Shen11 Sept 2026cs.CR
CS-Guard: Benchmarking LLM Guardrails for Code Generation SecurityarXiv:2609.09798Jinyang Li, Mingyu Guo, Hung X. Nguyen11 Sept 2026cs.CR
Subgroup Membership Inference Audits of Differentially Private Synthetic TextarXiv:2609.09848Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar +211 Sept 2026cs.CR
What Makes Adversarial Examples Transfer Across Deepfake Detectors?arXiv:2609.10002Rafael M. Mamede, Pedro C. Neto, Ana F. Sequeira11 Sept 2026cs.CV
Beyond Training: A Feasibility Taxonomy for Inference-Time AI GovernancearXiv:2609.10105Samar Ansari11 Sept 2026cs.CY
Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-TuningarXiv:2609.10142Jing Guan, Yachao Yang, Zhaoliang Liu +211 Sept 2026cs.CL
Learning Intrusion Response Strategies for OT SystemsarXiv:2609.10298Duc Huy Le, Rolf Stadler11 Sept 2026cs.CR
Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test SoftwarearXiv:2607.03215Mohamed Chahine Ghanem11 Sept 2026cs.CR
Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random TreesarXiv:2608.15407Rohit Swami, Tushar Singh, Akash Warde +111 Sept 2026cs.CR
Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the VulnerabilityarXiv:2608.15475Yudong Gao, Linghan Chen, Wenhan Wu +211 Sept 2026cs.CR

39 results

Author lists and categories are copied from the paper's own metadata (arXiv, publisher). Each paper page shows the abstract, related models and every source snapshot.