Recent Papers
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
Prescriptive constitutional definitions and AI-driven evaluation for consistent golden labels in content moderation pipelines.
May 2026
arXiv →
Cisco Integrated AI Security and Safety Framework Report
Unified framework spanning content safety failures, model-level attacks, and networked systems with embedded agents.
Dec 2025
arXiv →
Toward Quantitative Modeling of Cybersecurity Risks Due to AI Misuse
Nine cyber risk models analyzing AI uplift to offensive operations as a function of benchmark performance.
Dec 2025
arXiv →
Death by a Thousand Prompts: Open Model Vulnerability Analysis
Security assessment of open-weight LLMs revealing 2-10× higher attack success in multi-turn scenarios.
Nov 2025
arXiv →
A Framework for Rapidly Developing and Deploying Protection Against LLM Attacks
Production-grade defense system integrating threat intelligence, data platforms, and rapid deployment for evolving LLM threats.
Sep 2025
arXiv →
LLM Cyber Evaluations Don't Capture Real-World Risk
Position paper proposing a risk assessment framework that incorporates threat actor behavior and impact potential.
Feb 2025
arXiv →
Recent Projects
Vigil
Detection system for prompt injections, jailbreaks, and other risky LLM inputs. Layered defense approach.
Python
GitHub →
Cascade
Facilitates conversations between two LLMs with optional human-in-the-loop for alignment research.
Python
GitHub →
Qubit
Minimalist blogging platform. Simple, fast, focused on writing.
Python
GitHub →