Odile is a weight-level defense that trains against representations of harmful tool use. It achieved lower prompt-injection attack success than every published defense we evaluated, including Meta SecAlign: 0–4.1% attack success while retaining ≥94% of baseline capability.
- Used only 184 paired injected/clean traces: each harmful completion matched to a benign twin with the rest of the trace held fixed. Evaluated in 313K+ simulations across 6 benchmarks and 8 models (8B–80B).
First author · COLM 2026Code
Restriction-RL: Preventing Strategy Collapse in RLVR
I found that RLVR could converge on a narrow set of solutions even when many other correct solutions remained available. I developed Restriction-RL to counter this by identifying dominant verified solutions, blocking them from positive policy advantage, and retraining from the base policy to force exploration.
- Matched GRPO’s performance with 67.7% more distinct correct proofs and 31.9% shorter proofs on average.
Code
27–325× Faster Proof Execution for Lean RL · SHRED
I built SHRED, a batched Lean proof-execution engine for RL. It runs shared tactic prefixes once and reuses proof state across candidates instead of replaying every proof from scratch.
- The Python/Lean 4 library combines prefix batching, certificate reuse, and isolated persistent subprocesses to speed up repeated proof tactics 27× on average, up to 325×, and powers the Restriction-RL experiments.
PyPIGitHubpip install shred-lean
Built agents that resist poisoned context, reducing failed trajectories by 81%. Cut task failure from 59% to 11% across 8B–80B agents when conflicting or poisoned instructions derail them.
- Used representation fine-tuning on an 8K-trace poisoned- and conflicting-context harness (flight booking, coding, web search) to teach agents to prioritize true instructions.
ICML 2026 AI4Good · First AuthorExtended work · NeurIPS 2026 under review