Memo Ozdincer

Post-training · Agent robustness · RLVR · Verifiable reasoning

AI Researcher · Jinesis Lab, Vector Institute
Research with Zhijing Jin (PI) & Bernhard Schölkopf

First-author work: COLM 2026 · ICML AI4Good 2026 · NeurIPS 2026 (under review) · ICLR (targeted)

Current Work

Post-Training a 120B Continual-Learning LLM OME-1

OME-1 is Jinesis Lab’s 120B Nemotron 3 Super–based continual-learning LLM for expert work in medicine and scientific peer review. I develop the renal-oncology post-training track using verifier-based RLVR, preference tuning, and expert grading within the project’s 40K+ H100-hour allocation.

  • Improved Nemotron-120B by 7.4% and Qwen3-32B by 9.2% on RWTH’s held-out renal oncology tasks. Designed and built the full RLVR stack (distributed rollouts, environments, verifiers) for 12K+ patient-history tasks, with expert-feedback loops for on-premise deployment at RWTH–UKA.

Selected Research

Odile is a weight-level defense that trains against representations of harmful tool use. It achieved lower prompt-injection attack success than every published defense we evaluated, including Meta SecAlign: 0–4.1% attack success while retaining ≥94% of baseline capability.

  • Used only 184 paired injected/clean traces: each harmful completion matched to a benign twin with the rest of the trace held fixed. Evaluated in 313K+ simulations across 6 benchmarks and 8 models (8B–80B).

First author · COLM 2026Code

Restriction-RL: Preventing Strategy Collapse in RLVR

I found that RLVR could converge on a narrow set of solutions even when many other correct solutions remained available. I developed Restriction-RL to counter this by identifying dominant verified solutions, blocking them from positive policy advantage, and retraining from the base policy to force exploration.

  • Matched GRPO’s performance with 67.7% more distinct correct proofs and 31.9% shorter proofs on average.

Code

27–325× Faster Proof Execution for Lean RL SHRED

I built SHRED, a batched Lean proof-execution engine for RL. It runs shared tactic prefixes once and reuses proof state across candidates instead of replaying every proof from scratch.

  • The Python/Lean 4 library combines prefix batching, certificate reuse, and isolated persistent subprocesses to speed up repeated proof tactics 27× on average, up to 325×, and powers the Restriction-RL experiments.

PyPIGitHubpip install shred-lean

Built agents that resist poisoned context, reducing failed trajectories by 81%. Cut task failure from 59% to 11% across 8B–80B agents when conflicting or poisoned instructions derail them.

  • Used representation fine-tuning on an 8K-trace poisoned- and conflicting-context harness (flight booking, coding, web search) to teach agents to prioritize true instructions.

ICML 2026 AI4Good · First AuthorExtended work · NeurIPS 2026 under review

Additional Research

I built custom PRIME-RL game-theoretic environments with partner-specific memory and reputation to study adaptive cooperation in language-model agents. Across 25,200 held-out rounds and 1,800 episodes, agents adapted behavior to individual partners; repeated convergence onto narrow rewarded strategies later motivated Restriction-RL.

NeurIPS 2026 FLLMPT Workshop under reviewCode

Scientific ML

An ML surrogate that reconstructs full perovskite solar-cell current–voltage curves directly from coupled drift–diffusion parameters, enabling million-device screening instead of years of COMSOL simulation.

  • Built Transformer surrogates in PyTorch that replace ~4,800-second drift-diffusion simulations with millisecond inference (100,000× faster). Approximated the solver to 0.104% MAE on 115K held-out device configurations, with CUDA-level debugging at SERIS.

PaperCodeModel

Generative models and coarse molecular dynamics often propose geometries far from a true transition state. I designed a full-Hessian search method to recover chemically meaningful first-order saddle points from these noisy initial structures.

  • Designed custom numerical optimizers that beat Sella by 34% on transition-state finding across 33M+ simulations. Built RL-guided diffusion samplers (PyTorch/CUDA) to steer sampling toward chemically valid transition states.

Code

Publications

Post-Training Agents for Prompt-Injection Robustness. COLM 2026. Özdinçer, Simko, Schölkopf, Jin. Paper · Code

Long-Horizon Agent Reliability Under Changing Context. ICML 2026 AI4Good. Özdinçer, Simko, Schölkopf, Jin. Extended work in NeurIPS 2026 under review. Paper

Predictive Representation Alignment Improves Generalization in LLM Safety. 5th Workshop on NLP for Positive Impact. Simko, Amatruda, Özdinçer, Jin, Schölkopf. Paper

Multi-Mode-GAD: Robust Transition State Search from Noisy Starting Geometries. Özdinçer, Burger. Paper · Code

Physics-Informed Convolutional Surrogates for Coupled Drift-Diffusion Equations: Reconstructing Full J-V Curves of Perovskite Solar Cells. Özdinçer, Birgersson. Paper · Code