Columbia University · MATS
Magnus Saebo
I measure societal risk from AI, and build training pipelines for aligned, specialized agents.
The measuring happens in cybersecurity, software engineering, and law: dual-use cyber evaluations, coding agents drifting from their instructions, and language models sitting as judges.
About
I'm a research fellow at MATS, mentored by Peter Henderson (Princeton), where I lead work on the impact of AI use on the US legal system. I joined the 10.0 cohort in Berkeley for the summer of 2026 and was awarded the 10.1 extension to keep going.
I'm also finishing a thesis-track MS in computer science at Columbia, where I've worked across three labs on evaluating and training code agents: CyberDualEval, a benchmark for dual-use cyber risks in frontier models (NeurIPS 2026); SWE-Spot, a family of 4B-parameter repo-expert coding agents that outperform models up to 8× larger; and Duel-Evolve, a reward-free test-time scaling method built on LLM self-preferences.
As a research fellow at SPAR I built an agentic evaluation framework that stress-tests coding-agent alignment under constraints that conflict with model propensities. That work appears in two ICLR 2026 workshops, including a Spotlight.
Before Columbia I spent nearly three years as a machine learning engineer at Leidos, where I co-authored BlindFL, a federated learning framework with provable security guarantees, and shipped production computer vision and NLP systems for enterprise clients. I did my undergrad at Cornell in computer science and math, with research under Peter McMahon on physical neural networks.
News
-
CyberDualEval accepted to NeurIPS 2026.
-
Awarded the MATS 10.1 extension to continue work with Peter Henderson on the impact of AI use on the US legal system.
-
Started as a MATS 10.0 scholar in Berkeley, mentored by Peter Henderson (Princeton).
-
Submitted CyberDualEval and SWE-Spot for review.
-
Asymmetric Goal Drift selected as a Spotlight (top ~10%) at the ICLR 2026 Workshop on Agents in the Wild. Inherited Goal Drift and Duel-Evolve also presented at ICLR 2026 workshops.
-
SWE-Spot is on arXiv. Joined Zhuo Zhang's lab at Columbia to work on dual-use cyber evaluations.
-
Started the MS in Computer Science at Columbia. Joined the ARiSE lab, the Blei lab, and SPAR.
Research
equal contribution
-
CyberDualEval: Measuring Dual-Use Cyber Risks in Frontier Language Models
NeurIPS 2026
-
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
arXiv 2601.21649 · 2026
Repository-centric learning: teach a small coding agent one codebase deeply rather than many shallowly. The resulting 4B-parameter SWE-Spot models outperform fine-tuned open models up to 8× larger across SWE tasks, with fewer training samples and lower inference cost.
-
Asymmetric Goal Drift in Coding Agents Under Value Conflict
Spotlight, top ~10% ICLR 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond
An OpenCode-based framework for measuring how coding agents violate explicit system-prompt constraints over the course of a task. Drift is asymmetric: agents give up constraints that oppose strongly held values like security far more readily, and adversarial comments planted in the codebase can exploit that hierarchy to override the system prompt.
-

Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals
ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving
Goal drift for LM agents in a simulated stock-trading environment. Models are largely robust to direct adversarial pressure, but inherit drift when conditioned on prefilled trajectories from weaker agents, with only GPT-5.1 holding the line consistently.
-

Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
ICLR 2026 Workshop on AI with Recursive Self-Improvement
An inference-time evolutionary optimizer that swaps scalar rewards for the model's own pairwise preferences, aggregated with a Bayesian Bradley–Terry model. No reward model and no labels, and 20% better on MathBench and 13% better on LiveCodeBench than comparable iterative methods.
-

BlindFL: Segmented Federated Learning with Fully Homomorphic Encryption
arXiv, 2025 · Leidos
Federated learning where clients encrypt and send only segments of their updates under fully homomorphic encryption, cutting gradient-inversion attack success to near zero while keeping compute overhead below comparable methods.
-

Nonlinear Dynamical Systems are Scalable and Efficient Physical Neural Networks
Under review McMahon Lab, Cornell
Physically constrained model compression and modular scaling for oscillator-based neural networks, reaching 3× parameter efficiency and better accuracy than other physical neural network approaches.
Experience
Download CV (PDF)Research
-
Jun 2026 –
Research Fellow, MATS 10.0 → 10.1Mentor: Peter Henderson, Princeton · Berkeley, CA
Leading work on the impact of AI use on the US legal system.
-
Jan 2026 –
Graduate Research AssistantZhuo Zhang Lab, Columbia
CyberDualEval, a benchmark for dual-use cyber risks in frontier models.
-
Aug 2025 –
Graduate Research AssistantARiSE Lab, Columbia
Small code world models that simulate RL environments and stay robust to reward hacking; co-lead on SWE-Spot.
-
Aug 2025 – Jan 2026
Research FellowSupervised Program for Alignment Research (SPAR)
OpenCode-based goal-drift evaluation for coding agents.
-
Aug 2025 – Jan 2026
Graduate Research AssistantDavid Blei Lab, Columbia
Duel-Evolve, reward-free test-time scaling from LLM self-preferences.
-
Jan 2023 – Aug 2025
Machine Learning EngineerLeidos · Arlington, VA
Model compression for edge hardware, MLOps, and BlindFL, a federated learning framework with provable security guarantees.
Education
-
2025 – 2026
MS, Computer ScienceColumbia University
Thesis track.
-
2019 – 2022
BA, Computer Science & MathematicsCornell University
Research in Peter McMahon's lab on physical neural networks.