Magnus Saebo

Columbia University  ·  MATS

Magnus Saebo

I measure societal risk from AI, and build training pipelines for aligned, specialized agents.

The measuring happens in cybersecurity, software engineering, and law: dual-use cyber evaluations, coding agents drifting from their instructions, and language models sitting as judges.

me. click for more. 1 / 5

About

I'm a research fellow at MATS, mentored by Peter Henderson (Princeton), where I lead work on the impact of AI use on the US legal system. I joined the 10.0 cohort in Berkeley for the summer of 2026 and was awarded the 10.1 extension to keep going.

I'm also finishing a thesis-track MS in computer science at Columbia, where I've worked across three labs on evaluating and training code agents: CyberDualEval, a benchmark for dual-use cyber risks in frontier models (NeurIPS 2026); SWE-Spot, a family of 4B-parameter repo-expert coding agents that outperform models up to 8× larger; and Duel-Evolve, a reward-free test-time scaling method built on LLM self-preferences.

As a research fellow at SPAR I built an agentic evaluation framework that stress-tests coding-agent alignment under constraints that conflict with model propensities. That work appears in two ICLR 2026 workshops, including a Spotlight.

Before Columbia I spent nearly three years as a machine learning engineer at Leidos, where I co-authored BlindFL, a federated learning framework with provable security guarantees, and shipped production computer vision and NLP systems for enterprise clients. I did my undergrad at Cornell in computer science and math, with research under Peter McMahon on physical neural networks.

News

  1. CyberDualEval accepted to NeurIPS 2026.

  2. Awarded the MATS 10.1 extension to continue work with Peter Henderson on the impact of AI use on the US legal system.

  3. Started as a MATS 10.0 scholar in Berkeley, mentored by Peter Henderson (Princeton).

  4. Submitted CyberDualEval and SWE-Spot for review.

  5. Asymmetric Goal Drift selected as a Spotlight (top ~10%) at the ICLR 2026 Workshop on Agents in the Wild. Inherited Goal Drift and Duel-Evolve also presented at ICLR 2026 workshops.

  6. SWE-Spot is on arXiv. Joined Zhuo Zhang's lab at Columbia to work on dual-use cyber evaluations.

  7. Started the MS in Computer Science at Columbia. Joined the ARiSE lab, the Blei lab, and SPAR.

Research

* equal contribution

  1. CyberDualEval: Measuring Dual-Use Cyber Risks in Frontier Language Models

    Magnus Saebo, Francesco Piccoli, Maxwell Watson, Pranjali Thakur, Yangruibo Ding, Zhuo Zhang

    NeurIPS 2026

  2. SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning

    Jinjun Peng*, Magnus Saebo*, Tianjun Zhong, Yi-Jie Cheng, Junfeng Yang, Baishakhi Ray, Simin Chen, Yangruibo Ding

    arXiv 2601.21649 · 2026

    Repository-centric learning: teach a small coding agent one codebase deeply rather than many shallowly. The resulting 4B-parameter SWE-Spot models outperform fine-tuned open models up to 8× larger across SWE tasks, with fewer training samples and lower inference cost.

  3. Asymmetric Goal Drift in Coding Agents Under Value Conflict

    Magnus Saebo, Spencer Gibson, Tyler Crosse, Achu Menon, Eyon Jang, Diogo Cruz

    Spotlight, top ~10% ICLR 2026 Workshop on Agents in the Wild: Safety, Security, and Beyond

    An OpenCode-based framework for measuring how coding agents violate explicit system-prompt constraints over the course of a task. Drift is asymmetric: agents give up constraints that oppose strongly held values like security far more readily, and adversarial comments planted in the codebase can exploit that hierarchy to override the system prompt.

  4. Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals

    Achu Menon, Magnus Saebo, Tyler Crosse, Spencer Gibson, Eyon Jang, Diogo Cruz

    ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving

    Goal drift for LM agents in a simulated stock-trading environment. Models are largely robust to direct adversarial pressure, but inherit drift when conditioned on prefilled trajectories from weaker agents, with only GPT-5.1 holding the line consistently.

  5. Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences

    Sweta Karlekar*, Carolina Zheng*, Magnus Saebo*, Nicolas Beltran-Velez, Shuyang Yu, John Bowlan, Michal Kucer, David Blei

    ICLR 2026 Workshop on AI with Recursive Self-Improvement

    An inference-time evolutionary optimizer that swaps scalar rewards for the model's own pairwise preferences, aggregated with a Bayesian Bradley–Terry model. No reward model and no labels, and 20% better on MathBench and 13% better on LiveCodeBench than comparable iterative methods.

  6. BlindFL: Segmented Federated Learning with Fully Homomorphic Encryption

    Evan Gronberg*, Liv d'Aliberti*, Magnus Saebo*, Aurora Hook

    arXiv, 2025 · Leidos

    Federated learning where clients encrypt and send only segments of their updates under fully homomorphic encryption, cutting gradient-inversion attack success to near zero while keeping compute overhead below comparable methods.

  7. Nonlinear Dynamical Systems are Scalable and Efficient Physical Neural Networks

    Magnus Saebo*, Tyler King*, Maxwell Anderson, Tatsuhiro Onodera, Peter McMahon

    Under review McMahon Lab, Cornell

    Physically constrained model compression and modular scaling for oscillator-based neural networks, reaching 3× parameter efficiency and better accuracy than other physical neural network approaches.

Research

  1. Jun 2026 –
    Research Fellow, MATS 10.0 → 10.1
    Mentor: Peter Henderson, Princeton · Berkeley, CA

    Leading work on the impact of AI use on the US legal system.

  2. Jan 2026 –
    Graduate Research Assistant
    Zhuo Zhang Lab, Columbia

    CyberDualEval, a benchmark for dual-use cyber risks in frontier models.

  3. Aug 2025 –
    Graduate Research Assistant
    ARiSE Lab, Columbia

    Small code world models that simulate RL environments and stay robust to reward hacking; co-lead on SWE-Spot.

  4. Aug 2025 – Jan 2026
    Research Fellow
    Supervised Program for Alignment Research (SPAR)

    OpenCode-based goal-drift evaluation for coding agents.

  5. Aug 2025 – Jan 2026
    Graduate Research Assistant
    David Blei Lab, Columbia

    Duel-Evolve, reward-free test-time scaling from LLM self-preferences.

  6. Jan 2023 – Aug 2025
    Machine Learning Engineer
    Leidos · Arlington, VA

    Model compression for edge hardware, MLOps, and BlindFL, a federated learning framework with provable security guarantees.

Education

  1. 2025 – 2026
    MS, Computer Science
    Columbia University

    Thesis track.

  2. 2019 – 2022
    BA, Computer Science & Mathematics
    Cornell University

    Research in Peter McMahon's lab on physical neural networks.