← Workshops Workshop

Foundations of Adversarial Interpretability

An introductory series taking the cohort from linear probes to their first working deflection-vector detector on an open-weights model.

Feb 10, 2026 · Cohort 01 · 3-session hands-on series · Remote

cohorthands-onprobes
Materials

Placeholder — replace with the real workshop write-up

A short description of what this workshop covered and who it was for.

Session breakdown

  1. Linear probes & activation basics — extracting and reading residual-stream activations.
  2. From probes to steering — constructing and applying steering vectors.
  3. Building a detector — a capstone where each participant ships a working deflection-vector detector.

What participants leave with

  • A runnable notebook and reusable tooling.
  • A mental model for when interpretability holds and when it breaks.