The course we wish existed when we started. Each module builds on the last — from reading activations to building detection that holds up under adversarial pressure.
01Foundations: Activations, Probes & Features
Week 1–2The groundwork — residual streams, linear probes, and the linear representation hypothesis. Everything later modules build on.
foundationsprobes
Steering & Control Vectors
Week 3–4From reading features to changing behaviour — constructing steering vectors, applying them, and measuring their effects.
steeringcontrol
Adversarial Settings: When Models Don't Cooperate
Week 5–6The core of the field — what breaks when a model is actively non-cooperative, and how to build detection that survives adversarial pressure.
adversarial interpdeception detection