Redarc Labs
  • Publications
  • Research
  • Community
  • About
  • Team
Contact
Publications Research Community About Team
Contact
← Community Curriculum

Curriculum

The course we wish existed when we started. Each module builds on the last — from reading activations to building detection that holds up under adversarial pressure.

01

Foundations: Activations, Probes & Features

Week 1–2

The groundwork — residual streams, linear probes, and the linear representation hypothesis. Everything later modules build on.

foundationsprobes
02

Steering & Control Vectors

Week 3–4

From reading features to changing behaviour — constructing steering vectors, applying them, and measuring their effects.

steeringcontrol
03

Adversarial Settings: When Models Don't Cooperate

Week 5–6

The core of the field — what breaks when a model is actively non-cooperative, and how to build detection that survives adversarial pressure.

adversarial interpdeception detection
Redarc Labs

Work

  • Toxin Feature Hierarchy, ICML 2026
  • Attractor Framework
  • Thinking Model Emotions
  • Bipolar Defense

Contact

  • hello@redarclabs.com
  • LinkedIn
Redarc Labs · Where the safety layer holds
redarclabs.com