← Workshops Workshop

Detecting Deception at the Activation Level

A one-day intensive on activation-level deception detection: building evaluation harnesses, red-teaming your own detectors, and reasoning about non-cooperative models.

May 19, 2026 · Cohort 02 · Full-day intensive · Remote

cohortdeception detectionevaluation

Placeholder — replace with the real workshop write-up

A short description of what this workshop covered and who it was for.

Session breakdown

  1. Threat modelling — what “non-cooperative” actually means for interpretability.
  2. Evaluation harnesses — measuring whether a detector works under adversarial pressure.
  3. Red-teaming — breaking your own detector, then hardening it.

What participants leave with

  • An evaluation harness they can point at their own models.
  • A checklist for stress-testing interpretability claims.