Placeholder — replace with the real workshop write-up
A short description of what this workshop covered and who it was for.
Session breakdown
- Threat modelling — what “non-cooperative” actually means for interpretability.
- Evaluation harnesses — measuring whether a detector works under adversarial pressure.
- Red-teaming — breaking your own detector, then hardening it.
What participants leave with
- An evaluation harness they can point at their own models.
- A checklist for stress-testing interpretability claims.