A running log of what our research community has been working through, week by week — newest first.
Week 3 Jun 15, 2026
Adversarial Robustness & Refusal Circuits
Turned interpretability tools on themselves — FGSM/PGD attacks — and started replicating refusal-circuit findings across new architectures.
adversarialrefusal circuits
Week 2 Jun 8, 2026
Sparse Autoencoders & Feature Geometry
Trained our first SAEs, pulled features out of a small transformer, and argued about superposition.
SAEfeatures
Week 1 Jun 1, 2026
Foundations & Onboarding
Kicked off with transformer circuits and got everyone's environments set up for activation-patching experiments.
onboardingcircuits