← Research log

Week 2 · Jun 8, 2026

Sparse Autoencoders & Feature Geometry

Trained our first SAEs, pulled features out of a small transformer, and argued about superposition.

SAEfeatures

Placeholder log entry — replace with the real recap

A short narrative of what the community worked through this week.

What we covered

  • SAE training and dictionary learning on a small transformer.
  • First pass at feature extraction; inspected what the learned features fire on.
  • Discussion: polysemanticity and superposition — when is a “feature” real?

Open threads

  • How to evaluate feature quality beyond eyeballing activations.
  • Whether our features transfer across two model sizes.