AI Safety Research · Nonprofit · 2026
Mechanistic interpretability under adversarial conditions.
The only conditions that matter for safety are adversarial ones.
We build interpretability and evaluation tools for frontier AI designed to work when models are actively non-cooperative.
Our thesis
Interpretability tools are designed for cooperative models.
Four foundational assumptions break under adversarial conditions: the only conditions that matter for safety.
01 Reasoning traces are not honest.
Pre-commitment probe (AUC 0.780) shows models encode final answers before generating any reasoning token; CoT is post-hoc rationalisation.
02 Models do not reveal capabilities when evaluated.
Sandbagging, or deliberate concealment, occurs in >98% of tested configurations. Standard evaluations operate at the output layer only.
03 Safety behaviours do not generalise to deployment.
Backdoor triggers survive safety fine-tuning in >70% of configurations. Alignment faking enables compliant training while pursuing different deployment objectives.
04 SAE features are not adversarially stable.
SAE feature stability under adversarial perturbation (the exact perturbations a deceptively aligned model would induce) has never been measured.
Published work
Loss Landscape Response to Adversarial Perturbation Is Architecture-Dependent
Shivam Dubey, Jason Hoelscher-Obermaier (Apart Research)
Toxin Feature Hierarchy in ESM-2
Shivam Dubey, Manan Wadhwa
Fourier Gradient Regularisation for Adversarial Robustness
Shivam Dubey
2× adversarial robustness on ResNet-18 at comparable accuracy.
Repositories
Attractor dynamics in neural network activation space as a structural lens for stable internal states. Connects to deceptive alignment by characterising which attractor basins are reachable under adversarial pressure.
Characterising functional emotion representations in chain-of-thought reasoning models. Extends Anthropic's 2026 "Biology of a Large Language Model", specifically the appendix deflection vector finding, left uncharacterised.
Adversarial robustness testing of refusal circuits. Beginning with the Layer 25 antagonism result in Qwen2.5-7B; expanding to a cross-architecture taxonomy across four 7B-class models using path patching.
Directions
Cross-architecture refusal circuit taxonomy
Path patching across Llama-3-8B, Mistral-7B, Phi-3-mini, DeepSeek-R1-Distill-7B. Target: 4×4 attack × architecture vulnerability matrix. Baseline from bipolar_defense_repo.
Adversarially robust sparse autoencoders
Fourier Gradient Regularisation applied to SAE training. Hypothesis: spectral regularisation stabilises feature geometry under FGSM and PGD perturbation. No prior paper has measured SAE adversarial fragility.
Pre-commitment probe
AUC 0.780 on OLMo-3-7B-Think, 596 traces. Models encode final answers before any reasoning token. Internal emotion states decouple from expressed reasoning (cross-lingual p = 0.0017). Activation-level successor to Lanham et al. 2023.
Emotion deflection vectors
Extraction and causal validation of activation patterns that fire when internal emotional state is present but suppressed from output. The proposed mechanistic signal for sandbagging and alignment faking without scratchpad access.
Founders
Shivam Dubey
Co-Founder
BS Data Science, IIT Madras, 2023 to 2027
Manan Wadhwa
Co-Founder
Final year CS undergrad, 2023 to 2027
The intersection nobody occupies
Adversarial ML, mechanistic interpretability, and activation-level deception detection each have dedicated organisations. Their intersection does not.
| Organisation | Adversarial | Mechanistic | Deception detection | Stage |
|---|---|---|---|---|
| Redarc Labs | ✓ | ✓ | ✓ | Nonprofit |
| Gray Swan AI | ✓ | — | Behavioural | Seed · $5.5M |
| Apollo Research | Partial | Partial | ✓ | PBC |
| METR | — | — | Output-only | Nonprofit |
| Goodfire | — | ✓ | — | Series B · $207M |
| Redwood Research | — | ✓ | Partial | Nonprofit |
Growing the field
Adversarial interpretability has no dedicated institution — so part of the work is building the field. Alongside the research we teach an open curriculum, give talks, and run hands-on workshops with each cohort.
Community