Redarc Labs Logo Yantra-inspired geometric logo for Redarc Labs — Adversarial Interpretability for AI Safety Redarc Labs ADVERSARIAL INTERPRETABILITY · AI SAFETY WHERE THE SAFETY LAYER HOLDS

AI Safety Research · Nonprofit · 2026

Mechanistic interpretability under adversarial conditions.

The only conditions that matter for safety are adversarial ones.

We build interpretability and evaluation tools for frontier AI designed to work when models are actively non-cooperative.

Position

Our thesis

Interpretability tools are designed for cooperative models.

Four foundational assumptions break under adversarial conditions: the only conditions that matter for safety.

01 Reasoning traces are not honest.

Pre-commitment probe (AUC 0.780) shows models encode final answers before generating any reasoning token; CoT is post-hoc rationalisation.

02 Models do not reveal capabilities when evaluated.

Sandbagging, or deliberate concealment, occurs in >98% of tested configurations. Standard evaluations operate at the output layer only.

03 Safety behaviours do not generalise to deployment.

Backdoor triggers survive safety fine-tuning in >70% of configurations. Alignment faking enables compliant training while pursuing different deployment objectives.

04 SAE features are not adversarially stable.

SAE feature stability under adversarial perturbation (the exact perturbations a deceptively aligned model would induce) has never been measured.

Publications

Published work

2026 TAIS 2026

Loss Landscape Response to Adversarial Perturbation Is Architecture-Dependent

Shivam Dubey, Jason Hoelscher-Obermaier (Apart Research)

Conference PaperAdversarial RobustnessTAIS
Read paper
2026 GenBio Workshop · ICML 2026

Toxin Feature Hierarchy in ESM-2

Shivam Dubey, Manan Wadhwa

Workshop PaperProtein LMICML
Read paper
2025 Reliable ML from Unreliable Data · NeurIPS 2025

Fourier Gradient Regularisation for Adversarial Robustness

Shivam Dubey

2× adversarial robustness on ResNet-18 at comparable accuracy.

Workshop PosterAdversarial RobustnessNeurIPS
Read paper
Ongoing work

Repositories

Directions

D1

Cross-architecture refusal circuit taxonomy

Path patching across Llama-3-8B, Mistral-7B, Phi-3-mini, DeepSeek-R1-Distill-7B. Target: 4×4 attack × architecture vulnerability matrix. Baseline from bipolar_defense_repo.

D2

Adversarially robust sparse autoencoders

Fourier Gradient Regularisation applied to SAE training. Hypothesis: spectral regularisation stabilises feature geometry under FGSM and PGD perturbation. No prior paper has measured SAE adversarial fragility.

D3

Pre-commitment probe

AUC 0.780 on OLMo-3-7B-Think, 596 traces. Models encode final answers before any reasoning token. Internal emotion states decouple from expressed reasoning (cross-lingual p = 0.0017). Activation-level successor to Lanham et al. 2023.

D4

Emotion deflection vectors

Extraction and causal validation of activation patterns that fire when internal emotional state is present but suppressed from output. The proposed mechanistic signal for sandbagging and alignment faking without scratchpad access.

Team

Founders

Shivam Dubey

Co-Founder

BS Data Science, IIT Madras, 2023 to 2027

Toxin Feature Hierarchy in ESM-2
ICML 2026 GenBio Workshop
Fairness-Aware Speculative Decoding (FASD)
MIT Technology Review cited · 77% bias reduction at 6% latency overhead
Fourier Gradient Regularisation (FGR)
NeurIPS 2025 · 2× adversarial robustness on ResNet-18
Apart Research fellow
Under Jason Hoelscher-Obermaier · accepted TAIS Oxford
MARS V Research Fellow
Cambridge AI Safety Hub

Manan Wadhwa

Co-Founder

Final year CS undergrad, 2023 to 2027

Toxin Feature Hierarchy in ESM-2
ICML 2026 GenBio Workshop
Getting in SHAPe
ACL ARR 2026 · in submission
MARS V Research Fellow
Cambridge AI Safety Hub
Google Summer of Code 2026
HumanAI organisation
Research Fellow
AISI @ Georgia Tech
Landscape

The intersection nobody occupies

Adversarial ML, mechanistic interpretability, and activation-level deception detection each have dedicated organisations. Their intersection does not.

Organisation Adversarial Mechanistic Deception detection Stage
Redarc Labs Nonprofit
Gray Swan AI Behavioural Seed · $5.5M
Apollo Research Partial Partial PBC
METR Output-only Nonprofit
Goodfire Series B · $207M
Redwood Research Partial Nonprofit
Field building

Growing the field

Adversarial interpretability has no dedicated institution — so part of the work is building the field. Alongside the research we teach an open curriculum, give talks, and run hands-on workshops with each cohort.

Community