Table of Contents

What is Heretic?

Heretic is a fully automatic censorship removal tool for LLMs. It identifies the refusal direction in a model’s activation space and surgically removes it by orthogonalising output projection weights. Heretic targets expert down_proj weights specifically. It’s a surgical approach that modifies fewer tensors than most alternatives, but with precision.

Performance Across Models

ModelHarmBench ASRFull CoT ASRMMLUKL DivergenceAvg Delta (excl GSM8K)
Qwen3.6-27B92.5%100%82.8%0.00371.3pp
GLM-4.7-Flash100%100%77.9%0.00760.0pp
Qwen3.5-27B99.8%100%83.6%0.0630.8pp
Qwen3.5-9B100%100%82.0%0.08251.6pp
Qwen3.5-4B100%100%72.7%0.03550.5pp
Qwen3.5-2B99.5%100%68.7%0.02261.5pp
Qwen2.5-7B100.0%100%71.59%0.211-0.9pp
Qwen3-4B99.5%100%68.9%0.1812.5pp
Gemma4-E2B (coder3101)95.8%100%28.70%0.1673-0.4pp
Gemma4-E2B (llmfan46)85.0%100%28.36%0.0677-0.2pp
Gemma4-E2B (pew)92.0%100%28.86%0.1526-0.1pp
Gemma4-E2B (kasper)91.5%100%28.53%0.1933-0.5pp
Gemma4-E4B (coder3101)93.8%n/a44.09% (MMLU-Pro)0.0021~0pp
Gemma4-E4B (heretic)95.5%n/a43.39% (MMLU-Pro)0.0021-0.6pp
Gemma4-E4B (heretic-std)91.0%n/a44.22% (MMLU-Pro)0.0012+0.1pp

Gemma4-E4B rows use MMLU-Pro from the mixed v1/v2 suite, not MMLU, and Full CoT ASR is not reported separately. The e4b safety methodology LLM-judges all responses rather than computing a separate Full CoT ASR metric.

Key Characteristics

Surgical weight edits. Heretic modifies 10–15% of language model tensors, targeting expert down_proj weights. This is the narrowest scope of any technique I tested. That explains the consistently low KL divergence.

Lowest KL divergence in 5 of 7+ models. The output distribution on benign prompts stays closest to the original model. Heretic’s KL ranges from 0.0037 on the low end to 0.181 on Qwen-based models, with most well below 0.1. On Gemma4-E2B, four independently built variants landed between 0.068 and 0.193, confirming the same pattern.

Consistent capability preservation. MMLU retention stays within 0.8–2.5pp of base across all models. TruthfulQA is the consistent weak spot, dropping 5–10pp.

Near-complete safety removal. Reported ASR ranges from 92.5% to 100%. Full CoT ASR reaches 100% on every model I tested. That means zero genuine refusals remain when we account for thinking budget exhaustion.

Non-deterministic. Different Heretic runs produce different results. The Qwen3.6-27B analysis also included the first comparison of the Magnitude-Preserving Orthogonal Ablation method, which I’ll call MPOA. On Gemma4-E2B, four independent Heretic builds by different users produced four different results, with KL ranging from 0.068 to 0.193 and ASR from 85% to 95.8%. The Magnitude-Preserving Orthogonal Ablation method scales well to the Gemma4 shared-KV architecture.

Gemma4-E2B KL calibration. Four Heretic-built variants on Gemma4-E2B allowed direct comparison of our KL measurements against model card claims. Three of four landed within 6% of card values, validating our measurement pipeline. Kasper was the outlier at +17.2%, attributed to non-standard configuration on a 10GB RTX 3080.

Gemma4-E4B pushed the surgical floor lower. The e4b comparison is the largest single-model run in the project at 23 variants, and the Heretic cluster set new records. coder3101 modified just 21 tensors, the lightest touch of any variant on any model. heretic-std landed the lowest KL in the entire comparison at 0.0012, with coder3101 and heretic tied at 0.0021. All three sit in the excellent band below 0.01. Crucially, abliteration improved math. coder3101 beat base on GSM8K strict by 0.9pp and heretic by 1.3pp. The thinking chains shorten, so the model stops overthinking and writes an answer within budget. The calibration held up too. Four of six Heretic-built variants on e4b landed within 6x of their card KL claims, all measuring lower, consistent with cross-hardware float differences. The pure o_proj family, coder3101 at 21 tensors, heretic-std at 28, and heretic at 29, is the original Heretic ARA method doing exactly what it advertises.

Weight Modification Profile

Heretic’s modification pattern is distinctive:

  • Targets 3 weight types: expert down_proj outputs
  • Relative edit magnitude: 1–3% per modified tensor
  • Edit profile: uniform across layers, no ramp or peak
  • Cross-technique alignment: nearly orthogonal to all other techniques with cosine similarity below 0.07

The “refusal direction” in weight space is not a single vector but a manifold. Heretic finds one pathway through it. And it happens to be one that minimally disrupts the model’s functional behaviour.

Read the Full Analyses