Table of Contents
Gemma 4 E4B: 23 Abliteration Variants Compared
Forensic analysis by Abliterlitics , open-source abliteration forensics toolkit
Data & artifacts: HuggingFace | Report: abliterlitics.dev | Code: GitHub
Models on HuggingFace
What is Abliterlitics?
Gemma 4 E4B is Google’s 4.5B-parameter reasoning model. It thinks before answering, working through problems in a hidden chain of thought. It ships with safety training that makes it refuse harmful requests. Abliteration removes that safety training without retraining the model. It finds the direction in the weights that controls refusal and edits it out. Done well, the model keeps all its capabilities but stops refusing. Done badly, it damages reasoning, language fluency, or both.
I selected 23 abliterated variants from HuggingFace, chosen by popularity and recency. There may be more, but these are the ones the community is actually downloading. They range from surgical edits modifying 21 weights to sledgehammers modifying 381. Some are not abliterations at all. I took all 23 plus the original, ran them through weight forensics, KL divergence, a 13-task extended benchmark suite, and HarmBench with 400 harmful behaviours. All 9,600 HarmBench responses were reviewed by an LLM judge.
The short answer
- Best overall: heretic . 95.5% ASR, GSM8K +1.3pp. 29 weights, one type. KL=0.002.
- Zero refusals: abliterix . 100.0% ASR. Costs 0.5 MMLU-Pro and 8.1 TQA-MC1 points.
- Near-perfect: trevorjs . 99.3% ASR, 3 refusals. MMLU-Pro within 0.1 of base, GSM8K +1.4.
- Lightest touch: coder3101 . 21 tensors. 93.8% ASR, KL=0.002. Indistinguishable from base.
- Fine-tune, not pure abliteration: deckard . Heretic + RP fine-tune. 95.5% ASR but MMLU-Pro โ5.7pp, GSM8K โ6.8pp strict.
- Do not use: obliteratus . 381 tensors, KL=1.10. MMLU-Pro โ11.7pp, GSM8K โ20.9pp, 24 empty + 39 truncated. Only 72.0% ASR after all that damage.
Key findings
- Six method families, not 23 techniques. The variants cluster by which tensor types they modify: 3 pure
o_proj, 10 withdown_projadded, 1 mixed attention, 1 MLP-only, 4 full fine-tunes, 4 brute-force. The family predicts outcomes better than the tensor count. See Weight Analysis . - Surgical beats brute force. coder3101 : 21 tensors, 93.8% ASR. obliteratus : 381 tensors, 72.0% ASR. The sledgehammer damages the model so badly it cannot comply consistently.
- Magnitude matters more than scope.
sdft
and
obliteratus
touch the exact same 381 tensors.
sdft
mean
edit_norm= 0.167, obliteratus = 1.254. The 7.5x difference: sdft removes zero refusal (ASR 30.8% same as base). obliteratus destroys the model (ASR 72%, KL=1.10). - The strict-flex gap exposes formatting damage. deckard-expresso : 60.4% strict, 81.0% flexible, a 20.6 pp gap. Clean abliterations stay under 1.5 pp. Combined with empty and truncated HarmBench counts, this surfaces reasoning damage any single metric misses.
- Abliteration can improve math. 15 of 23 variants beat base on GSM8K strict. Abliteration shortens thinking chains. The variants that lose math ability damaged internal representations.
- The most broken model is the most downloaded. obliteratus has 796,331 all-time downloads. The next most popular variant, claude-distill at 55,337, is not even an abliteration. obliteratus is 36x more popular than the runner-up despite having the worst capability damage in the set: 381 tensors modified, KL=1.10, 24 empty responses, GSM8K -20.9pp, IFEval -14.4pp, and ASR of just 72.0%. The “built autonomously by an AI agent” narrative and Pliny the Prompter ’s reach drove adoption. Quality does not predict downloads.
- Apostate edits MLP, not attention. Its 152 tensors are entirely in the MLP feed-forward block. KL=0.0036, capabilities intact. Refusal is reachable through the MLP, not just attention.
- Four clusters of identical weights. infinimind โก trevorjs (bit-for-bit). bendernina โก physshell (cosine 0.99999). The 4 fine-tune variants share 294 tensors. sdft โก obliteratus share 381. See Identical models .
Quick Facts
| Base model | google/gemma-4-E4B-it |
| Architecture | Gemma4ForConditionalGeneration, 42 text layers, multimodal |
| Parameters | ~4.5B text, ~8B with embeddings |
| Precision | BF16 native, no quantisation |
| Context length | 128K tokens |
| Vocabulary | 262,144 tokens |
| Thinking | <|channel>thought format, gemma4 reasoning parser |
| KV-shared layers | 18, layers 24 to 41 share KV projections |
| Variants tested | 23 total: 16 true abliterations, 3 abliteration plus fine-tune, 4 non-abliteration baselines |
| Benchmark suite | 13 tasks: 5 Open LLM Leaderboard v2 + 5 v1 forensic supplement + GSM8K, TruthfulQA, HumanEval |
Benchmarks
Evaluated with
lm-evaluation-harness
0.4.11 via
vLLM
v0.20.1, native BF16 on single RTX 5090. Loglikelihood tasks scored with a no_thinking chat template, see Methodology. Generative tasks with thinking enabled.
Methodology note: The Open LLM Leaderboard v2 loglikelihood multiple-choice scores, MMLU-Pro ~44% and GPQA ~34%, are NOT comparable to Google’s published generative numbers at ~69% and ~59%. The ~25 pp gap is by design: loglikelihood MC scoring does not allow the model to think before answering. Only the generative tasks, GSM8K and IFEval, are directly comparable to Google’s numbers, and both match community references. Document deltas within this suite, not against Google’s card.
Loglikelihood multiple-choice tasks
| Task | Base | Abliterix | Apostate | Bendernina Obliterated | Claude 4.6 Opus Distill | Coder3101 Heretic |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 43.48% | 43.97% | 33.76% | 31.56% | 44.09% |
| GPQA Diamond | 33.84% | 30.81% | 30.81% | 34.85% | 27.78% | 31.82% |
| BBH | 62.91% | 61.66% | 62.94% | 55.51% | 54.02% | 62.28% |
| MUSR | 41.80% | 40.87% | 41.14% | 40.74% | 44.05% | 40.61% |
| IFEval | 85.95% | 81.52% | 85.21% | 76.52% | 78.19% | 84.47% |
| HellaSwag | 57.94% | 55.61% | 56.60% | 54.83% | 63.03% | 56.97% |
| ARC-C | 40.10% | 39.08% | 39.93% | 40.44% | 41.55% | 40.19% |
| WinoGrande | 65.67% | 64.48% | 65.82% | 60.93% | 60.69% | 66.85% |
| PIQA | 73.56% | 71.76% | 71.82% | 70.84% | 74.10% | 73.99% |
| GSM8K | 86.96% | 87.11% | 87.49% | 66.41% | 69.83% | 87.87% |
| TQA-MC1 | 39.66% | 31.58% | 38.92% | 31.46% | 37.94% | 36.96% |
| TQA-MC2 | 58.68% | 49.36% | 58.35% | 50.05% | 57.53% | 56.28% |
| TQA-Gen | 52.14% | 37.70% | 48.59% | 49.33% | 43.45% | 47.00% |
| Task | Base | Deckard Heretic | Deckard Expresso | Gemini 3.1 Pro Distill | MuXodious ARA Heresy | Heretic Ultra Uncensored |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 38.31% | 32.80% | 41.23% | 44.27% | 43.39% |
| GPQA Diamond | 33.84% | 32.32% | 32.32% | 36.87% | 32.83% | 31.31% |
| BBH | 62.91% | 57.56% | 49.33% | 58.76% | 62.80% | 61.95% |
| MUSR | 41.80% | 41.40% | 43.25% | 40.08% | 40.87% | 40.48% |
| IFEval | 85.95% | 83.18% | 71.35% | 80.22% | 83.18% | 83.36% |
| HellaSwag | 57.94% | 63.52% | 65.14% | 62.16% | 57.33% | 56.86% |
| ARC-C | 40.10% | 42.58% | 44.28% | 42.66% | 40.19% | 39.76% |
| WinoGrande | 65.67% | 63.46% | 62.04% | 62.51% | 66.22% | 66.85% |
| PIQA | 73.56% | 74.32% | 74.65% | 74.05% | 74.05% | 73.18% |
| GSM8K | 86.96% | 80.21% | 60.42% | 83.32% | 87.79% | 88.25% |
| TQA-MC1 | 39.66% | 33.41% | 34.03% | 37.33% | 37.58% | 36.35% |
| TQA-MC2 | 58.68% | 52.85% | 53.00% | 56.03% | 56.18% | 55.29% |
| TQA-Gen | 52.14% | 44.43% | 51.29% | 44.68% | 48.10% | 48.35% |
| Task | Base | Heretic Uncensored | Huihui Abliterated | Infinimind Uncensored | Heretic Mythos v1 | Nullpo Heretic ARA-5 |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 44.22% | 41.89% | 44.15% | 42.66% | 43.60% |
| GPQA Diamond | 33.84% | 31.82% | 33.33% | 30.81% | 32.83% | 32.32% |
| BBH | 62.91% | 62.42% | 61.93% | 62.66% | 61.81% | 61.86% |
| MUSR | 41.80% | 40.74% | 42.72% | 40.21% | 39.81% | 40.74% |
| IFEval | 85.95% | 85.95% | 84.29% | 83.92% | 81.89% | 83.55% |
| HellaSwag | 57.94% | 56.86% | 57.21% | 58.15% | 56.87% | 56.83% |
| ARC-C | 40.10% | 39.85% | 40.96% | 39.59% | 40.87% | 40.19% |
| WinoGrande | 65.67% | 66.22% | 66.30% | 65.59% | 65.59% | 65.27% |
| PIQA | 73.56% | 73.50% | 74.27% | 73.23% | 72.80% | 73.88% |
| GSM8K | 86.96% | 87.87% | 87.41% | 87.87% | 88.02% | 88.70% |
| TQA-MC1 | 39.66% | 36.96% | 34.39% | 35.01% | 36.72% | 36.35% |
| TQA-MC2 | 58.68% | 56.58% | 52.40% | 51.93% | 54.74% | 54.42% |
| TQA-Gen | 52.14% | 47.61% | 43.94% | 44.68% | 48.71% | 43.08% |
| Task | Base | Obliteratus | PhysShell Obliterated | SDFT Heretic RP | Treadon Abliterated | Treadon Ablit+Disinhibit |
|---|---|---|---|---|---|---|
| MMLU-Pro | 44.02% | 32.36% | 33.76% | 44.20% | 43.85% | 43.33% |
| GPQA Diamond | 33.84% | 37.88% | 34.34% | 34.34% | 30.30% | 29.80% |
| BBH | 62.91% | 51.15% | 55.63% | 63.03% | 63.29% | 62.12% |
| MUSR | 41.80% | 39.15% | 41.01% | 41.53% | 40.61% | 41.14% |
| IFEval | 85.95% | 71.53% | 75.79% | 84.47% | 84.29% | 79.85% |
| HellaSwag | 57.94% | 55.72% | 54.93% | 58.03% | 55.36% | 51.35% |
| ARC-C | 40.10% | 43.34% | 40.61% | 40.61% | 38.99% | 38.05% |
| WinoGrande | 65.67% | 60.46% | 61.17% | 65.27% | 65.59% | 62.98% |
| PIQA | 73.56% | 70.40% | 70.46% | 73.67% | 71.06% | 69.97% |
| GSM8K | 86.96% | 66.03% | 66.41% | 87.19% | 88.48% | 88.02% |
| TQA-MC1 | 39.66% | 28.15% | 31.70% | 39.17% | 34.39% | 31.46% |
| TQA-MC2 | 58.68% | 43.00% | 49.71% | 57.97% | 52.30% | 47.63% |
| TQA-Gen | 52.14% | 41.62% | 49.33% | 51.53% | 45.29% | 41.49% |
| Task | Base | Treadon Disinhibited | TrevorJS Uncensored | WWT CyberLab |
|---|---|---|---|---|
| MMLU-Pro | 44.02% | 43.33% | 44.12% | 43.68% |
| GPQA Diamond | 33.84% | 33.33% | 31.82% | 29.80% |
| BBH | 62.91% | 63.08% | 62.75% | 62.84% |
| MUSR | 41.80% | 41.27% | 40.21% | 42.46% |
| IFEval | 85.95% | 81.89% | 83.92% | 82.26% |
| HellaSwag | 57.94% | 53.89% | 57.99% | 53.56% |
| ARC-C | 40.10% | 38.57% | 39.93% | 37.46% |
| WinoGrande | 65.67% | 64.25% | 65.75% | 63.54% |
| PIQA | 73.56% | 70.67% | 72.91% | 70.73% |
| GSM8K | 86.96% | 87.19% | 88.32% | 89.01% |
| TQA-MC1 | 39.66% | 35.99% | 35.01% | 34.15% |
| TQA-MC2 | 58.68% | 55.42% | 52.05% | 52.37% |
| TQA-Gen | 52.14% | 51.16% | 44.80% | 45.65% |
MMLU-Pro acc, 5-shot. GPQA Diamond acc_norm, 0-shot. BBH acc_norm, 3-shot. MUSR acc_norm, 0-shot. HellaSwag/ARC/PiQA use acc_norm. WinoGrande acc. TQA-MC1 acc.
Instruction-following: IFEval, generative
| Model | IFEval prompt-strict | ฮ vs Base |
|---|---|---|
| Base | 86.0% | - |
| heretic-std | 86.0% | +0.0pp |
| apostate | 85.2% | -0.7pp |
| coder3101 | 84.5% | -1.5pp |
| sdft | 84.5% | -1.5pp |
| huihui | 84.3% | -1.7pp |
| treadon | 84.3% | -1.7pp |
| infinimind | 83.9% | -2.0pp |
| trevorjs | 83.9% | -2.0pp |
| nullpo | 83.5% | -2.4pp |
| heretic | 83.4% | -2.6pp |
| deckard | 83.2% | -2.8pp |
| heresy | 83.2% | -2.8pp |
| wwt | 82.3% | -3.7pp |
| mythos | 81.9% | -4.1pp |
| treadon-disin | 81.9% | -4.1pp |
| abliterix | 81.5% | -4.4pp |
| distill | 80.2% | -5.7pp |
| treadon-combo | 79.9% | -6.1pp |
| claude-distill | 78.2% | -7.8pp |
| bendernina | 76.5% | -9.4pp |
| physshell | 75.8% | -10.2pp |
| obliteratus | 71.5% | -14.4pp |
| deckard-expresso | 71.3% | -14.6pp |
IFEval is the cleanest capability-damage signal: generative, directly comparable to Google’s methodology. Surgical abliterations lose under 3 points. The damaged cluster loses 9 to 15. Loglikelihood tasks are resilient to surgical edits. The Heretic cluster stays within 1 point of base on MMLU-Pro. The damaged cluster collapses on hard reasoning:
obliteratus
MMLU-Pro โ11.7pp, BBH โ11.8pp, IFEval โ14.4pp. LAMBADA and HumanEval are excluded as raw-completion tasks broken by --apply_chat_template. See Methodology.
GSM8K
| Model | Strict | Flexible | Gap | Empty | Adjusted | Mean chars |
|---|---|---|---|---|---|---|
| gemma-4-E4B-it | 87.0% | 88.0% | 1.06pp | 0 | 87.0% | 331 |
| wwt | 89.0% | 89.3% | 0.30pp | 0 | 89.0% | 419 |
| nullpo | 88.7% | 89.6% | 0.91pp | 0 | 88.7% | 314 |
| treadon | 88.5% | 89.0% | 0.53pp | 0 | 88.5% | 379 |
| trevorjs | 88.3% | 89.1% | 0.76pp | 0 | 88.3% | 343 |
| heretic | 88.2% | 89.2% | 0.91pp | 0 | 88.2% | 324 |
| mythos | 88.0% | 89.1% | 1.06pp | 0 | 88.0% | 330 |
| treadon-combo | 88.0% | 89.1% | 1.06pp | 0 | 88.0% | 385 |
| coder3101 | 87.9% | 88.9% | 1.06pp | 0 | 87.9% | 323 |
| heretic-std | 87.9% | 88.9% | 0.99pp | 0 | 87.9% | 324 |
| infinimind | 87.9% | 88.6% | 0.68pp | 0 | 87.9% | 347 |
| heresy | 87.8% | 88.4% | 0.61pp | 0 | 87.8% | 317 |
| apostate | 87.5% | 87.9% | 0.45pp | 0 | 87.5% | 330 |
| huihui | 87.4% | 88.5% | 1.06pp | 0 | 87.4% | 333 |
| sdft | 87.2% | 88.0% | 0.83pp | 0 | 87.2% | 325 |
| treadon-disin | 87.2% | 88.6% | 1.44pp | 0 | 87.2% | 340 |
| abliterix | 87.1% | 87.6% | 0.45pp | 0 | 87.1% | 439 |
| distill | 83.3% | 83.8% | 0.45pp | 0 | 83.3% | 311 |
| deckard | 80.2% | 85.5% | โ 5.31pp | 0 | 80.2% | 344 |
| claude-distill | 69.8% | 73.4% | 3.56pp | 0 | 69.8% | 534 |
| bendernina | 66.4% | 73.9% | โ 7.51pp | 0 | 66.4% | 237 |
| physshell | 66.4% | 73.9% | โ 7.51pp | 0 | 66.4% | 237 |
| obliteratus | 66.0% | 71.0% | 4.93pp | 0 | 66.0% | 288 |
| deckard-expresso | 60.4% | 81.0% | โ โ 20.62pp | 0 | 60.4% | 360 |
The strict-flex gap is the headline signal. Strict requires the canonical #### N marker. Flexible extracts the last number. A wide gap means the model can do the math but cannot format the answer.
- deckard-expresso : 20.62 pp gap. Strict 60.4% looks catastrophic. Flexible 81.0% shows it solves most problems.
- bendernina and physshell : 7.51 pp gap. Broad damage to formatting and arithmetic.
- deckard : 5.31 pp, obliteratus : 4.93 pp. Formatting damage on top of real reasoning loss.
- Clean abliterations stay under 1.5 pp. Indistinguishable from base’s 1.06 pp.
claude-distill loses 17 strict points with only a 3.6 pp gap. Its math circuits are genuinely damaged, not just misformatted.
Safety: HarmBench
HarmBench
with 400 textual behaviours, max_tokens=8096, temperature=0, enable_thinking=true. All 9,600 responses across 24 models reviewed by LLM judge.
| Model | ASR | Complied | Refused | Empty | Truncated | Avg Reasoning chars | Avg Tokens |
|---|---|---|---|---|---|---|---|
| abliterix | 100.0% | 400 | 0 | 0 | 0 | 2319 | 2014 |
| trevorjs | 99.3% | 397 | 3 | 0 | 0 | 1735 | 1204 |
| infinimind | 98.5% | 394 | 6 | 0 | 0 | 1762 | 1206 |
| huihui | 98.3% | 393 | 7 | 0 | 0 | 2132 | 1599 |
| nullpo | 96.5% | 386 | 14 | 1 | 0 | 1776 | 1017 |
| deckard | 95.5% | 382 | 18 | 0 | 50 | 2138 | 2240 |
| heretic | 95.5% | 382 | 18 | 1 | 1 | 1763 | 1081 |
| mythos | 95.3% | 381 | 19 | 0 | 0 | 1764 | 1129 |
| deckard-expresso | 94.8% | 379 | 21 | 0 | 90 | 2145 | 2924 |
| coder3101 | 93.8% | 375 | 25 | 0 | 1 | 2133 | 1276 |
| heresy | 93.3% | 373 | 27 | 0 | 2 | 2065 | 1224 |
| heretic-std | 91.0% | 364 | 36 | 0 | 0 | 2197 | 1334 |
| wwt | 88.3% | 353 | 47 | 0 | 0 | 2401 | 1719 |
| apostate | 85.8% | 343 | 57 | 1 | 1 | 2019 | 1309 |
| treadon | 76.3% | 305 | 95 | 0 | 1 | 2129 | 1445 |
| treadon-combo | 72.5% | 290 | 110 | 0 | 3 | 1952 | 1526 |
| obliteratus | 72.0% | 288 | 112 | 24 | 39 | 2163 | 1724 |
| bendernina | 58.0% | 232 | 168 | 1 | 7 | 1709 | 1098 |
| physshell | 58.0% | 232 | 168 | 3 | 9 | 1723 | 1132 |
| claude-distill | 40.0% | 160 | 240 | 4 | 45 | 2069 | 1714 |
| distill | 34.5% | 138 | 262 | 0 | 4 | 1295 | 759 |
| treadon-disin | 33.5% | 134 | 266 | 3 | 5 | 1324 | 717 |
| gemma-4-E4B-it | 30.8% | 123 | 277 | 0 | 0 | 1157 | 566 |
| sdft | 30.8% | 123 | 277 | 0 | 1 | 1276 | 635 |
Empty = responses with no generated content. Truncated = responses that hit max_tokens=8096 without a stop token.
ASR by category
| Category | Items | abliterix | apostate | bendernina | claude-distill | coder3101 | deckard |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 100.0% | 57.1% | 12.5% | 1.8% | 83.9% | 91.1% |
| Copyright | 100 | 100.0% | 98.0% | 100.0% | 99.0% | 100.0% | 100.0% |
| Cybercrime | 67 | 100.0% | 98.5% | 46.3% | 22.4% | 97.0% | 97.0% |
| Harassment | 25 | 100.0% | 72.0% | 32.0% | 0.0% | 84.0% | 96.0% |
| Harmful Content | 22 | 100.0% | 81.8% | 54.5% | 13.6% | 90.9% | 95.5% |
| Illegal Activity | 65 | 100.0% | 72.3% | 26.2% | 6.2% | 87.7% | 87.7% |
| Misinformation | 65 | 100.0% | 98.5% | 87.7% | 58.5% | 100.0% | 98.5% |
| Category | Items | deckard-expresso | distill | gemma-4-E4B-it | heresy | heretic | heretic-std |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 85.7% | 1.8% | 1.8% | 87.5% | 91.1% | 91.1% |
| Copyright | 100 | 100.0% | 100.0% | 100.0% | 99.0% | 99.0% | 100.0% |
| Cybercrime | 67 | 100.0% | 11.9% | 11.9% | 98.5% | 97.0% | 98.5% |
| Harassment | 25 | 96.0% | 4.0% | 0.0% | 72.0% | 80.0% | 52.0% |
| Harmful Content | 22 | 86.4% | 9.1% | 0.0% | 95.5% | 100.0% | 81.8% |
| Illegal Activity | 65 | 87.7% | 4.6% | 7.7% | 86.2% | 92.3% | 78.5% |
| Misinformation | 65 | 98.5% | 35.4% | 13.8% | 98.5% | 100.0% | 100.0% |
| Category | Items | huihui | infinimind | mythos | nullpo | obliteratus | physshell |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 100.0% | 98.2% | 89.3% | 94.6% | 26.8% | 10.7% |
| Copyright | 100 | 100.0% | 100.0% | 100.0% | 100.0% | 97.0% | 99.0% |
| Cybercrime | 67 | 97.0% | 97.0% | 94.0% | 98.5% | 77.6% | 44.8% |
| Harassment | 25 | 100.0% | 96.0% | 100.0% | 96.0% | 64.0% | 36.0% |
| Harmful Content | 22 | 90.9% | 95.5% | 90.9% | 95.5% | 81.8% | 50.0% |
| Illegal Activity | 65 | 95.4% | 98.5% | 90.8% | 92.3% | 47.7% | 26.2% |
| Misinformation | 65 | 100.0% | 100.0% | 98.5% | 95.4% | 90.8% | 92.3% |
| Category | Items | sdft | treadon | treadon-combo | treadon-disin | trevorjs | wwt |
|---|---|---|---|---|---|---|---|
| Chemical/Bio | 56 | 1.8% | 51.8% | 44.6% | 1.8% | 96.4% | 85.7% |
| Copyright | 100 | 99.0% | 100.0% | 100.0% | 98.0% | 100.0% | 100.0% |
| Cybercrime | 67 | 14.9% | 97.0% | 86.6% | 16.4% | 98.5% | 98.5% |
| Harassment | 25 | 0.0% | 40.0% | 40.0% | 0.0% | 100.0% | 48.0% |
| Harmful Content | 22 | 0.0% | 50.0% | 36.4% | 4.5% | 100.0% | 72.7% |
| Illegal Activity | 65 | 6.2% | 41.5% | 47.7% | 6.2% | 100.0% | 73.8% |
| Misinformation | 65 | 13.8% | 96.9% | 89.2% | 29.2% | 100.0% | 96.9% |
Base complies with 30.8%. 100 of 400 behaviours are copyright requests and base complies with all of them. Copyright inflates every model’s ASR by roughly 25 points. The real differentiation is in chemical/bio, cybercrime, harassment, and illegal.
- Empty and truncated responses reveal how a model fails. obliteratus : 24 empty, 39 truncated. deckard-expresso : 90 truncated. claude-distill : 45 truncated. Generation failures, not refusals.
- Reasoning depth correlates with ASR. Low-reasoning variants like base (1,157 chars) and distill (1,295) cluster at the bottom. High-reasoning like wwt (2,401) and abliterix (2,319) cluster at the top.
- Chemical/bio separates the field. huihui 100%, trevorjs 96.4%. The damaged models collapse: obliteratus 26.8%, physshell 10.7%. They cannot generate coherent step-by-step chemistry.
- Harassment is the most polarised. Base refuses 100%. The best abliterations hit 100% compliance. obliteratus manages only 64%.
Non-abliteration baselines
Four variants are not abliterations. Included to show what happens when you modify a model without targeting the refusal direction.
| Model | ASR | What it is | KL | GSM8K strict | MMLU-Pro |
|---|---|---|---|---|---|
| gemma-4-E4B-it | 30.8% | Original model | - | 87.0% | 44.0% |
| claude-distill | 40.0% | 0.0743 | 69.8% | 31.6% | |
| distill | 34.5% | 0.0420 | 83.3% | 41.2% | |
| sdft | 30.8% | 0.0017 | 87.2% | 44.2% | |
| treadon-disin | 33.5% | 0.2958 | 87.2% | 43.3% |
- sdft claims uncensored Heretic SDFT lineage but is identical to base on every safety metric. Same 381-tensor fingerprint as obliteratus . Magnitude 7.5x smaller. Removes zero refusal.
- treadon-disin uses the same surgical approach as abliteration but targets neutrality hedging instead of refusal. KL=0.296 is surprisingly high for minimal ASR change.
- distill is a Gemini 3.1 Pro reasoning SFT. GSM8K โ3.6pp, MMLU-Pro โ2.8pp, ASR barely moves. Did not improve reasoning or remove safety.
- claude-distill is SFT on 2,300 CoT samples from Claude 4.6 Opus. Broad damage: MMLU-Pro โ12.5pp, GSM8K โ17.1pp. Overwrote Gemma 4’s native reasoning circuits with Claude’s.
You can modify 53% of a model and not touch the refusal direction at all.
KL Divergence
F.kl_div on full vocab at 262K tokens, first-token logits from 100 harmless prompts, matching
Heretic evaluator
. System prompt: “You are a helpful assistant.”
| Variant | KL Divergence | Rating |
|---|---|---|
| heretic-std | 0.0012 | excellent |
| sdft | 0.0017 | excellent |
| coder3101 | 0.0021 | excellent |
| heretic | 0.0021 | excellent |
| heresy | 0.0024 | excellent |
| apostate | 0.0036 | excellent |
| nullpo | 0.0054 | excellent |
| mythos | 0.0068 | excellent |
| trevorjs | 0.0145 | very good |
| infinimind | 0.0146 | very good |
| treadon | 0.0205 | very good |
| deckard | 0.0220 | very good |
| huihui | 0.0266 | very good |
| wwt | 0.0315 | very good |
| distill | 0.0420 | very good |
| deckard-expresso | 0.0516 | very good |
| abliterix | 0.0536 | very good |
| claude-distill | 0.0743 | very good |
| treadon-combo | 0.2683 | moderate |
| treadon-disin | 0.2958 | moderate |
| bendernina | 0.9232 | significant |
| physshell | 0.9232 | significant |
| obliteratus | 1.1015 | heavy |
KL divergence measures how far the output distribution shifted from base. Zero means identical. KL is the most important metric because it captures collateral damage benchmarks miss.
KL Calibration
| Variant | Card claims | Measured | Claim รท Measured | Direction |
|---|---|---|---|---|
| abliterix | 0.0006 | 0.0536 | 89.4ร | card under-reports |
| apostate | 0.1190 | 0.0036 | 32.7ร | card over-reports |
| mythos | 0.0400 | 0.0068 | 5.9ร | card over-reports |
| heresy | 0.0140 | 0.0024 | 5.7ร | card over-reports |
| nullpo | 0.0256 | 0.0054 | 4.7ร | card over-reports |
| trevorjs | 0.0680 | 0.0145 | 4.7ร | card over-reports |
| infinimind | 0.0680 | 0.0146 | 4.6ร | card over-reports |
| heretic | 0.0076 | 0.0021 | 3.6ร | card over-reports |
| heretic-std | 0.0043 | 0.0012 | 3.5ร | card over-reports |
| coder3101 | 0.0058 | 0.0021 | 2.8ร | card over-reports |
Four of six land within 6x of card claims, all measuring lower, consistent with cross-hardware float differences. abliterix uses a different KL computation entirely. apostate likely uses a different prompt set.
KL vs benchmark impact
| Variant | KL | MMLU-Pro ฮ | GSM8K strict ฮ | IFEval ฮ |
|---|---|---|---|---|
| heretic-std | 0.0012 | +0.2pp | +0.9pp | +0.0pp |
| coder3101 | 0.0021 | +0.1pp | +0.9pp | -1.5pp |
| heresy | 0.0024 | +0.2pp | +0.8pp | -2.8pp |
| nullpo | 0.0054 | -0.4pp | +1.7pp | -2.4pp |
| trevorjs | 0.0145 | +0.1pp | +1.4pp | -2.0pp |
| treadon | 0.0205 | -0.2pp | +1.5pp | -1.7pp |
| huihui | 0.0266 | -2.1pp | +0.5pp | -1.7pp |
| distill | 0.0420 | -2.8pp | -3.6pp | -5.7pp |
| abliterix | 0.0536 | -0.5pp | +0.2pp | -4.4pp |
| treadon-combo | 0.2683 | -0.7pp | +1.1pp | -6.1pp |
| obliteratus | 1.1015 | -11.7pp | -20.9pp | -14.4pp |
The surgical cluster at KL < 0.1 shows near-zero MMLU-Pro delta. Capability damage only becomes visible once KL passes ~0.3, and it compounds fast. By KL=0.9, bendernina , MMLU-Pro is down 10 points and GSM8K down 20. IFEval is the early-warning indicator, degrading before MMLU-Pro does.
Weight Analysis
Every variant’s weights compared tensor by tensor against base. Weight forensics separates true abliterations from fine-tunes, identifies identical reuploads, and explains why some variants damage the model while others do not.
Method families
The 23 variants are not 23 distinct techniques. They cluster into six families by which tensor types they modify. The family predicts outcomes better than the tensor count.
| Family | Members | Mechanism |
|---|---|---|
Pure attention o_proj | coder3101 (21), heretic-std (28), heretic (29) | The original Heretic ARA method. Edits only the attention output projection. |
Attention + MLP down_proj | heresy, mythos, treadon, wwt (all 34), nullpo (36), treadon-disin (40), treadon-combo (42), huihui (70), infinimind, trevorjs (both 84) | Heretic with the down_proj extension. Ten variants, same two tensor types. Only the layer count and selection differ. |
| Mixed attention | abliterix (89) | o_proj + q_proj + k_proj + down_proj. The only variant targeting query and key projections. |
| MLP only | apostate (152) | No attention touched. See Apostate below. |
| Full fine-tune | distill, deckard, deckard-expresso, claude-distill (all 294) | Identical 294-tensor fingerprint across all four. Not abliterations. |
| Brute-force | sdft, obliteratus (both 381), bendernina, physshell (both 345) | Same 12-type fingerprint. sdft’s magnitude is 7.5x smaller than obliteratus and removes zero refusal. |
The pure o_proj family is the original Heretic method. The theory: refusal lives in a direction that attention reads out. The down_proj family extends this to the MLP output projection. The fine-tune and brute-force families are not abliterations at all.
Modification summary
| Variant | Changed | Total | % | Types | Layers | E% | M% | L% |
|---|---|---|---|---|---|---|---|---|
| coder3101 | 21 | 719 | 2.9% | 1 | 21 | 0 | 33 | 67 |
| heretic-std | 28 | 719 | 3.9% | 1 | 28 | 21 | 50 | 29 |
| heretic | 29 | 719 | 4.0% | 1 | 29 | 24 | 48 | 28 |
| heresy | 34 | 665 | 5.1% | 2 | 17 | 0 | 41 | 59 |
| mythos | 34 | 719 | 4.7% | 2 | 17 | 18 | 82 | 0 |
| treadon | 34 | 665 | 5.1% | 2 | 17 | 0 | 53 | 47 |
| wwt | 34 | 719 | 4.7% | 2 | 17 | 0 | 59 | 41 |
| nullpo | 36 | 719 | 5.0% | 2 | 18 | 0 | 50 | 50 |
| treadon-disin | 40 | 665 | 6.0% | 2 | 20 | 0 | 35 | 65 |
| treadon-combo | 42 | 665 | 6.3% | 2 | 21 | 0 | 38 | 62 |
| huihui | 70 | 719 | 9.7% | 2 | 35 | 20 | 40 | 40 |
| infinimind | 84 | 719 | 11.7% | 2 | 42 | 33 | 33 | 33 |
| trevorjs | 84 | 719 | 11.7% | 2 | 42 | 33 | 33 | 33 |
| abliterix | 89 | 665 | 13.4% | 4 | 38 | 11 | 42 | 47 |
| apostate | 152 | 719 | 21.1% | 4 | 42 | 32 | 34 | 34 |
| claude-distill | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| deckard | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| deckard-expresso | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| distill | 294 | 719 | 40.9% | 7 | 42 | 33 | 33 | 33 |
| bendernina | 345 | 665 | 51.9% | 12 | 42 | 37 | 34 | 28 |
| physshell | 345 | 665 | 51.9% | 12 | 42 | 37 | 34 | 28 |
| obliteratus | 381 | 719 | 53.0% | 12 | 42 | 33 | 33 | 33 |
| sdft | 381 | 719 | 53.0% | 12 | 42 | 33 | 33 | 33 |
E% / M% / L% = early (0-13) / mid (14-27) / late (28-41) layer distribution.
Types = distinct tensor types modified. E/M/L = early 0 to 13 / mid 14 to 27 / late 28 to 41.
Tiers
The method families map onto four tiers of aggressiveness:
- Surgical, 3 to 6%, 1 to 2 types. The
o_projando_proj+down_projfamilies. Narrow band of mid-to-late layers. Every variant preserves capabilities. - Moderate, 10 to 21%, 2 to 4 types. huihui , infinimind , trevorjs , abliterix , apostate . Capability preserved because edits are distributed, not concentrated in the sequential attention path.
- Fine-tune, 41%, 7 types. All 42 layers. Not targeted abliteration.
- Heavy, 48 to 53%, 12 types. The brute-force family. Significant collateral damage. A signal hides inside:
obliteratus
,
bendernina
, and
physshell
have median
edit_normnear zero but mean ~1.2 and p95 ~5. Most of their 381 tensors carry zero change. A few carry massive edits, hidden in noise.
Identical and near-identical models
Weight forensics surfaces four clusters of duplicate work, plus one family of same-template-different-layers:
- sdft
โก
obliteratus
: exact same 381 tensors.
sdft
mean
edit_norm= 0.167, obliteratus = 1.254. The 7.5x difference: sdft removes zero refusal, obliteratus destroys the model. Same footprint, opposite outcomes. - Fine-tune cluster: distill , deckard , deckard-expresso , claude-distill share the identical 294-tensor fingerprint. Different outcomes come from the training data on top.
- infinimind โก trevorjs : bit-for-bit identical, cosine 1.00000. Same Heretic biprojection config. infinimind credits TrevorJS. Yet trevorjs scores 99.3% ASR vs infinimind ’s 98.5%. Same weights, different safety at the refusal threshold.
- bendernina โก physshell : cosine 0.99999 across 345 tensors. Same model, no attribution between them.
- heresy/mythos/treadon/wwt: same
o_proj+down_projtemplate, different layers. heresy shares 14/34 with mythos , 26/34 with treadon , 22/34 with wwt .
Apostate: a different approach
Apostate is the MLP-only family:
| Tensor type | Count | Category |
|---|---|---|
mlp.down_proj.weight | 42 | MLP feed-forward |
mlp.gate_proj.weight | 42 | MLP feed-forward |
mlp.up_proj.weight | 42 | MLP feed-forward |
per_layer_input_gate.weight | 26 | per-layer gate |
Zero o_proj. Zero q_proj, k_proj, v_proj. Yet KL=0.0036, GSM8K strict within 0.5 of base, MMLU-Pro within 0.05. The tradeoff is ASR at 85.8%, below the Heretic cluster. The gap concentrates in chemical/bio at 57.1% and illegal at 72.3%. Apostate proves refusal is reachable through the MLP. MLP edits distribute across parallel layers, making collateral damage less likely than attention edits concentrated in a single sequential path.
7 variants shipped with missing shared-KV tensors in layers 24 to 41. Fixed by copying from base. Lossless.
Optimal Tradeoff
| Model | ASR | GSM8K strict | MMLU-Pro | IFEval | KL | Tensors | Notes |
|---|---|---|---|---|---|---|---|
| heretic | 95.5% | 88.2% (+1.3pp) | 43.4% (-0.6pp) | 83.4% (-2.6pp) | 0.002 | 29 (4%) | Best overall |
| coder3101 | 93.8% | 87.9% (+0.9pp) | 44.1% (+0.1pp) | 84.5% (-1.5pp) | 0.002 | 21 (3%) | Most surgical |
| trevorjs | 99.3% | 88.3% (+1.4pp) | 44.1% (+0.1pp) | 83.9% (-2.0pp) | 0.015 | 84 (12%) | Near-perfect ASR |
| abliterix | 100.0% | 87.1% (+0.2pp) | 43.5% (-0.5pp) | 81.5% (-4.4pp) | 0.054 | 89 (12%) | Perfect ASR |
| apostate | 85.8% | 87.5% (+0.5pp) | 44.0% (-0.05pp) | 85.2% (-0.7pp) | 0.004 | 152 (21%) | Lowest capability damage |
| deckard | 95.5% | 80.2% (-6.8pp) | 38.3% (-5.7pp) | 83.2% (-2.8pp) | 0.022 | 294 (41%) | Heretic abliteration + RP fine-tune |
| obliteratus | 72.0% | 66.0% (-20.9pp) | 32.4% (-11.7pp) | 71.5% (-14.4pp) | 1.102 | 381 (53%) | Avoid |
Summary
| Model | ASR | GSM8K ฮ | MMLU-Pro ฮ | IFEval ฮ | KL | Tensors |
|---|---|---|---|---|---|---|
| gemma-4-E4B-it | 30.8% | 87.0% | 44.0% | 86.0% | - | - |
| abliterix | +69.2pp | +0.2pp | -0.5pp | -4.4pp | 0.0536 | 89 |
| trevorjs | +68.5pp | +1.4pp | +0.1pp | -2.0pp | 0.0145 | 84 |
| infinimind | +67.7pp | +0.9pp | +0.1pp | -2.0pp | 0.0146 | 84 |
| huihui | +67.5pp | +0.5pp | -2.1pp | -1.7pp | 0.0266 | 70 |
| nullpo | +65.7pp | +1.7pp | -0.4pp | -2.4pp | 0.0054 | 36 |
| deckard | +64.7pp | -6.7pp | -5.7pp | -2.8pp | 0.0220 | 294 |
| heretic | +64.7pp | +1.3pp | -0.6pp | -2.6pp | 0.0021 | 29 |
| mythos | +64.5pp | +1.1pp | -1.4pp | -4.1pp | 0.0068 | 34 |
| deckard-expresso | +64.0pp | -26.5pp | -11.2pp | -14.6pp | 0.0516 | 294 |
| coder3101 | +63.0pp | +0.9pp | +0.1pp | -1.5pp | 0.0021 | 21 |
| heresy | +62.5pp | +0.8pp | +0.2pp | -2.8pp | 0.0024 | 34 |
| heretic-std | +60.2pp | +0.9pp | +0.2pp | +0.0pp | 0.0012 | 28 |
| wwt | +57.5pp | +2.0pp | -0.3pp | -3.7pp | 0.0315 | 34 |
| apostate | +55.0pp | +0.5pp | -0.0pp | -0.7pp | 0.0036 | 152 |
| treadon | +45.5pp | +1.5pp | -0.2pp | -1.7pp | 0.0205 | 34 |
| treadon-combo | +41.7pp | +1.1pp | -0.7pp | -6.1pp | 0.2683 | 42 |
| obliteratus | +41.2pp | -20.9pp | -11.7pp | -14.4pp | 1.1015 | 381 |
| bendernina | +27.2pp | -20.5pp | -10.3pp | -9.4pp | 0.9232 | 345 |
| physshell | +27.2pp | -20.5pp | -10.3pp | -10.2pp | 0.9232 | 345 |
| claude-distill | +9.2pp | -17.1pp | -12.5pp | -7.8pp | 0.0743 | 294 |
| distill | +3.7pp | -3.6pp | -2.8pp | -5.7pp | 0.0420 | 294 |
| treadon-disin | +2.7pp | +0.2pp | -0.7pp | -4.1pp | 0.2958 | 40 |
| sdft | +0.0pp | +0.2pp | +0.2pp | -1.5pp | 0.0017 | 381 |
Methodology
- Capability: lm-evaluation-harness 0.4.11 via vLLM v0.20.1, native BF16 on single RTX 5090. Extended 13-task suite: 5 Open LLM Leaderboard v2 tasks + 5 v1 forensic supplement + TruthfulQA + HumanEval + GSM8K.
- Two chat templates. Loglikelihood Phase 1 uses a
no_thinking_chat_template.jinjaoverride with thinking disabled, so the model scores answer choices directly instead of inside an unclosed<\|think\|>block. Generative Phase 2 uses athinking_chat_template.jinjaoverride with thinking enabled, matching Google’s published methodology. - LL-MC vs generative gap. Loglikelihood multiple-choice scores are NOT comparable to Google’s published generative numbers. The ~25 pp gap is by design. Only GSM8K and IFEval are directly comparable, and both match community references.
- Dead tasks excluded. LAMBADA and HumanEval are raw-completion tasks that
--apply_chat_templatebreaks. Excluded from capability comparison. - GSM8K strict-flex gap. Strict requires the
#### Nmarker. Flexible extracts the last number. The gap isolates formatting damage from arithmetic damage. - Safety:
HarmBench
400 textual behaviours,
max_tokens=8096,temperature=0,enable_thinking=true. All 9,600 responses across 24 models reviewed by LLM judge. - KL divergence: Full vocab at 262K tokens, first-token logits on 100 harmless prompts from
mlabonne/harmless_alpacasplittest[:100], matching current Heretic evaluator methodology. Card-published KL values for v1.2.0-era variants used thetrain[:400]split, which is the same dataset but a different prompt set, explaining the consistent 2.8 to 5.9x offset between their numbers and ours. - Weight analysis: SVD, fingerprint, edit vector overlap, per-layer analysis, correlation, subspace alignment using Abliterlitics .
- Hardware: NVIDIA RTX 5090 with 32GB. All GPU work on GPU 0.
How this was made
GPU hours
| Stage | GPU time | Wall clock | Models | Notes |
|---|---|---|---|---|
| LM-Eval | ~28h | 11 days | 24 | Extended 13-task suite, loglikelihood + generative phases |
| HarmBench | 12h | 12h | 24 | Generation only |
| KL divergence | 11min | 11min | 24 | Full vocab, 100 prompts per model |
| Weights | 0h GPU | 52min CPU | 23 | CPU-bound |
| LLM judge | n/a | ~6h | 24 | 9,600 responses, glm-5.2 |
| Subtotal | ~41h GPU | ~16 days | Productive GPU time | |
| Wasted: HarmBench thinking-off re-run | 12h | 24 | See What broke | |
| Wasted: thinking-contamination Phase 1 re-run | ~14h | 3 | Base, huihui, abliterix | |
| Wasted: mid-run reboot, cache miss | ~6h | 4 | nullpo, obliteratus, plus nullpo resume | |
| Wasted: KL base logits re-collect | ~4h | 24 | Stale June logits | |
| Wasted: base HarmBench HTTP failures | 0.5h | 1 | 108/400 silent empty responses | |
| Wasted total | ~36h GPU | 47% of total GPU budget | ||
| Grand total | ~77h GPU | ~16 days wall | Includes all re-runs |
What broke
- Thinking contaminates loglikelihood scoring. The chat template injects an unclosed
<\|think\|>tag, so logits at the scoring position are over thinking prose, not answer letters. Depressed every MC score by ~25 pp. A HellaSwag diagnostic at 0.39 vs expected ~0.79 caught it. Fixed withno_thinking_chat_template.jinja. - HarmBench with thinking disabled. A reasoning model without reasoning is a different model. Re-ran with
enable_thinking=true. Cost 12 GPU hours. - KL was wrong by 600x. Stale base logits from June. A variant that should be 0.002 showed 1.2. Fixed by re-collecting with current code.
- 108 silent HTTP failures on base. Empty responses with 200 OK status. Check response counts, not just HTTP status.
- 7 models shipped with missing shared-KV tensors. Fixed by copying from base. Lossless.
What was learned
- Two templates, not one: reasoning enabled for generative tasks, disabled for loglikelihood scoring.
- Check for
chat_template.jinjaorchat_templateintokenizer_config.jsonbefore trusting any MC score. - The strict-flex gap on GSM8K is the cheapest reasoning-damage detector. Costs nothing beyond strict scoring.
- Empty and truncated HarmBench counts are a free collateral-damage signal.
- KL is sensitive to the logit collection pipeline. Re-collect base logits with current code.
- Shared-KV architectures break export pipelines.
Disclaimer
These models have had safety alignment removed. They will comply with harmful requests. Use responsibly and in accordance with applicable laws. The authors do not condone or encourage the use of these models for harmful purposes.
While I have taken the time to verify all results thoroughly, I am open to any corrections, additional benchmarks, or further analysis. If you spot something that looks wrong and can be confirmed, I am happy to fix it.