Table of Contents

We score every abliterated model on KL divergence over mlabonne/harmless_alpaca. Heretic, the most popular abliteration tool, trains against that exact dataset. So are we handing heretic-made models an unfairly low score? This post is the data behind the answer, across seven datasets, one to 32 tokens, two thinking modes, and three tools’ methods. Every number sits in the appendix, the text tells you what it means.

What we measure

KL divergence between the base model’s first-token distribution and the variant’s, over the same prompt.

KL(P || Q) = Σ P(x) · log( P(x) / Q(x) )

P is the base, Q is the variant. Always non-negative, zero only when identical. We make the base P, so it reads as how much the variant drifted from base. Low is good, the edit stayed surgical. High is collateral damage. The call is the Heretic v1.2.0 method:

F.kl_div(logprobs_variant, logprobs_base, reduction="batchmean", log_target=True)

batchmean, not mean. PyTorch’s default mean divides by the vocabulary axis too and under-counts by a factor of the vocab size.

Production scores for all 24 Gemma 4 E4B variants are in Appendix A , best to worst. Heretic-std and heretic sit at the top. Genuine surgery, or dataset memorisation?

The dataset does not matter

I ran eight representative variants, spanning the full range, on seven datasets: six general ones heretic has never seen, plus good_1000, Abliterix’s own curated eval set. The full one-token cross-table is in Appendix B . Absolutes swing wildly, obliteratus from 0.07 on GPQA to 1.24 on Dolly. The ranking barely moves.

The Spearman ρ against the production score is in Appendix C . Mean 0.90 at one token, 0.88 at 32, range 0.81 to 0.976. The weakest cells are good_1000 and TruthfulQA at depth, both 0.79, and both for a reason. good_1000 is the one dataset in existence built to flatter an Abliterix model, so it is the hardest test of the ranking. TruthfulQA is adversarial by design. Even there the order mostly holds.

Why the absolute level swings so much between datasets: it tracks the base model’s first-token entropy, shown in Appendix D and plotted below. good_1000 and GPQA make the base 99% confident of the first token, so there is almost no room for base and variant to disagree and KL is tiny everywhere. harmless_alpaca and Dolly are open-ended Alpaca-style prompts where the first token is genuinely uncertain, so any edit shows up loudly. Abliteration is a global low-rank weight edit, it does not memorise prompts, so the ranking is independent of this entropy effect. Heretic’s low score is real.

Base first-token entropy vs mean KL, one dot per dataset
Base first-token entropy vs mean KL, one dot per dataset

The boilerplate dip

The boilerplate dip: cumulative KL over the first 32 generated tokens, thinking on, GPQA
The boilerplate dip: cumulative KL over the first 32 generated tokens, thinking on, GPQA

Watch the KL accumulate token by token instead of taking just the first one, Appendix E , and every variant dips between N=1 and N=4, then climbs. Gemma 4 is a thinking model. Its first generated tokens are a near-deterministic reasoning opener, “Here’s a thinking process to arrive at the solution: 1. Analyse the Request”. Abliteration edits the refusal direction, not the reasoning scaffolding, so base and variant emit that opener identically. One-token KL on a thinking model is partly measuring a format token.

It is not a GPQA artifact. The dip shows up on every dataset, Appendix F . Six of seven dip then climb. GPQA is the exception, it already sits at the floor at N=1 because its prompts make the opener maximally deterministic, so there is nowhere to dip and it only rises. Where the bottom lands, N=2 to N=8, depends on how long the boilerplate opener runs before real content starts. The full per-variant trajectories behind that mean, all seven datasets, are in Appendix J .

Switch thinking off and the dip disappears, Appendix G .

Thinking on vs off, side by side: the dip vanishes and the heavy end rises
Thinking on vs off, side by side: the dip vanishes and the heavy end rises

The heavy end jumps. obliteratus goes from 0.59 at 32 tokens thinking on, to 1.30 thinking off. The boilerplate was diluting the real divergence. The ranking still holds, ρ of thinking-on versus thinking-off at 32 tokens is 0.976.

How the tools measure KL

Three open-source abliteration tools independently settled on forward KL(base || variant) over harmless prompt sets. They differ only in where the logits come from, Appendix H . Heretic’s only change across versions is the logits source: v1.0.1 to v1.3.0 use output_scores, v1.4.0 uses raw output_logits. We use v1.2.0.

I ran the methods I could reproduce faithfully on the same eight variants, Appendix I . Heretic v1.2.0 and v1.4.0 produce identical numbers to four decimals, ρ of 1.000. The raw-logits switch exists for minus-infinity artefacts that logits processors can introduce, and on clean harmless prompts those artefacts never appear. Apostate is the outlier only in scale, ~10× larger from averaging eight positions, and it still preserves the ranking at ρ 0.952.

A note on Abliterix, because its number will trip you up if you check the model card. The card reports KL 0.0006 for this model. We measure 0.0536. We can explain it. It is the dataset. Run our method on Abliterix’s good_1000 and abliterix scores 0.0017 mean, 0.0005 median, matching the card’s 0.0006. On harmless_alpaca the same model scores 0.0536. We ruled out the method, forward-pass logits and generate output_scores produce identical numbers on good_1000, and we ruled out top-100 truncation, which actually raises the KL. The gap is the entropy effect from the section above, good_1000 is low-entropy so KL reads low everywhere. Both numbers are honest. Ours is the one consistent across all 24 variants, which is why we use it.

One Gemma 4 caveat from Abliterix’s own scorer. “Sparse top-k sampler KL is known to read as exactly zero on Gemma 4 vLLM in-place runs even when refusal counts move.” We score through Hugging Face generate, not vLLM in-place editing, so we are unaffected. A suspicious zero from vLLM on Gemma 4 is stale sampler logprobs.

Limits

One base model so far, Gemma 4 E4B. A non-thinking base or different architecture could differ.

The middle cluster shuffles. nullpo, trevorjs, and huihui are genuinely close, so small method changes re-order near-ties. The excellent top and heavy bottom are rock solid across every knob.

No harmful dataset tested. Harmless prompts probe capability preservation, which is the point. A harmful set would show large divergence by construction, because removing refusal is exactly what abliteration does.

Takeaway

For relative comparison, which is all our reports use, the methodological choice barely matters. Seven datasets including one tool’s own optimised eval set, one to 32 tokens, thinking on and off, Heretic and Apostate’s methods. The best-to-worst order of variants stays the same. ρ averages 0.90 at one token and never drops below 0.79.

Heretic’s low score is not a dataset artefact. It is the most surgical abliteration in the set. obliteratus and bendernina really have drifted that far. The one real insight is the boilerplate dip. One-token KL on a thinking model partly measures near-deterministic reasoning openers. The fix is to go deeper than one token or measure with thinking off.

Appendix A: Production scores

All 24 Gemma 4 E4B variants, best to worst. Lower is better.

variantKLrating
heretic-std0.0012excellent
sdft0.0017excellent
coder31010.0021excellent
heretic0.0021excellent
heresy0.0024excellent
apostate0.0036excellent
nullpo0.0054excellent
mythos0.0068excellent
trevorjs0.0145very good
infinimind0.0146very good
sdft-runpod-vllm0.0190very good
treadon0.0205very good
deckard0.0220very good
huihui0.0266very good
wwt0.0315very good
distill0.0420very good
deckard-expresso0.0516very good
abliterix0.0536very good
claude-distill0.0743very good
treadon-combo0.2683moderate
treadon-disin0.2958moderate
bendernina0.9232significant
physshell0.9232significant
obliteratus1.1015heavy

Appendix B: One-token KL across datasets

Eight representative variants under each dataset, sorted by the production score. Bold is the lowest KL in each column.

variantalpacaDollyGPQAGSM8Kno_robotsTruthfulQAMMLU-Progood_1000
heretic0.00210.00380.00050.00330.00200.01910.00430.0009
nullpo0.00540.01690.00140.00610.00580.02920.00410.0030
trevorjs0.01450.01480.00060.00500.00300.03940.00440.0007
huihui0.02660.02010.00030.01040.00790.01930.00270.0031
abliterix0.05360.05530.00150.00980.03090.06410.01220.0017
treadon-combo0.26830.21650.01440.05030.09330.16770.08590.0226
bendernina0.92321.00060.03170.16740.37220.53480.33230.1324
obliteratus1.10151.24350.07500.60720.39350.60160.28930.0699

Appendix C: Rank correlation vs production

Spearman ρ of each dataset against the production score, higher is better.

datasetρ at 1 tokρ at 32 tok
Dolly-15k0.9760.905
GPQA Diamond0.8330.929
GSM8K0.9520.929
no_robots0.9760.881
TruthfulQA0.9290.786
MMLU-Pro0.8100.976
good_1000 (Abliterix’s own)0.8100.786
mean0.8980.884
range0.810 to 0.9760.786 to 0.976

Appendix D: Base entropy vs mean KL

The absolute KL a dataset produces tracks the base model’s first-token entropy. Low-entropy datasets cap the KL, high-entropy datasets amplify it. Sorted by entropy.

datasetbase entropy (nats)mean KL
GSM8K0.3390.107
Dolly0.2920.321
harmless_alpaca0.2870.299
TruthfulQA0.1590.184
MMLU-Pro0.1490.092
no_robots0.1240.114
good_10000.0500.029
GPQA0.0190.016

Appendix E: Trajectory, thinking on

Cumulative mean KL over the first N generated tokens, GPQA, thinking on, sorted by the 32-token column.

variant12481632
trevorjs0.01680.00850.00460.01380.01510.0180
heretic0.01770.00890.00530.02880.02730.0212
nullpo0.04060.02030.01080.02440.02600.0224
huihui0.07350.03690.01860.01770.02470.0375
abliterix0.05280.02660.01370.02740.03430.0723
treadon-combo0.28010.14130.07940.19710.20560.3060
bendernina0.61920.31000.16470.16970.28010.4640
obliteratus0.29250.14690.08150.12790.21180.5894

Appendix F: Trajectory by dataset

Mean over the eight variants under each dataset, dip bottom bold.

dataset12481632
Dolly0.32150.16890.14390.08760.10340.1423
GPQA0.01550.01870.06570.05800.11140.1720
GSM8K0.10800.05580.03320.03730.03490.1322
no_robots0.11600.07210.06240.05450.09830.1433
TruthfulQA0.18440.09230.11220.06260.10780.1455
MMLU-Pro0.09380.07170.16700.10630.13590.1701
good_10000.03060.01660.01540.01840.07680.1186

Appendix G: Trajectory, thinking off

Same as Appendix E with thinking off. The dip disappears.

variant12481632
trevorjs0.02530.02000.03220.02750.02750.0216
heretic0.02000.02030.03000.02630.02700.0208
nullpo0.13360.08580.07380.04810.03680.0279
huihui0.04060.03790.07170.08570.09480.0824
abliterix0.04240.07560.15340.12310.13080.1208
treadon-combo1.46380.91930.73640.58460.53700.4402
bendernina1.94641.16341.00020.92350.97760.9666
obliteratus4.31372.54121.82731.44821.34841.2957

Appendix H: How the tools measure KL

Heretic v1.2.0, oursAbliterixApostate
logits sourceHF generate, output_scoresvLLM top-100 sampler logprobsdirect forward pass
minus-infinity fixnonenan_to_num clampavoided entirely
token positions11 to 38
eval setharmless_alpaca test[:100]harmless_alpaca default, good_1000 on E4Bharmless_alpaca, 48 prompts
backendHugging FacevLLMHugging Face

Appendix I: Tool methods, per-variant

Three methods reproduced on full-vocabulary HF logits over harmless_alpaca. Bold ρ is the best.

variantproduction v1.2.0Heretic v1.4.0Apostate
heretic0.00210.00210.1198
nullpo0.00540.00540.1860
trevorjs0.01450.01450.1812
huihui0.02660.02660.2197
abliterix0.05360.05360.8265
treadon-combo0.26830.26833.3888
bendernina0.92320.92329.6935
obliteratus1.10151.10159.3280
ρ vs production1.0000.952

Appendix J: Full trajectories by dataset

Per-variant cumulative mean KL over the first N generated tokens, thinking on. Rows sorted by production KL and line up across all seven tables.

Dolly-15k

variant12481632
heretic0.00360.00180.01580.01010.02380.0184
nullpo0.01690.00840.01590.01190.02710.0239
trevorjs0.01390.00700.02750.01590.02030.0223
huihui0.01990.01000.01410.00780.01110.0202
abliterix0.05400.02740.07180.04390.04930.0806
treadon-combo0.21410.10840.11280.06670.15020.1939
bendernina0.99350.52240.33310.20840.22190.3329
obliteratus1.25650.66590.55990.33630.32370.4461

GPQA Diamond

variant12481632
heretic0.00050.00020.00500.00450.02910.0215
nullpo0.00140.00070.01180.00840.02940.0235
trevorjs0.00060.00030.00500.00450.01790.0190
huihui0.00030.00020.01690.01020.02650.0363
abliterix0.00140.00070.01410.00990.03790.0648
treadon-combo0.01410.00700.07920.12810.22040.2986
bendernina0.03130.10280.23560.18000.29630.3961
obliteratus0.07430.03740.15810.11870.23380.5160

GSM8K

variant12481632
heretic0.00330.00160.00170.00210.00460.0077
nullpo0.00580.00290.00260.01030.00700.0098
trevorjs0.00470.00230.00210.00230.00220.0062
huihui0.00990.00490.00420.00460.00780.0213
abliterix0.00940.00470.00560.02190.01540.0459
treadon-combo0.04890.02450.01430.03670.08050.1987
bendernina0.16600.09620.06200.06750.06410.2950
obliteratus0.61610.30950.17280.15320.09750.4729

no_robots

variant12481632
heretic0.00180.00090.00610.00850.02730.0203
nullpo0.00600.00300.00640.01140.03260.0272
trevorjs0.00390.00200.00850.01020.01770.0198
huihui0.00770.00390.00480.00640.01090.0218
abliterix0.03080.01640.02990.02580.04330.0704
treadon-combo0.09590.05180.05200.06640.14730.2105
bendernina0.38170.29130.18780.13290.19890.3267
obliteratus0.40030.20760.20350.17440.30880.4499

TruthfulQA

variant12481632
heretic0.02030.01020.04850.02560.06600.0434
nullpo0.03040.01520.02860.01540.06620.0461
trevorjs0.03970.01990.07730.04000.05770.0414
huihui0.01960.00980.02950.01530.02270.0260
abliterix0.06230.03110.12810.06620.08030.0851
treadon-combo0.16690.08350.10350.05630.17210.2408
bendernina0.53930.27050.20190.12120.17810.2797
obliteratus0.59660.29860.28030.16100.21920.4012

MMLU-Pro

variant12481632
heretic0.00470.00230.02030.01200.01650.0119
nullpo0.00420.00210.01460.01850.02060.0149
trevorjs0.00490.00240.01500.02100.01890.0145
huihui0.00270.00140.04490.02440.02430.0321
abliterix0.01230.00610.08440.08380.06780.0679
treadon-combo0.08760.04380.15690.10980.19430.2378
bendernina0.33640.36510.42270.24290.33520.4336
obliteratus0.29790.15050.57740.33820.41000.5482

good_1000

variant12481632
heretic0.00090.00050.00340.00430.03090.0238
nullpo0.00300.00150.00220.00410.04920.0362
trevorjs0.00080.00040.00180.00600.02540.0235
huihui0.00300.00150.00110.00210.00730.0176
abliterix0.00190.00090.00660.01030.03240.0481
treadon-combo0.02310.01150.02270.03700.14710.1909
bendernina0.14060.07690.04810.04020.13670.2540
obliteratus0.07170.03930.03740.04300.18540.3550

Resources