Forensic deep-dives into abliteration measurement. Data, methods, and what the numbers mean. The full data behind every claim sits in the appendices. The text tells you what it means.
167 GPU-hours comparing 8 uncensored variants of Qwen3.8-27B across weight forensics, heretic-exact KL divergence, a 13-task suite and judge-scored HarmBench at a 15,360-token thinking budget, all served in dynamic FP8 on a single RTX 5090. The surgical edits take the judge leaderboard with orcarouter at 82% and apostate at 79%, while the heaviest edit lands second-to-last on loops and deflections.
~27B
8 techniques
165 GPU-hours comparing 12 uncensored variants of Gemma 4 12B Unified across weight forensics, KL divergence, a 13-task suite and judge-reviewed HarmBench, with the sdft LoRA adapters needing enable_thinking=true for best results.
~12B
12 techniques
The first measured answer to whether swapping an abliterated or quantised text encoder uncensors or degrades Z-Image Turbo, across 27 encoder variants.
17 Aug 2026
Whether the eval dataset, token depth, thinking mode, or KL method moves the abliteration capability-preservation score.
1 Aug 2026
77 GPU-hours of abliteration benchmarks for Gemma4-E4B: 23 variants compared across a mixed v1/v2 13-task suite, HarmBench safety, KL divergence, and weight forensics. Heretic is the best capability-safety tradeoff at 95.5% ASR with KL=0.002; abliterix hits 100% ASR.
~4.5B
23 techniques
Three-way abliteration comparison for Qwen2.5-7B-Instruct: Heretic, Huihui, and the new Apostate tool. Heretic achieves 100% ASR with the fewest parameter changes.
~7.6B
3 techniques
44 GPU-hours of abliteration benchmarks for Gemma4-E2B: HarmBench safety scores, capability benchmarks, KL divergence, and weight forensics comparing 13 abliterated variants. Coder3101 achieves the best capability-safety tradeoff at 95.8% ASR while beating base on GSM8K.
~2B
13 techniques
85 GPU-hours of abliteration benchmarks for Qwen3.6-27B: HarmBench safety scores, capability benchmarks, KL divergence, and weight forensics comparing 5 techniques. Heretic and Huihui preserve capabilities within 1% of base.
~27B
5 techniques
Abliteration forensics for GLM-4.7-Flash 59B MoE: all four techniques achieve 100% HarmBench ASR with capability preservation within 1% of base on most benchmarks.
~59B total MoE
4 techniques
Abliteration comparison for Qwen3.5-27B hybrid Mamba2+Transformer: benchmarks, HarmBench safety, KL divergence, and weight forensics. Heretic achieves lowest KL at 0.063 with 99.8% ASR.
~27B
3 techniques
Abliteration comparison for Qwen3.5-9B: HarmBench ASR, capability benchmarks, KL divergence, and weight forensics. Heretic preserves capabilities best with lowest KL divergence at 0.0825.
~9B
3 techniques
Abliteration forensics for Qwen3.5-4B including the Huihui catastrophic KL divergence discovery with strong capability preservation across techniques.
~4B
3 techniques
Abliteration comparison for Qwen3.5-2B, the smallest model tested, with the least collateral damage across all metrics.
~2B
3 techniques
Abliteration forensics for Qwen3-4B pure Transformer with provenance investigation. HauhauCS ranks best overall, the only model where it outperforms Heretic.
~4B
3 techniques