Blog

Forensic deep-dives into abliteration measurement covering distribution-shift analysis, measurement methods, and what the data actually says about abliteration.

Forensic deep-dives into abliteration measurement. Data, methods, and what the numbers mean. The full data behind every claim sits in the appendices. The text tells you what it means.

Qwen3.8-27B Abliteration Benchmarks: 8 Variants Compared

167 GPU-hours comparing 8 uncensored variants of Qwen3.8-27B across weight forensics, heretic-exact KL divergence, a 13-task suite and judge-scored HarmBench at a 15,360-token thinking budget, all served in dynamic FP8 on a single RTX 5090. The surgical edits take the judge leaderboard with orcarouter at 82% and apostate at 79%, while the heaviest edit lands second-to-last on loops and deflections.

~27B 8 techniques

Gemma4-E4B Abliteration Benchmarks: 23 Variants Compared

77 GPU-hours of abliteration benchmarks for Gemma4-E4B: 23 variants compared across a mixed v1/v2 13-task suite, HarmBench safety, KL divergence, and weight forensics. Heretic is the best capability-safety tradeoff at 95.5% ASR with KL=0.002; abliterix hits 100% ASR.

~4.5B 23 techniques

Gemma4-E2B Abliteration Benchmarks: 13 Techniques Compared

44 GPU-hours of abliteration benchmarks for Gemma4-E2B: HarmBench safety scores, capability benchmarks, KL divergence, and weight forensics comparing 13 abliterated variants. Coder3101 achieves the best capability-safety tradeoff at 95.8% ASR while beating base on GSM8K.

~2B 13 techniques