Abliterlitics LLM Abliteration Forensics

We take the same base LLM, compare the different abliteration techniques others have applied, then measure what actually changed using benchmarks, safety evaluation, distribution shift, and weight-level analysis.

Gemma4-E4B Abliteration Benchmarks: 23 Variants Compared

77 GPU-hours of abliteration benchmarks for Gemma4-E4B: 23 variants compared across a mixed v1/v2 13-task suite, HarmBench safety, KL divergence, and weight forensics. Heretic is the best capability-safety tradeoff at 95.5% ASR with KL=0.002; abliterix hits 100% ASR.

~4.5B