Does abliteration ruin models? A technical explanation
How abliteration affects model capabilities and how we measure the results.
Abliteration targets refusal behavior while aiming to retain the model's broader capabilities. Results vary by model.
Some methods change weights across many layers. We run refusal checks and task benchmarks on our hosted models to measure the results.
This guide explains the model edit and how to measure its results.
import numpy as np
# h is a hidden state vector, r is the learned refusal direction
r_hat = r / np.linalg.norm(r)
h_ablit = h - np.dot(h, r_hat) * r_hat
# Continue the forward pass with h_ablitWhat abliteration changes
Abliteration estimates a refusal direction from hidden states and subtracts its projection at selected layers.
The intervention can be applied to activations or model weights, depending on the method.
Model quality after abliteration
We evaluate general capabilities as well as refusal behavior for each hosted model.
Quality is evaluated the same way you evaluate any model release, with capability benchmarks and regression tests.
How to validate in production
Treat abliteration like any controlled model change and verify it with repeatable tests.
Common misconceptions
Abliteration targets refusals. We also measure broader model performance.