NewGLM-5.3 is live as Abliterated Large v2.Try it now
GlossaryReviewed 2026-01-24

Refusal vector ablation

How refusal vector ablation removes refusal behavior while preserving core model capability.

Refusal vector ablation removes the refusal direction from hidden states.

It is the core operation behind abliteration.

Definition

Refusal vector ablation

Refusal vector ablation is the process of subtracting a learned refusal direction from a model's hidden states to reduce refusals without retraining the entire model.

Why it matters
  • Provides a way to reduce refusals and evaluate the resulting model.
  • Can be applied during inference or through weight edits, depending on the method.
  • Lets teams tune refusal behavior with transparent, testable edits.
How it works
  1. 01Learn a refusal direction from hidden-state examples.
  2. 02Choose which layers to apply the ablation.
  3. 03Subtract the projection onto the refusal vector at those layers.
  4. 04Validate with benchmarks and refusal-rate checks.
Ablation formula
h_ablit = h - (h · r_hat) r_hat
FAQ

Frequently asked questions.

Is refusal vector ablation the same as fine-tuning?

No. It does not require gradient-based training; the edit can be applied to activations or weights.

Can I reverse the ablation?

A runtime edit can be removed. For a weight edit, restore the source weights.