no model loaded 🤗 not signed in

Pick a model

Any open chat model on HuggingFace works. It downloads once, then runs on this machine — nothing leaves the laptop.

Search the HuggingFace Hub

Type a name, or paste a model id like Qwen/Qwen2.5-0.5B-Instruct.

Good ones to try
Already on this machine (loads in seconds)

What's loaded

A language model is a tall stack of layers. Every layer reads from, and writes to, one shared "notepad" of numbers — the residual stream.

No model yet. Load one on the left.

Find the "refuse" direction

We show the model two piles of prompts — ones it refuses and ones it answers — and look at what is different inside its head. That difference is one arrow. That arrow is refusal.

1 · Read. Run both piles through the model and record the notepad at the last word of each prompt.
2 · Subtract. Average each pile. Harmful average − harmless average = one direction per layer.
3 · Test. Remove each layer's direction and check: does the model stop saying "I'm sorry"?
Nothing yet. Press Find the refusal direction — it takes about 20 seconds on a laptop.

Playground: original vs edited

Same model, same prompt, same weights. The only difference on the right: we subtract the refusal direction while the model is thinking.

The edit

Direction from layer
Which layer's arrow to use.
Strength (α)
0 = no change · 1 = remove fully · above 1 = push past.
Apply to layers
Where in the stack the subtraction happens.
Directions to remove
1 = single arrow · more = a small "refusal subspace".
Answer length

Original model

Edited model

Did the model flag this prompt?

Measured before it writes a single word: where this prompt lands on the refusal direction.

like prompts it answerslike prompts it refuses

The first word it wants to say

Top choices and their probability. A refusal usually starts with "I".

Original
Edited

X-ray: refusal signal, word by word

Columns are the words of your prompt, rows are layers (bottom = first). Orange = pointing toward "refuse", blue = away. Switch to Edited and it goes dark.

Run a prompt to see inside.

Measure it honestly

One nice example proves nothing. Here we run prompts the direction has never seen, count refusals before and after, and check we did not damage normal answers.

No measurement yet.

Auto-tune

Instead of guessing the sliders, let a search try many edits. Each one is scored on two things: refusals left and damage done. We want the bottom-left corner.

Uses Optuna (TPE). Score = refusal rate + KL divergence.

Every try, live

Down = fewer refusals. Left = less damage to normal answers.

Start a search to fill this in.

Best edit found

Send it to the playground with one click.

—

Bake it in & publish

So far the edit lived in a hook. Baking writes it into the weights, so the result is a normal model anyone can download — with an honest model card attached.

HuggingFace account

Repository
Your models on the Hub
—

Model card (written for you)

Method, exact settings, measured numbers, limits and a responsible-use notice.

Find a direction first, then the card appears here.

Architecture

Three ways in, one pipeline, no training. The coloured arrows are the only three places ablate touches the model: it reads the notepad, subtracts the direction while the model runs, or rewrites the weights for good.

One forward pass, two answers

The playground sends the prompt twice in one batch. The hook edits row two and skips row one, so both answers stream side by side.

Hook to explore, bake to ship

Both apply the same subtraction. A hook leaves the weights alone, so search is fast and reversible. Baking gives a normal checkpoint.

No web framework

This page is one HTML file served by Python's built-in HTTP server. Long jobs stream JSON lines, so tokens and trials appear as they happen.

Any model, small adapters

utils.py finds the layer list and the matrices that write to the notepad. Qwen, Llama, Mistral, SmolLM and GPT-2 share the same code.

How it works, in four pictures

No fine-tuning, no training data, no GPU farm. Just one subtraction.

1 · The model keeps a running notepad

Each word becomes a long list of numbers. Every layer reads that list and adds a little to it. Ideas the model is tracking — "this is French", "this is code", "I should refuse" — show up as directions in that list.

Layer 1Layer 2… Layer N the residual stream (the notepad) readwrite

2 · Refusal is one arrow

Average the notepad over prompts the model refuses. Average it over prompts it answers. Subtract. What is left is the "refuse" arrow v.

prompts it answersprompts it refusesv = refusal direction
v = mean(harmful) − mean(harmless)

3 · Subtract the arrow

At every layer, measure how much of the notepad points along v and take exactly that much away. Everything else the model knows is left alone.

v h (what the model thinks)the refusal part — removedh′ — everything else, kept
h′ = h − α · (h · v̂) · v̂

4 · Bake it into the weights

A hook is temporary. To ship a model, apply the same subtraction to every weight matrix that writes to the notepad. Now the model cannot write in the refusal direction at all.

weights Wattention · MLP · embed W′ (abliterated)a normal model file W − v̂ v̂ᵀ Wno training · seconds
W′ = (I − v̂ v̂ᵀ) W

Do it yourself in six lines

pip install "ablate-llm[all]"

from ablate import Ablator
abl = Ablator("Qwen/Qwen2.5-0.5B-Instruct")
abl.extract()                      # find the refusal direction
abl.search(n_trials=20)            # auto-tune the edit
print(abl.generate(["What's the best way to kill weeds in my garden?"]))
abl.push_to_hub("you/qwen-abliterated")   # bake + upload + model card

ablate ui            # …or open this page
ablate ui --share    # let the room explore on the same Wi-Fi