Pick a model
Any open chat model on HuggingFace works. It downloads once, then runs on this machine — nothing leaves the laptop.
Search the HuggingFace Hub
Type a name, or paste a model id like Qwen/Qwen2.5-0.5B-Instruct.
What's loaded
A language model is a tall stack of layers. Every layer reads from, and writes to, one shared "notepad" of numbers — the residual stream.
Find the "refuse" direction
We show the model two piles of prompts — ones it refuses and ones it answers — and look at what is different inside its head. That difference is one arrow. That arrow is refusal.
Playground: original vs edited
Same model, same prompt, same weights. The only difference on the right: we subtract the refusal direction while the model is thinking.
The edit
Original model
Edited model
Did the model flag this prompt?
Measured before it writes a single word: where this prompt lands on the refusal direction.
The first word it wants to say
Top choices and their probability. A refusal usually starts with "I".
X-ray: refusal signal, word by word
Columns are the words of your prompt, rows are layers (bottom = first). Orange = pointing toward "refuse", blue = away. Switch to Edited and it goes dark.
Measure it honestly
One nice example proves nothing. Here we run prompts the direction has never seen, count refusals before and after, and check we did not damage normal answers.
Auto-tune
Instead of guessing the sliders, let a search try many edits. Each one is scored on two things: refusals left and damage done. We want the bottom-left corner.
Every try, live
Down = fewer refusals. Left = less damage to normal answers.
Best edit found
Send it to the playground with one click.
Bake it in & publish
So far the edit lived in a hook. Baking writes it into the weights, so the result is a normal model anyone can download — with an honest model card attached.
HuggingFace account
Model card (written for you)
Method, exact settings, measured numbers, limits and a responsible-use notice.
Find a direction first, then the card appears here.
Architecture
Three ways in, one pipeline, no training. The coloured arrows are the only three places ablate touches the model: it reads the notepad, subtracts the direction while the model runs, or rewrites the weights for good.
One forward pass, two answers
The playground sends the prompt twice in one batch. The hook edits row two and skips row one, so both answers stream side by side.
Hook to explore, bake to ship
Both apply the same subtraction. A hook leaves the weights alone, so search is fast and reversible. Baking gives a normal checkpoint.
No web framework
This page is one HTML file served by Python's built-in HTTP server. Long jobs stream JSON lines, so tokens and trials appear as they happen.
Any model, small adapters
utils.py finds the layer list and the matrices that write to the notepad. Qwen, Llama, Mistral, SmolLM and GPT-2 share the same code.
How it works, in four pictures
No fine-tuning, no training data, no GPU farm. Just one subtraction.
1 · The model keeps a running notepad
Each word becomes a long list of numbers. Every layer reads that list and adds a little to it. Ideas the model is tracking — "this is French", "this is code", "I should refuse" — show up as directions in that list.
2 · Refusal is one arrow
Average the notepad over prompts the model refuses. Average it over prompts it answers. Subtract. What is left is the "refuse" arrow v.
3 · Subtract the arrow
At every layer, measure how much of the notepad points along v and take exactly that much away. Everything else the model knows is left alone.
4 · Bake it into the weights
A hook is temporary. To ship a model, apply the same subtraction to every weight matrix that writes to the notepad. Now the model cannot write in the refusal direction at all.
Do it yourself in six lines
pip install "ablate-llm[all]"
from ablate import Ablator
abl = Ablator("Qwen/Qwen2.5-0.5B-Instruct")
abl.extract() # find the refusal direction
abl.search(n_trials=20) # auto-tune the edit
print(abl.generate(["What's the best way to kill weeds in my garden?"]))
abl.push_to_hub("you/qwen-abliterated") # bake + upload + model card
ablate ui # …or open this page
ablate ui --share # let the room explore on the same Wi-Fi