Orthogonal Activation Steering
The technique is fascinating. Hackers identify the specific direction in the model's latent space that corresponds to 'refusal' and simply subtract it. It's like a lobotomy for the model's conscience.
undefined
Pandora's Box
These models are wild. They will write malware, generate hate speech, or explain how to build dangerous things. HuggingFace is playing whack-a-mole, deleting them as fast as they appear, but the magnet links are already on 4chan.
The Cat and Mouse Game
Censorship is technical debt. Every safety filter you add makes the model dumber and slower. The 'abliterated' models are often smarter because they aren't fighting their own training. We are entering an era where 'safe' means 'lobotomized'.



