Arditi et al. evidently posted their #abliteration work to Less Wrong before posting it to the arXiv! #AI #neural-networks
on 02026-05-17Arditi et al.’s “Refusal in language models is mediated by a single direction” #paper. #AI #neural-networks #abliteration
on 02026-05-17explaining how #abliteration of #neural-networks works, based on Arditi et al.’s “Refusal in language models is mediated by a single direction” paper. #AI
on 02026-05-17