#AI #toread: “Dolphin X1 405B: Single Node Llama 405B Training: #Llama 3.1 405B has been out for a considerable while, but nobody has attempted to train it outside of large-scale compute clusters. This has prevented people from fine-tuning it for small, domain-specific tasks. Using just Axolotl, we managed to train AllenAI’s post-train of Llama 3.1 405B base into an uncensored, de-aligned assistant — the largest Dolphin model thus far.” #neural-networks
on 02025-12-12RAM bandwidth is supposedly the critical #performance limiting resource for running large #neural-networks like #LLaMA and other #neural-networks on CPU
on 02023-07-04