“We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, DeepSeek-Coder-V2 is further pre-trained from an intermediate checkpoint of DeepSeek-V2 with additional 6 trillion tokens. Through this continued pre-training, DeepSeek-Coder-V2 substantially enhances the coding and mathematical reasoning capabilities of DeepSeek-V2, while maintaining comparable performance in general language tasks.” #neural-networks #DeepSeek #AI
on 02026-01-12supposedly #DeepSeek v3.1 beats Claude Opus 4 on some coding benchmark or other and is “open-source”. “671B parameters (37B active per token).” Lots of stock footage in the #video. Optimized for "UE8M0 FP8", which I guess is an 8-bit floating-point format supported by domestic (SMIC?) GPUs like the recently-released Xi Yun C600 and a speculated-to-be-near-future Huawei chip. Apparently GPT-4 slightly beats it on coding benchmarks. #AI #news. Video is 2 hours 42 minutes. #toread
on 02025-11-06#AI #news from WaPo: “AI firm #DeepSeek writes less secure code for groups China disfavors”
on 02025-10-16“#DeepSeek says that the main restriction on their development is lack of compute, and the PRC responds not by helping them get better chips but by advising them to not use the chips that they have, greatly slowing things down at least for a while. (...) It’s also not fair to fully pin this on DeepSeek when they were forced to do a lot of their training this year on #Huawei Ascend chips rather than Nvidia chips.” (The #NVIDIA chips they had been using in the past.) #China #politics #neural-networks #AI #news
on 02025-08-23Gizmodo’s #news summary of the #DeepSeek controversy today
on 02025-01-29#Economist #news summary of the #DeepSeek controversy today
on 02025-01-29Jiayi Pan’s replication of the #DeepSeek #neural-networks training approach through RL finetuning of the Qwen2.5 base model. “Clean, accessible reproduction of DeepSeek R1-Zero”
on 02025-01-29Andrej #Karpathy reports Jiayi Pan (@jiayi_pirate) has replicated the #DeepSeek #neural-networks training approach through RL finetuning of a base model, showing an example of “the CountDown game”
on 02025-01-29