News Score: Score the News, Sort the News, Rewrite the Headlines

>10x More Efficient Pretraining — Magic

Research update on compute-efficient pretraining and scaling to trillion-parameter models.Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. After compounding for … a while …, our pretraining recipe is now >10x more compute-efficient than that of leading open-weight base models. We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200. We continue...

Read more at magic.dev

© News Score  score the news, sort the news, rewrite the headlines