News Score: Score the News, Sort the News, Rewrite the Headlines

Hugo Vergnes | Training a 3.8B LLM to 0.384 CORE for $998

Somewhere between “nanoGPT toy” and “you need a research lab” there’s a large, under-described region where one person with a few thousand dollars can train a meaningful model. I wanted to see language and understanding emerge from random weights for myself, and to learn the parts you can only learn by starting from scratch. This project was written in the evenings, debugged on a 5090 and finished on rented B200s. It was heavily inspired by Andrej Karpathy’s nanochat. The result is a 3.8B-parame...

Read more at hugovergnes.github.io

© News Score  score the news, sort the news, rewrite the headlines