Stop Thinking of LLMs as Next-Token Predictors
gmcgoldr's blog
View My GitHub Profile
Strictly speaking, the statement “LLMs are next-token predictors” isn’t wrong, but it’s incomplete. It’s a fine zeroth-order approximation, and it is grounded in something real: transformer-based language models emit tokens autoregressively:
while not done:
tokens.append(model.sample_next_token(tokens))
This certainly has the shape of something you might call a next-token predictor. During pre-training, the model repeatedly takes some prior tokens, looks at...
Read more at gmcgoldr.github.io