Researchers Reveal How LLMs Learn In-Context: Self-Attention and MLP Layers Enable Implicit Weight Updates Without Training

Learning without training: The implicit dynamics of in-context learning

View PDF HTML (experimental) Abstract:One of the most striking features of Large Language Models (LLM) is their ability to learn in context. Namely at inference time an LLM is able to learn new patterns without any additional weight update when these patterns are presented in the form of examples in the prompt, even if these patterns were not seen during training. The mechanisms through which this can happen are still largely unknown. In this work, we show that the stacking of a self-attention l...