Frame selection is the whole game: engineering notes from making LLMs watch video
A video is not an article about the video
Why feed a model a video at all, when a writeup of the same thing costs way fewer tokens?
Because an article is someone's compression of an event. A human watched it, decided what mattered, threw the rest away. The framing, the timing, the stuff on screen they didn't think was relevant. When an LLM reads the article, it learns inside that author's choices. It can't recover what got cut, and it can't disagree with a selection it never saw.
Give the model ...
Read more at leoaido.com