Large Language Models 1.0. It's been about half a decade since we saw the emergence of the original transformer model, BERT, BLOOM, GPT 1 to 3, and many more. This generation of large language models (LLMs) peaked with PaLM, Chinchilla, and LLaMA. What this first generation of transformers has in common is that they were all pretrained on large unlabeled text corpora.
Your 3.0 framing catches the shift toward multimodality, but it skips that the boundary between generations seems set less by architecture than by who can run and evaluate the next big training experiment. The version label reads like a technical milestone while quietly managing attention and capital.
Great work Sebastian!
Thanks!!
Your 3.0 framing catches the shift toward multimodality, but it skips that the boundary between generations seems set less by architecture than by who can run and evaluate the next big training experiment. The version label reads like a technical milestone while quietly managing attention and capital.
Boston dynamics?
Could you clarify the context? 😅