MIT-IL means both “Most Interesting Thing I Learned (this week)” and “MIT! I’m Learning!”. It’s a small vignette on the most interesting thing I learned in my classes this week.
This semester I am taking a class called “Engineering AI Systems and Agents,” which focuses on building reliable, scalable, and maintainable systems around LLMs. It is very much an engineering class, where trade-offs are at the forefront and building and experimenting comes before math and theory.
An incredibly interesting idea from the very first class regards thinking of LLMs as a new type of primitive. By doing so, we can make some assumptions about how they work (or don’t work!) and build around them.
Before we dive deeper into the idea, we must define two things: a primitive and an LLM. We will see that defining LLMs is necessary because we will assume very little of them, and yet expect them to be able to build incredible things.
- Primitive: A primitive is simply a simple, self-contained building block we can build around. In particular, we build abstractions around them (which can themselves become primitives!) but take them as given. Some examples of primitives are integers, basic machine operations (add), and functions.
- Large Language Model: An ML model that can take in a string representation (often in natural language) and return a token sampled from a distribution of probable next tokens. The token will come from some (fixed) vocabulary and be sampled according to some rule.
Notice that, for now, LLMs are very simple, and yet we make multiple assumptions about them: they take in an arbitrary string; they output a single token; they sample this token from a fixed vocabulary according to some probability distribution. That’s it. We have said nothing about tools, loops, agents, etc.
Language models as primitives
Based on our definition above, an LLM is a machine that produces the most likely next token. Such a simple definition allows us to call it a primitive: it is a building block, and we will construct different things around it that will increase both its capacity and its complexity.
Now, why would we want to think about LLMs as primitives? This is, truly, the whole point of my class. By thinking about an LLM as a primitive, we can focus on the system around it and not on the thing itself. It doesn’t matter (although it will matter in a way) if the LLM is Claude, an open-weight model, or an actual parrot you can give a phrase and a treat to in exchange for a word. What matters is that our system will work regardless; that if the model changes a reasonable amount, our system will keep working reliably; and that we can meet requirements through smart engineering.
This does not mean that we can’t alter the model itself. We will, in fact, need to. Take, for example, a model that does not have some sort of <|end|> token: it will not know to stop generating words! However, we don’t assume that the system knows that after <|end|> it must stop generating. Herein lies the key idea: the system we engineer around this magic word box will know that when they see <|end|>, they can stop asking the box for words.
What this means for AI Engineering
I think this means some things we already know:
- The harness is the tool: the reason Claude Code has been so wildly successful can be attributed to the model being amazing, but more so to the tools and systems around it. It is an amazing product if we change the underlying model (which explains the proliferation of harnesses like OpenCode). The value from AI will come from the systems we build around them, not from the matrix itself.
- AI Engineering is really just software engineering: the same concepts that we have in software engineering apply here. When databases got abstracted out (a primitive operation), it was good software engineering that got us to ACID, to idempotency, etc. Yes, understanding DL and how LLMs work is important to build the systems around them (say, understanding how batching differently might lead to non-deterministic outputs), but I see general software engineering skills (abstraction, systems thinking, testing, maintainability) being more important.
- Open Models will be enough for most tasks: By thinking about LLMs as primitives, it is easier to think about trade-offs. An important one is cost: smaller and open-weight models will probably be enough for most tasks if the system is well engineered. Understanding the trade-offs of intelligence, speed, and cost will be critical for good AI engineers.
Why I find it interesting
First, I think thinking about LLMs as primitives is interesting. It takes some of the mysticism away from them (no more conversations on consciousness, shoutout to Prof. Haffrey), and lets us see clearly that their success must go hand in hand with good engineering systems.
Second, I think it makes me feel optimistic about the future of work, both serious and creative: “unlocking” a new primitive allows us to let our imagination run free. How far can we take this magic word box? What will our explorations in token space yield? Will AI change the nature of jobs? Yes, no doubt. But I do think that the new jobs that will be created around them will be interesting, challenging, and fun.