VFD Group
Group Technology,
Products & Platforms

A small bright object outperforming a much larger dim one
Drop 02 of 5 · Tuesday

The models are getting smaller, and the small ones are winning.

If the first drop was about where intelligence runs, this one is about how big it has to be. The industry spent three years insisting the answer was bigger. The evidence has been quietly going the other way, and the implication for a group like ours is direct.

540B PaLM, the teacher ANLI score: lower TAUGHT 770M T5, the student ANLI score: higher
Seven hundred times smaller. Trained on less data. Scored higher.
Google Research, distilling step-by-step. The smaller model scored higher on the ANLI reasoning benchmark. Diagram is illustrative, not to scale.

Google Research trained a 770 million parameter model that outperformed a 540 billion parameter model on a reasoning benchmark, using less data than the standard method requires. Seven hundred times smaller. What changed was not the size of the student but the quality of the teaching.

NVIDIA, a company with every commercial reason to argue the opposite, published a paper last year making the case that agent systems are mostly narrow, repetitive, tightly specified calls, and that small models are sufficiently capable, better suited and necessarily cheaper for them. Note who is saying it, and what it costs them to say it.

This is the part that matters for us. If advantage no longer comes from model size, it comes from what the model is taught on. Fifteen years of credit decisions, investment calls, customer conversations, reconciliations and recoveries. Nobody else can buy that. It is the one input into an AI system that is genuinely, unambiguously ours.

The memorandum says advantage will not come from access to models the whole market can buy. The technical literature agrees, and points at the teaching material instead.

LLM vs. SLM vs. FM: Choosing the Right AI ModelWatch

LLM vs. SLM vs. FM: Choosing the Right AI Model

IBM Technology
Eight minutes, no jargon, and the clearest explanation of when a smaller model is the correct engineering choice rather than the cheap one.