
The Group Managing Director's memorandum sets the destination: an intelligent enterprise, and eventually an intelligence business. This series is about one question the memorandum leaves open, which is where that intelligence physically runs. It starts with a finding that almost nobody outside the research community has absorbed.
Inside any large AI model there is a small group of neurons that fires on almost every input. Researchers call them hot neurons. The rest sit idle, waking only for particular kinds of question. The split is dramatic, and in late 2023 a team at Shanghai Jiao Tong University did the obvious thing with it: keep the busy neurons resident in graphics memory, push the idle ones onto the ordinary processor, and stop paying to hold the whole model in the most expensive memory you own.
A single consumer graphics card then reached 82 per cent of the token rate of a data-centre A100, and ran up to 11.7 times faster than the standard open-source engine. Not a laboratory curiosity. Working code, openly licensed.
The hardware has moved in the same direction. Apple's own measurements show the M5 returning a first token up to four times faster than the M4, with a 24GB laptop comfortably holding a 30 billion parameter model. Those are the machines our people already carry. No purchase order, no procurement cycle, no vendor.
Most of the machine is doing nothing, most of the time. That idle capacity is the cheapest compute this Group will ever have access to.
Tomorrow: why the models themselves are getting smaller, and why that is the best news we have had all year.