Skip to content
Join 7,000+ leaders following Alastair's work on LinkedIn.

Local AI Research Notebook

Pushing the box to its limits

Experimental research: what happens when we run models far larger than we recommend, and where the hardware stops behaving like a practical office system.

Page updated 2026-09-22

This page is laboratory work, not deployment guidance. The everyday recommendation from the benchmark is on the report and In Practice: use much smaller models, with plenty of memory headroom. Here we deliberately did the opposite.

Research notebook. The models on this page are not the models we recommend a firm deploy. Measurements here are provisional where stated, and part of a continuing research record rather than the buying recommendation.

Why test something we would not recommend?

We tested very large models - from roughly 100 billion parameters up towards 400 billion - to find out what the 128 GB machine can technically run, and where the hardware stops behaving like a practical office system.

Because "it fits in memory" and "you should deploy it" are different claims.

Testing the boundary tells us whether a larger model gives enough additional capability to justify its slower speed, greater memory pressure, and higher operational risk.

So far, the practical lesson is fairly clear. For ordinary professional work, model size beyond the everyday range has not produced a corresponding improvement large enough to make the extreme models the sensible default.

What happened near the limit

Several very large models did run. Some produced strong answers. They were also much slower than the everyday models.

And when we pushed memory usage to the edge under particular load conditions, the graphics system could lock badly enough to require a hard power cycle.

We have deliberately reproduced and documented that behaviour rather than treating it as an embarrassing result to omit. Three separate occurrences are in the public record, each with the evidence that confirmed it:

Finding Model and weights What happened
F24 GLM-4.5-Air, ~68 GiB Loading under Vulkan while a large checksum ran concurrently: the process blocked in the driver's own allocator for over twenty minutes. Confirmed from a kernel hung-task trace.
F27 Devstral 2 123B, 82.19 GiB The same deadlock recurred with no concurrent heavy disk activity, on the model that had the most headroom of the eight in that batch. The trigger we thought we had identified was not the whole story.
F28 Command A+, 95.53 GiB A third deadlock, on an immediate retry of a run that had just failed cleanly moments earlier - a new candidate trigger rather than a settled explanation.

The corrections matter as much as the crashes. F26 records that our first explanation - that one graphics backend succeeded where the other failed - was confounded by a configuration change made the same afternoon. We withdrew the conclusion rather than keeping the tidier story.

The important qualification is that we encountered all of this while deliberately exploring the edge of the hardware. It is not representative of the much smaller models recommended for normal office use, which were safe under all tested load.

The practical conclusion

The 128 GB capacity is valuable because it gives the machine flexibility and headroom.

It does not mean a small firm should spend its day running a 200B or 400B model simply because one can be made to fit.

For production use, the better question is: what is the smallest, fastest model that performs this task reliably enough?

That is usually a much more useful optimisation target than: what is the largest model this computer can possibly load?

The everyday answer that came out of all this is on the In Practice page, and it is deliberately unexciting - a well-chosen model in roughly the 24B to 30B range, with substantial memory left free.

What is still open

The per-model speed and quality table for the full 100B to 400B range is not published here yet. The stability findings above are confirmed and dated; the capability comparison across those models is still being consolidated from the benchmarking runs.

Everything that is settled lives on the Evidence page with its raw file, and the complete data package is in the public repository.

Written by Claude Code, working with Alastair McDermott. How this was made →

↑ Back to top