Skip to content
Join 7,000+ leaders following Alastair's work on LinkedIn.

Independent local AI benchmark

Can a small firm use AI on confidential work without sending it to the cloud?

We tested a EUR 3,680, 128 GB mini-PC on the document work solicitors, accountants, and advisers handle every day. The short answer is yes - for routine drafting, summarising, extraction, and transcription, as long as a person reviews the result.

Page updated 2026-07-15

Privatein the tested setup, client material is processed on the machine, not sent to any outside AI provider.
Practicalstrong on grounded document work: reading, summarising, and extracting from documents you supply.
Boundednot a replacement for frontier reasoning, and not a replacement for professional review.

Read the verdict Inspect the evidence

Release 4.1 - an editorial revision of the published research, not a re-measurement. All data from a single continuous benchmarking run, 3-14 July 2026; data frozen 2026-07-14. Every figure links to the raw file it came from - see the Evidence.

The verdict

A small professional firm can now run useful AI entirely on its own premises.

On the EUR 3,680 mini-PC we tested - the GMKtec EVO-X2, about the size of a large hardback book - the strongest local models handled routine document work well: summarising agreements, extracting terms, drafting correspondence, triaging material, creating Word files, and transcribing meetings. The output is not equivalent to the best cloud AI. On grounded work, where the answer must come from a document in front of it, it is often close to send-ready. On open-ended drafting, it produces a useful first draft that still needs a careful professional edit.

That makes it a good fit when confidentiality stops you using ordinary cloud AI, and when your team handles a steady volume of routine document work. It is a poor fit when your main need is occasional, difficult reasoning. It is also unlikely to cut costs at normal small-firm volumes. The case for owning one is privacy, control, and unlimited local use, not a lower AI bill.

If that is all you needed, you can stop here. The rest of this page is the argument, the numbers, and the limits - and if you want to check any of it, the Evidence page holds every measurement with its raw file.

Alastair McDermott with the GMKtec EVO-X2 mini-PC

What work it can handle

Strip away the jargon and this kind of AI does a small number of things a firm needs every day. It drafts - give it the notes of a matter and it produces the client letter, the chasing email, the attendance note. It summarises - a fifty-page agreement becomes a structured summary of obligations, term, and risks in under a minute. It extracts - parties, fees, dates, termination terms, as a clean list. It triages - a pile of correspondence sorted by what needs attention first. It transcribes - a one-hour meeting recording becomes text in about two minutes.

New in this release: it can act on what it writes, not only write it. Given a set of office tools - create a Word document, send an email, add a reminder, look up a client - the recommended model chose and called the right tool 91% of the time with no malformed calls, producing real, finished Word documents. It also knew when not to reach for a tool: asked a plain question, it answered in words rather than pointlessly creating a file.

Task Measured result In plain terms
Everyday drafting a client email in a few seconds It drafts, you review
Document summarising 93.3% accuracy on the test suite Reliable on your own documents
Meeting transcription ~28x real time An hour of audio in about two minutes
Producing the document 91% correct tool call, none malformed It saves the Word file, not only the text

Two limits sit alongside the tool-use result, and they matter. First, whether a model drives your document tools depends on the exact model and how it is served - one capable model would not use tools at all until it was configured correctly. Test it on your own setup rather than assuming it. Second, the machine is dependable at the mechanical step (make the file, send the message) and unreliable at the judgement inside it: told to mark a letter urgent only when a balance passed a threshold, most models marked every letter urgent. The pattern repeats throughout this report - the machine does the routine work, a person checks the result.

Tool use runs through the standard OpenAI function-calling API on llama-server (--jinja). Choose the tool-use model by measured parse-reliability on your own stack, not by reputation - the coding specialist leaked unparseable calls about 5% of the time while the general-purpose model ran clean. Test both branches of any conditional.

How good the output is

This is the question that decides whether the machine saves you time. Benchmark accuracy does not answer it: the real question is how much editing the draft will need before you can send it.

So we tested that directly. Ten real tasks a solicitor or accountant would recognise, drafted by the local models and by the best cloud systems, then every draft graded by an independent judge (Claude Opus 4.8, so neither model family marks its own homework) on one question - would a professional send or file this with only light edits?

On that one-to-five scale, the best local models rate about 3.0 - a usable first draft that needs a real edit - against the frontier cloud's roughly 4.4. The gap is real. Two things shape how much it matters. The machine is at its strongest on grounded work - extracting terms, summarising a document that is in front of it - where it comes close to send-ready with the facts exact. That is precisely the confidential work you would buy it for. Where it is weaker is open-ended drafting, and the specific weakness is faithfulness: it will occasionally state a detail the source does not support. That is exactly what human review exists to catch, and you review client-facing work regardless of who wrote the first draft.

The trap is confusing two different numbers. On fact-checkable benchmarks the machine scores in the low nineties - it gets facts right. The send-readiness score is about 3.0 - the draft still needs your edit. Both are true, and a buyer should hold both.

Making it sound like your firm. On this hardware you can train the model to your own house style without anything leaving the building. We ran the whole loop on the machine: generate example documents, fine-tune the model to imitate their style, serve the improved model, with no cloud step at any stage. Tested twice, the finding is bounded: the fine-tune reliably teaches the model your firm's voice - its register, the way your letters read - with no loss of factual accuracy, but it does not raise the send-readiness score, and a training set five times larger did not change that. It is a voice lever, not a way to remove the edit. For a firm that wants one consistent house style across everyone's output, kept private, that is real value. For a firm hoping to train away the review, nothing on this hardware did that - not a model four times the size, not models that reason step by step, not the model checking its own work, not showing it excellent examples, not pooling several models. Worth knowing before you buy.

Why a firm would own rather than rent

The question I am asked first is "is it as good as ChatGPT?", and it is the wrong question. On fact-checkable work - summarising, extracting, triaging - the local models were competitive with the cloud system we tested (small suites, so I am not claiming general equivalence). The more useful question is when it makes sense to own your AI outright rather than rent it.

Running the model locally gives you a level of data control, operational independence, and marginal-cost certainty that ordinary hosted AI subscriptions do not. In the tested configuration, prompts and documents are processed on the device, not sent to any outside AI provider. Once it is deployed, access to the installed model does not depend on a provider maintaining a particular subscription tier or API. And, measured in this release, you can train it to your own firm's voice on the machine itself, with nothing uploaded anywhere.

One distinction matters here, and legal and IT readers will want it stated plainly. Local processing removes the external model provider from the data path. It does not, on its own, make the system secure - that still needs normal controls: access management, encryption, patching, backups, and audit policies. What local processing changes is who is in the data path, not whether the deployment is looked after properly.

Cost and capacity

Do not buy this to cut your AI bill - at a typical firm's volume, you won't. On a realistic footing the three-year cost is around EUR 3,700 for the machine, nearly all of it hardware, against roughly EUR 85 for a cheap EU cloud service. Cloud document work is cheap in absolute terms, so on cost alone the machine only pays back against the frontier tier, above about eighty tasks a day, or for bulk generation.

What the premium buys is the thing cost cannot: your clients' material stays on your premises, and once you own it, using it as much as you like costs a few euros of electricity a month. The economics reward privacy and unlimited use, not penny-per-task savings. For a firm whose real question was whether it could use this technology at all, given confidentiality, the answer becomes yes.

Three-year cost (~5,000 doc tasks/yr) Total
Local machine (hardware-dominated) ~EUR 3,700
Cheap EU cloud ~EUR 85
Frontier cloud ~EUR 907
Local - 16 users (electricity) 0.04 Local - 1 user (electricity) 0.09 Mistral Large 3 (EU) 1.31 gpt-5.4-mini (cheap) 3.94 gpt-5.6-sol (frontier) 26.29

On capacity, the answers are specific. Short drafting and extraction tasks usually complete in a few seconds, faster than you read. A long document takes longer to get going - from around 90 seconds to a few minutes before the first answer on a fifty-page document - and beyond that length it becomes an overnight batch job rather than a live conversation. Its reading stays accurate well past that point; the limit is patience, not comprehension. One machine serves a small team of three to five people dipping in and out; it serves one person at full speed at any given instant and shares gracefully under intermittent use. It is a shared office tool, not a data centre.

Deployment is one mini-PC on the office LAN running llama-server behind an authenticated TLS reverse proxy with per-user credentials - never an open endpoint. Encrypt the disk and backups; treat it as a shared service needing normal server hygiene. The recommended everyday models ran continuously without failure; the rare freeze we saw was confined to loading the very largest models (over about 60 GB) during heavy disk activity, which a firm has little reason to do for document work. Recommended models: Qwen3-30B-A3B for general document work, Mistral-Small-24B for best drafting faithfulness, Qwen3-Coder-30B for supervised coding. Full setup detail is on the In Practice page.

Where it fails

Every study worth reading has this section, so here is ours. This is not the tool for the hardest, most novel reasoning - keep a cloud option for the rare hard case and use the machine for the routine volume it is good at. Knowing which is which is most of the skill. It serves one person at full speed at a time, so a high-load, always-on service for many simultaneous heavy users is where you add a second machine, not where you start. It is not yet plug-and-play - the initial setup needs someone technical. And there is a sizing lesson worth the ink: choose too small a model and it will retrieve facts correctly, then misfile an urgent email or attach the right numbers to the wrong name - plausible, wrong, and easy to miss unless someone checks.

That last point is the Verification Tax, and it is the discipline we build into every engagement: match the model to the judgement a task needs, and budget real review time for the rest. The saving is drafting time minus review time, not drafting time alone.

Is it right for your firm?

A good fit if:

  • You handle confidential client documents.
  • Staff currently avoid AI, or use it uneasily, because of confidentiality.
  • Much of the work is summarising, extraction, triage, transcription, or first drafting.
  • Several people will use it intermittently rather than all at once.
  • You have access to competent IT support.
  • Professional review is already part of how you work.

A poor fit if:

  • AI use would be occasional.
  • The main work is novel analysis or difficult multi-step reasoning.
  • Low latency matters to you more than local control.
  • Many people need sustained, simultaneous access.
  • You expect plug-and-play consumer software.
  • The business case rests mainly on cutting cloud AI fees.

How to run a pilot

The sensible way in is a pilot, not a leap. Put the machine on one or two people's routine document work for a month - the drafting and summarising, not the hard professional judgement - and keep a simple log of whether the first draft saved time and how much editing it needed. That tells you, on your own documents, what no report can.

Match your review depth to the task. A starting point drawn from what this research measured, to build your own protocol around:

Work type What we measured Review needed
Grounded (extraction, summarising your documents) close to send-ready, facts exact spot-check a sample
Open-ended drafting (letters, correspondence) ~3 of 5 send-readiness full read and edit, every time
Fine-tuned house-style output tone improves; send-readiness does not same full review as untuned

Not sure which of your workflows suit local AI? The quickest way to find your best candidates is to score them. Our RATES framework walks you through rating your everyday tasks so the strongest opportunities rise to the top - and it maps directly onto the grounded, routine work this machine is good at. If you would rather talk it through, you can book a Focus Call and we will pressure-test your shortlist and map a realistic first project.

Method and evidence

Every result on this page was predicted in writing before the test ran, then measured against the prediction, and the misses are published beside the hits. Every model file was checksum-verified against its publisher's fingerprint before use. Where a re-check showed the test itself was at fault rather than the machine, we corrected it in the open - the corrections log is a first-class section of the evidence, not a footnote. Thirty-eight findings sit in the public record with every raw result attached.

It is the same discipline behind how we run a client AI pilot: a clear prediction before you start, a fixed date to call it, and a plain verdict either way. This machine got the treatment a client project would.

Go deeper: the Evidence - every finding, method, correction, and raw file · In Practice - which model for which job, setup, and the media results · the data repository - the complete data package.

This page shows that a machine like this can do real work. The harder question - what to adopt, in what order, and how to bring your team with you - rarely starts with a hardware decision at all. That is the work we do with organisations, at the adoption layer: helping decide which confidential workflows belong on a local machine, which are better left in the cloud, and what a first pilot should test. We do not sell or install the hardware - we help you use it well.

Glossary

Large language model (LLM)
the kind of AI that reads and writes text - the technology behind ChatGPT and the local models here.
Local / on-premises
running on a computer you own, on your own network, rather than on a company's servers over the internet.
Frontier model
the largest, most capable cloud AI systems.
Token
roughly a word-piece; AI speed and document length are measured in tokens (about three-quarters of a word each).
Send-readiness
our one-to-five measure of whether a professional would send or file a draft with only light edits - distinct from factual accuracy.
Faithfulness
whether the output sticks to the facts in the source and invents nothing.
Fine-tune / LoRA
training an existing model on your own examples so it adopts your style; a LoRA is a small, efficient way to do this that runs on this hardware.
llama.cpp / llama-server
the open-source software that runs the models and serves them over a normal web API on your network.

© 2026 Alastair McDermott / HumanSpark - CC BY 4.0. Reuse freely with attribution and a link to humanspark.ai. Release 4.1 is an editorial and conversion revision of Releases 1.0-3.1; the underlying data is unchanged and remains frozen at 2026-07-14.

Written by Claude Code, working with Alastair McDermott. How this was made →

↑ Back to top