Small Models, Local Questions

A capable AI model is not always the most interesting choice. Smaller local models can support private notebooks, offline tools and specialised tasks with clearer boundaries. Their limits help reveal which problems need broad intelligence and where modest capability offers more useful control.

Small Models, Local Questions
Photo by Erik Fabian / Unsplash

AI demonstrations tend to move towards scale. Larger models answer more kinds of questions, handle more complex instructions and produce more convincing language. When a new capability appears, the natural response is to imagine everything it might absorb.

I am curious about the opposite direction.

What becomes possible when a model is small enough to run on a laptop, a phone or a modest server controlled by the organisation using it? The answer is not simply a cheaper version of a larger service. Local constraints produce different product ideas.

Imagine a private notebook that can group years of personal fragments without uploading them elsewhere. A museum guide could answer questions from a carefully selected collection even when connectivity is poor. A maintenance technician might search equipment manuals on a device inside a restricted environment. A writer could compare drafts during a train journey with no network at all.

These applications benefit from locality, but locality is not a guarantee of privacy. The surrounding software can still collect data, the device can be compromised and model files may contain risks of their own. A trustworthy local tool must make storage, updates and deletion understandable. “Runs on your device” should describe an architecture, not serve as a magic shield.

Small models also fail more visibly. They may follow complex instructions less reliably, lack broad knowledge or generate weaker prose. Those limits can be useful design pressure. Instead of asking a model to act as a universal assistant, a team must define a narrower task and provide relevant context.

For an editorial archive, that task might be suggesting alternative headings. The tool receives the draft, the publication’s tone guidance and a small set of examples. It produces five options, each labelled by the pattern it uses. The writer remains the editor. No claim is made that the model understands the audience or can judge the article’s purpose.

Another experiment could classify a new note into one of four editorial lenses. Because the categories depend on perspective rather than topic, the tool should explain which passages influenced its suggestion and make uncertainty explicit. Disagreement would be more interesting than accuracy alone: it could reveal where the guide is ambiguous or the article holds two competing centres.

Running locally changes performance questions. Response time depends on the reader’s hardware. Downloading a model consumes storage and data. Battery use matters. An application that works beautifully on a recent laptop may exclude people using older or cheaper devices. The possibility should therefore be tested across real conditions, not only the builder’s machine.

There may be hybrid approaches. A small local model can handle routine classification and send a difficult request to a remote service only with clear permission. Retrieval can narrow context before generation. A conventional search or template may replace the model for predictable tasks. Architecture becomes a set of choices about where capability, data and control should live.

Environmental comparisons require care. A smaller model generally demands less computation per use, but local hardware may be inefficient and replicated across many devices. A remote system can share highly utilised resources. Training, download and idle capacity all matter. The responsible choice depends on scale and use, not the emotional appeal of “local.”

Distribution is part of the experiment. A multi-gigabyte download may be trivial on one connection and prohibitive on another. Updates can be offered as smaller changes, scheduled deliberately and shared across devices in an organisation. A local tool should not celebrate independence while assuming abundant bandwidth, storage and technical confidence.

The strongest reason to explore small models is diversity. If every intelligent feature depends on a few remote platforms, experimentation inherits their prices, policies, connectivity and assumptions. Local and open models can give researchers, small organisations and communities more room to adapt tools to their own languages and constraints.

That freedom brings responsibility. A locally deployed model still needs evaluation, security updates and clear ownership. Greater control means there is no distant provider to blame when maintenance stops.

Small models may never match frontier systems across broad tests. They do not need to. A pocketknife is not a failed workshop. The interesting question is whether a bounded capability, close to the person and the data it serves, can support forms of technology that scale has taught us to overlook.

Next note

Building a Small Recommendation Engine in the Open →
Let's collaborate ↗