“Local-first” can sound like a position on cloud versus self-hosting. It isn’t one for me. When I start a personal AI project, I usually want the first working loop to run on hardware I control, so that I can see what the thing is actually doing. Once I understand the loop, I can decide which parts deserve a cloud service. That order is the whole idea.
APIs make starting incredibly easy
I use hosted models, and they are excellent. If I want to find out whether an idea works at all, an API call is the fastest honest test there is. I’m not going to pretend otherwise, and this isn’t an ideological stand. Even Afterlab, which is mostly self-hosted, leans on external services: Supabase for authentication and data, a separate server for the web panel, Cloudflare in front of it.
Here is roughly how that plays out inside Afterlab:
| Piece | Where it runs |
|---|---|
| Core services and the Agent | On the mini PC, under my control |
| Web panel | On a separate server, behind Cloudflare Access |
| Authentication and application data | Supabase, hosted |
| Wake word and speech transcription for Pixel | Local |
| Models used by the planner benchmark | Several configured models behind one endpoint, at least one of them hosted |
Privacy is a genuine benefit of running things locally, and I’m not dismissing it. It just isn’t the reason I keep coming back to this pattern. The reason is visibility.
But local changes what you can see
When a model runs on your own machine, the constraints stop being someone else’s problem. You can see inference time, memory use, the cost of loading a model, what moves between the CPU and the GPU, which quantisation you can afford, and which parts of the system quietly depend on the network. A hosted API gives you an answer and a latency figure. A local model gives you the whole shape of the cost.
Local speech recognition for Pixel is a good example. I wanted it running on a Radeon RX 590 I already owned, so the choice of model size was a real trade-off rather than a setting. I compared a small and a larger model, and then tested giving the small one the vocabulary of the lab, because names like Meteolab and Docker mattered more than fluent Spanish. The result was modest, a handful of recordings from one speaker, and the experiment page says so. But I learned things about that deployment that I would not have learned from an API: the domain-context speech recognition test in AI.002 has the details.
Local is especially useful when software meets hardware
This is where the pattern pays off most. The devices in Afterlab use MQTT to reach the lab without a trip through the public internet. Local reads can inspect internal services and notice when a node stops reporting. The infrastructure map in the Kubernetes note shows where those responsibilities sit.
On the CrowPanel, the board listens and the Agent on the mini PC thinks. The device streams short audio frames, the wake word and transcription run locally, and the answer comes back to the speaker. A cloud model can still take part in that loop, but it doesn’t have to own it.

Small models force better architecture
If you can’t assume unlimited intelligence, you get more deliberate about what each component is allowed to do. The Agent selects state and runs tools; the model gets a bounded reasoning task. Simple questions need a fast path, and rejected plans need a fallback. The memory note explains why evidence and current state stay outside the model.
The Composite Planner in Pixel is built this way. Its job is to decide whether a question needs one source, several, or a plain direct answer, instead of blindly querying everything. It is constrained on purpose: a short allowlist of read-only sources, a maximum of three calls, a fallback when the plan is rejected, and a benchmark before anything is switched on more widely. Some simple queries never reach it, because a deterministic fast path is simply the better engineering.
The benchmark is a good illustration of why. Two tests, two different question sets, so these aren’t a before and after:
| Test | Valid JSON | Correct source choice | p95 latency |
|---|---|---|---|
| Baseline, 25 questions (Gemma, two passes) | 50/50 | 46/50 | 1.035 s |
| Extended catalogue with operational history, 30 questions | 30/30 | 19/30 | 0.66 s |
With the wider catalogue the planner kept returning valid JSON and still chose the right sources in only 19 of 30 cases. I treated that as a stop, kept operational history out of the planner, and wrote a closed deterministic composition layer for it instead. That layer’s own live checkpoint is still pending. The source-selection benchmark in AI.003 has the full results. The experiment benchmarks one deployment path, the models I had configured behind the same endpoint, rather than local models in general, and it says so.
Local-first doesn’t mean local-only
This is the part I most want to get right. Hybrid is the normal case, not a compromise.
I use external models when capability matters more than latency or control, when the hardware can’t reasonably run the model I need, or when occasional high-quality reasoning is more useful than running something permanently on my own machine. The planner benchmark above went through a single endpoint with several configured models, and one provider rejected requests outright, so a hosted model is clearly part of the picture. The point isn’t to avoid the cloud. It’s to put it where it earns its place, after I know what the local loop does.
Hardware limitations can improve the design
A 4 GB GPU and a tiny ESP32 are both limits, and so is any small local model or CPU-only machine. Each one pushes me to separate things that a single big model would happily blur together: reasoning, state, tools, storage and interfaces. The microcontroller doesn’t reason. The Agent owns state and tools. A model reads selected context. The panel just renders.
The AI.001 training run is the extreme version of this: a 22M-parameter model that is useless as an assistant and very good at showing where the mechanics are. The Cardputer-based Afterlab ADV is the opposite end. Its portable core compiles and runs on a normal computer, and it queues notes on a microSD card when the lab can’t be reached. Neither design would exist if I had assumed there was always a large model and a fast connection on the other side.
Why I keep coming back to this pattern
It gives me a system I can inspect, change and understand. When something goes wrong, there’s a finite number of places to look, and most of them are mine. The cloud becomes an extension: a capable collaborator for the jobs that need it, rather than the foundation that everything else sits on.
I still don’t know where the line between “run it here” and “send it out” should sit for every case, and I expect it to keep moving as models and hardware change. Starting local is what lets me notice when it should. The system this all lives in is on the Afterlab project page.