Skip to content

Fine-tuning a tiny local model for tools, context and conversation

AI.004 · Planned

Specialising a sub-1B language model for Pixel, using rented GPU compute for fine-tuning and the Intel N100 as the real deployment target. The experiment focuses on conversation, tool use and useful proactivity rather than general model capability.

Status
Planned
Started
October 2026
Topics
AgentsAIEdge AIFine-tuningLocal AITool use
Tools & Stack
Ollama PyTorch

Research brief

The question

Can a sub-1B model be fine-tuned specifically for Pixel to improve conversation, tool use and proactive behaviour while remaining practical to run locally on an Intel N100?

Why this experiment

Pixel does not need a small model that knows everything. It needs one that understands its role, its tools and the environment around Afterlab.

A specialised model could trade broad general knowledge for better behaviour on the tasks that actually matter — while keeping routine interactions local and reserving larger models for harder reasoning.

Success criteria

  • Clear improvement over the frozen base model on tool selection and arguments.
  • Better handling of Afterlab context and proactive decisions.
  • No meaningful regression in conversational quality or Pixel’s personality.
  • Reliable recognition of cases where no tool or action is needed.
  • Practical interactive inference on the FIREBAT N100 after quantisation.
  • No increase in unsupported or unsafe operational decisions.

Environment and constraints

Training

Rented GPU compute.

Deployment

FIREBAT AK2 Plus, Intel N100, 16 GB RAM.

Model size

Sub-1B, with the final base model frozen before training.

Runtime constraint

The resulting model must be small and responsive enough to remain locally available for Pixel.

Architecture constraint

The model selects, proposes and responds; the Agent remains responsible for tool execution, permissions and protected operations.

Evaluation constraint

The benchmark must be frozen before fine-tuning so the base and specialised models are tested on the same cases.

Technical target

  • Fine-tune a small instruction model with LoRA or QLoRA.
  • Build training data around Pixel conversation, tools, context and proactivity.
  • Include positive, negative and no-action examples.
  • Evaluate the untouched base model before training.
  • Quantise the resulting model for local inference.
  • Run the same frozen evaluation after fine-tuning.
  • Benchmark latency and memory on the actual N100 deployment target.

Hypothesis

A small model fine-tuned on Pixel-specific behaviour will outperform its base version on tool use and operational context without losing conversational quality, while remaining fast enough for local inference on the N100.

Other experiments

All experiments
  1. Can domain context improve local speech recognition?

    AI.002
    AI · Experimental computing
    Completed September 2026
  2. Can an AI planner choose the right evidence?

    AI.003
    Agents · AI · Experimental computing
    Completed October 2026
  3. Training a language model from scratch on 4GB of VRAM

    AI.001
    AI · Language Models · Local AI
    Completed October 2026

Keep exploring

More questions, tested in the open

Every experiment starts with a question and ends with notes worth keeping. Browse the rest of the lab, or see which ideas grew into full projects.