Skip to content

AI Workflows

Training my first language model on a GTX 1650

A 4 GB GPU, a small corpus and a lot of things that went wrong in instructive ways. The story around AI.001, with a few examples of what the model actually produced.

Using ChatGPT, or pulling a model through Ollama, makes modern AI feel deceptively simple. You type a name, wait for a download, and you’re talking to something that writes. Training even a tiny model removes that illusion very quickly.

That is why I started AI Lab, a separate repository for research code, datasets, experiment definitions, models and results. Its first experiment, AI.001 — training a language model from scratch, trains a small language model from random weights on a laptop. This post is the story around it, with a few generation examples. The full methodology, corpus breakdown and training curves stay on the experiment page.

I wanted to understand what happens before inference

APIs and pretrained local models are wonderful, and they hide almost everything I wanted to see. What does the dataset actually look like? How much does cleaning matter? What happens between a document and a token? What is a checkpoint, really? What does loss tell me, and what does it quietly not tell me? What does generation look like before a model is useful at all?

I couldn’t answer those by calling something that already worked. The only way I know to find out is to build the unglamorous version and watch it misbehave.

The constraint: 4 GB of VRAM

The machine is a Lenovo IdeaPad Gaming laptop: an Intel i7-10750H, 16 GB of RAM and a GTX 1650 with 4 GB of VRAM, running PyTorch with CUDA under WSL2. It is not a good machine to train on, and I’m not presenting it as one.

That’s the point. With a big GPU you can hide a lot of sloppy decisions behind headroom. With 4 GB you have to know why the context length is what it is, why the batch is split into microbatches, and what the optimizer state costs. I settled on a 256-token context and microbatches of 16 with gradient accumulation, so each update still sees a reasonably large batch. Peak memory for the main run was 1.54 GiB, which is comfortably under the card’s limit. I don’t read that as the card being generous. I read it as the result of designing the pipeline around the limit from the start rather than discovering it halfway through.

Building the corpus

Spanish and English Wikipedia, Python documentation and a handful of other technical documentation sets: collecting text is the easy part. Anyone can download a dump. The interesting part is deciding what is allowed into the training set and what that decision does later.

I chose a controlled corpus over a large unfiltered web dataset, and judged sources on quality, licensing, language balance, duplication risk and how painful they were to preprocess. Cleaning and filtering is where the corpus stops being “everything I could download” and becomes something I made on purpose. I’m not going to dump corpus statistics here; the experiment page has them. What I’d rather say is that those choices came back later, in the evaluation, as behaviour I could point at.

AI.001 workflow: select source text, clean and filter, tokenise, train from random weights, save a resumable checkpoint, then evaluate held-out text and inspect generated output.
The workflow behind AI.001. Checkpoints make recovery possible; held-out loss and generated samples answer different questions.

The first training run changes your perspective

Validation loss at the start of the main run was 9.48. With a vocabulary of 12,000 tokens, a model guessing uniformly would score about 9.39, so the untrained network was effectively guessing. By the end of a single pass over the corpus it had come down to 4.27. Watching a number like that fall is satisfying for exactly as long as it takes to realise it doesn’t say whether the model is any good.

Checkpoints stopped being an abstract idea at around the same time. I had tested resuming from a checkpoint on a short run before committing to the long one, and it printed exactly what I hoped to see:

Resumed from: results/EXP-001/resume-test/latest.pt
Resume step: 10

That turned out to be sensible. Shortly after step 3,540 of the main run, CUDA reported an unknown sticky error and training stopped. It wasn’t an out-of-memory failure and the loss had been behaving. I restarted from the step-3,500 checkpoint, with the optimizer, mixed-precision scaler and random state restored, and it ran to the end.

The more instructive surprise was generation. Sampled with some randomness, later checkpoints produced sentences that looked like documentation, while an early one gave an incoherent mix of WordPress, MDN and Docker material. Under greedy decoding, always picking the most likely next token, the final model fell into loops. Validation loss never suggested that. It measures prediction on real text. Generation feeds the model its own output, and a small model finds a very probable rut and stays in it.

PromptDecodingWhat came out
A WordPress plugin canSampled, temperature 0.8, final checkpoint“A WordPress plugin can be used when the plugin is loaded” and “A plugin contains a PHP file”
Barcelona es una ciudadGreedyRepeated variations of “de la ciudad”
A language model learnsGreedy“the language is the language language…”
def fibonacci(n):Every checkpoint I testedNo valid implementation

Evaluation was more interesting than training

Training is mostly waiting. Evaluation is where the questions start. When I measured the final checkpoint separately by corpus component, technical documentation scored better than general prose, and the spread within technical sources was wide. MDN and WordPress were more predictable to the model than Docker, Python or Git documentation. The source-by-source evaluation and loss chart in AI.001 show the measurements.

I want to be careful with that ordering. It is an observation from one checkpoint of one run, and it is not evidence that the model is good. What it does is raise better questions than the aggregate loss did. Is a source easy because it is well-written, or because it is repetitive? How much of a low loss is the model knowing something and how much is a predictable format? Did I give some sources more room than they deserved? Is there overlap between domains that I haven’t noticed? Loss can’t separate those on its own.

Tiny models are brutally honest

A 22-million-parameter model, trained for about one pass, has nowhere to hide. It learned how technical documentation looks far better than what it means. It writes convincing headings and parameter lists for APIs that don’t exist, drifts between domains within a paragraph, and never produced a working def fibonacci(n): at any checkpoint I tried.

That is useful for learning precisely because it is bad. A strong model smooths over its mechanics. A weak one shows them: capacity, data mix, repetition, the gap between predicting text and knowing things. This is research, not product development, and I try to keep that straight in my own head.

What I actually got from it

Not a competitive model, and not something I’d ask a question. What I got is a research workflow: repeatable corpus tooling, training scripts that can pause and resume, evaluation tools that report by source, and a much better mental model of how a language model is trained and where it goes wrong.

Where it goes next

I’m not promising a bigger model. The questions the first run left open are about corpus selection, better evaluation, and changing one decision at a time so I can tell what each change did. The experiment page lists a more focused follow-up, with more source code in the mix and a fixed evaluation I can compare against this one. I’d rather make a small, controlled step that I understand than a large one that I can only describe.

If you want the rest of the reasoning behind working with constrained hardware, Why most of my AI projects start local-first is the broader version.

Keep exploring

The questions behind these notes

Each note starts from something I tested. The experiments hold the method, the evidence and what did not work.