Skip to content

Training a language model from scratch on 4GB of VRAM

AI.001 · Completed

Training a small language model from random initialization to explore how much linguistic structure can emerge under severe consumer-hardware constraints.

Status
Completed
Started
September 2026
Completed
October 2026
Topics
AILanguage ModelsLocal AIMachine Learning
Tools & Stack
CUDA Python PyTorch WSL2

Research brief

The question

Can I train a language model entirely from random initialization on a consumer laptop with only 4 GB of VRAM, and get it to learn measurable language structure?

Why this experiment

Most local AI experiments begin with an existing pretrained model. This experiment goes one layer deeper: starting from random initialization to understand the full training pipeline — dataset construction, tokenization, architecture, optimization, checkpointing, evaluation and generation.

The goal is not to build a competitive general-purpose language model, but to explore how much language structure can emerge under severe compute constraints and to build a reproducible foundation for future AI Lab experiments.

Success criteria

The experiment will be considered successful if:

  • Training remains numerically stable without persistent divergence or out-of-memory failures.
  • Training and validation loss show meaningful improvement over the initial baseline.
  • Held-out evaluation demonstrates learning beyond simple memorization.
  • Generated samples evolve from near-random token sequences toward recognizable lexical, syntactic or structural patterns.
  • Training can be paused and resumed from checkpoints reproducibly.
  • The full pipeline can run within the available 4 GB VRAM budget.

Environment and constraints

Hardware

Lenovo IdeaPad Gaming
NVIDIA GeForce GTX 1650 — 4 GB VRAM
Intel Core i7-10750H
16 GB system RAM

Environment

WSL2 / Linux
Python
PyTorch
NVIDIA CUDA

Constraints

The complete training workflow must remain viable within 4 GB of GPU memory.

AI-001 uses no pretrained model weights, LoRA adapters or external training compute for the core experiment.

Technical target

Train a compact decoder-only Transformer entirely from random initialization using a custom corpus and tokenizer.

Current AI-001 configuration

  • Context length: 256 tokens
  • Training corpus: 62,171,404 tokens
  • Validation corpus: 1,194,089 tokens
  • Microbatch size: 16
  • Gradient accumulation: 4
  • Effective tokens per optimization step: 16,384
  • Mixed-precision GPU training
  • Periodic validation
  • Automatic checkpointing
  • Resumable optimizer and training state

The final architecture, parameter count and training schedule will be recorded from the frozen configuration used for the primary AI-001 run.

Hypothesis

A small decoder-only Transformer trained from scratch on a carefully constructed corpus can learn measurable language structure on a consumer GPU with 4 GB of VRAM, provided that model size, context length, batch strategy and dataset size are adapted to the hardware constraints.

18 Sep 2026Research environment initialized

AI-001 starts as the first formal experiment inside AI Lab.

Initial constraint The complete training workflow must be viable on an NVIDIA GTX 1650 with 4 GB of VRAM.

The objective is to train a small language model entirely from random initialization using consumer hardware, rather than fine-tuning or adapting an existing pretrained model.

A dedicated ai-lab repository was created to keep research code, datasets, experiment definitions, models and results separated from other local AI projects.

Decision Treat hardware limitations as part of the experiment rather than moving training to external compute.

18 Sep 2026GPU training environment validated

The local training environment was prepared under WSL2 with PyTorch and CUDA support.

PyTorch successfully detected the NVIDIA GTX 1650 and GPU execution was validated.

This established that AI-001 could proceed using the laptop GPU without requiring cloud infrastructure.

Result Local CUDA training is operational.

Sep 2026Dataset strategy defined

The first corpus strategy was designed around a relatively small but varied text dataset that could be processed and trained locally.

The goal was not simply to maximize token count, but to provide enough linguistic variety for a small model to learn observable structure.

Sources were evaluated according to:

  • text quality;
  • licensing and accessibility;
  • language balance;
  • duplication risk;
  • preprocessing complexity;
  • usefulness for a small general-purpose language model.

The corpus design includes general prose alongside educational and technical material.

Decision Build a controlled corpus rather than downloading an unrestricted large-scale web dataset.

Sep–Oct 2026Training pipeline implemented

A dedicated AI-001 training pipeline was built around a compact decoder-only Transformer.

The trainer supports:

  • GPU training;
  • gradient accumulation;
  • periodic validation;
  • checkpoint creation;
  • model snapshots;
  • deterministic experiment configuration;
  • interruption and resume from saved training state.

The pipeline was designed specifically around the available memory budget rather than using default configurations intended for larger GPUs.

Key constraint 4 GB VRAM.

03 Oct 2026AI-001 training configuration established

The current AI-001 dataset contains:

MetricValue
Training tokens62,171,404
Validation tokens1,194,089
Context length256
Microbatch size16
Gradient accumulation4
Effective tokens / update16,384

Training runs on the NVIDIA GTX 1650.

This configuration represents the first practical training setup that fits the experiment’s hardware constraints while maintaining a useful effective batch size.

Observation Gradient accumulation allows the effective training batch to be substantially larger than what can physically fit in GPU memory at once.

03 Oct 2026Checkpoint and resume test

Before committing to a long training run, checkpoint recovery was tested independently.

A short AI-001 run was stopped after optimization step 10 and then restarted using the generated latest.pt checkpoint.

The trainer correctly restored the saved state:

Resumed from: results/AI-001/resume-test/latest.pt
Resume step: 10

Training then continued with the expected schedule.

Result Checkpoint resume works correctly.

This is an important requirement for AI-001 because training on laptop hardware makes uninterrupted long-running sessions unreliable.

Decision Checkpoint/resume is considered validated for the main experiment.

03 Oct 2026End-to-end trainer validated

The complete basic training loop has now been exercised successfully:

dataset → batches → forward pass → optimization → validation → checkpoint → resume

At this stage, the experiment has demonstrated that the infrastructure required to train the model from scratch works within the available hardware constraints.

This does not yet validate the main hypothesis.

The remaining question is whether sustained training produces measurable language learning.

Current state AI-001 is ready for longer training and model-quality evaluation.

03 Oct 2026Main training run started

With the training pipeline validated, AI-001 moved from infrastructure testing to the primary learning run.

The model is a compact decoder-only Transformer containing 22,451,712 trainable parameters, initialized entirely from random weights.

ParameterValue
Parameters22,451,712
Vocabulary12,000 tokens
Context length256
Transformer layers10
Model dimension384
Attention heads6
MLP dimension1,536
Microbatch size16
Gradient accumulation4
Effective tokens / update16,384
OptimizerAdamW
Peak learning rate3e-4
Minimum learning rate3e-5
Warmup200 steps
PrecisionFP16 AMP

The training target was 3,795 optimization steps, corresponding to 62,177,280 processed tokens.

This is approximately one complete pass through the 62.17M-token training corpus.

Baseline validation loss 9.4776

From this point forward, the experiment measures whether training from random initialization produces consistent improvement on held-out validation data.

03 Oct 2026Validation loss falls throughout training

Validation was performed regularly during the main run.

StepValidation loss
09.4776
2506.4093
5005.8517
7505.5165
1,0005.2391
1,2505.0322
1,5004.8696
1,7504.7362
2,0004.6319
2,2504.5452
2,5004.4689
2,7504.4067
3,0004.3613
3,2504.3237
3,5004.2948
3,7504.2720
3,7954.2680

Validation loss decreased at every recorded evaluation point, from an initial 9.4776 to a final 4.2680.

This represents a reduction of approximately 55% from the randomly initialized baseline.

There is no sustained validation regression visible during the run. The final checkpoint also produced the lowest validation loss observed during AI-001.

Observation The model continued improving on held-out data throughout the complete training run.

This provides quantitative evidence that training is producing generalizable predictive structure rather than merely reducing loss on the active training batches.

03 Oct 2026Training remains viable within 4 GB VRAM

The complete model and training configuration remained comfortably inside the experiment’s primary hardware constraint.

Peak CUDA memory usage during the run was 1.54 GiB on the NVIDIA GTX 1650 4 GB.

MetricResult
Peak CUDA memory1.54 GiB
Available GPU memory4 GB
Final training loss4.4347
Final validation loss4.2680
Tokens processed62,177,280
Optimization steps3,795

No NaN divergence or GPU out-of-memory failure was observed during optimization.

Result A 22.45M-parameter Transformer can be trained from scratch using this configuration on the available GTX 1650.

The hardware constraint therefore influenced the design of the experiment, but did not prevent completion of the training run.

03 Oct 2026CUDA interruption and recovery

Shortly after step 3,540, training was interrupted by a CUDA runtime failure reporting:

CUDA_ERROR_UNKNOWN / Sticky error detected

The interruption was not preceded by an out-of-memory condition, NaN loss or sustained validation regression.

A valid checkpoint was already available from step 3,500.

The run was restarted from this checkpoint, restoring:

  • model weights;
  • optimizer state;
  • AMP scaler state;
  • random number generator state;
  • training progress.

Training resumed successfully from step 3,500 and continued normally through the planned final step.

Result Checkpoint recovery worked under a real training interruption, not only during the earlier controlled resume test.

The incident did not prevent the resumed run from surpassing the previous best validation loss.

03 Oct 2026Primary training run completed

AI-001 Run 001 completed successfully at step 3,795.

Final metricValue
Final step3,795
Tokens processed62,177,280
Initial validation loss9.4776
Final validation loss4.2680
Best validation loss4.2680
Final training loss4.4347
Peak CUDA memory1.54 GiB
Best checkpointFinal checkpoint

The final checkpoint was also the best validation checkpoint of the complete run.

Validation loss was still decreasing at the end of training, moving from 4.2948 at step 3,500, to 4.2720 at step 3,750, and finally 4.2680 at step 3,795.

No late-run validation deterioration was observed.

Result The optimization phase of AI-001 completed successfully.

The model started from random initialization and processed approximately one complete corpus pass while continuously improving its performance on held-out validation data.

Large runtime artifacts — including checkpoints, intermediate model snapshots, metrics and console logs — are retained locally and intentionally excluded from Git.

03 Oct 2026Quantitative outcome

AI-001 asked whether a small language model could be trained entirely from random initialization on consumer hardware limited to 4 GB of VRAM and learn measurable structure from the prepared corpus.

The completed run establishes several results:

  • a 22.45M-parameter decoder-only Transformer can be trained locally on the GTX 1650;
  • the full training workflow remains well within the available GPU memory budget;
  • checkpoint recovery can survive a real interruption and preserve training continuity;
  • validation loss decreased consistently from 9.4776 to 4.2680;
  • the lowest validation loss occurred at the final checkpoint;
  • no sustained validation regression or optimization divergence was observed.

The quantitative results therefore support the central claim that meaningful learning can occur from random initialization under the experiment’s compute constraints.

They do not, by themselves, establish how coherent or useful the model’s generated language is.

03 Oct 2026Model behaviour evaluated

The final model and selected intermediate checkpoints were evaluated using fixed prompts across general English, Spanish, web development, WordPress and programming topics.

The first evaluation used stochastic sampling with a temperature of 0.8.

A clear progression is visible between early and late checkpoints.

For example, with the prompt A WordPress plugin can, the step-1,000 model produces largely incoherent text that mixes WordPress, MDN-like documentation and Docker-related material.

By the final checkpoint, the model generates structures such as:

A WordPress plugin can be used when the plugin is loaded

and:

A plugin contains a PHP file

These are not reliable technical explanations, but they demonstrate that WordPress-specific vocabulary and relationships emerged during training.

A similar progression appears in web-development prompts.

For HTML is used to, early checkpoints mix HTML with unrelated material, while later checkpoints increasingly associate the prompt with browsers, pages and content.

CSS generations also begin to associate the prompt with layout and HTML structure, although longer continuations eventually lose coherence.

Observation The improvement measured by validation loss is visible in generated text and is not limited to the numerical training objective.

03 Oct 2026Per-corpus evaluation

The final checkpoint was evaluated separately across corpus categories to determine what parts of the dataset the model had learned most effectively.

The resulting global document-aware validation loss was:

4.3015

This is close to the 4.2680 validation loss measured during training.

The values are not expected to be identical: the trainer evaluates sampled windows from the token stream, while the per-corpus evaluation respects individual document boundaries.

The two measurements are therefore consistent.

At the broadest level:

Corpus groupLoss
English general4.5171
Spanish general4.3132
Technical3.4788

The strongest result is immediately visible: technical material was learned substantially better than general-language material.

Within the technical corpus:

SourceLoss
MDN2.9566
WordPress2.9995
Docker3.8273
Python documentation3.8702
Git4.1085

MDN and WordPress are the two most predictable technical sources for the final model.

Their perplexities are approximately:

SourcePerplexity
MDN19.2
WordPress20.1

This helps explain a behaviour already visible during generation: the model frequently produces MDN- and WordPress-like structures, including headings, examples, APIs, command references and documentation conventions.

General-language corpora remain considerably harder.

Observed perplexities include:

SourcePerplexity
Wikipedia EN93.6
Wikibooks EN90.6
Wikipedia ES81.6
Wikibooks ES64.3

The 22.45M-parameter model therefore appears substantially better at modelling constrained, repetitive technical documentation than broad natural-language distributions.

Finding AI-001 learned the technical corpus significantly more strongly than the general-language corpus.

Bilingual learning

The per-corpus results also provide an important result about language balance.

General Spanish performed slightly better than general English:

LanguageGeneral loss
English4.5171
Spanish4.3132

When measured more broadly, the languages are nearly tied:

LanguageLoss
English4.2942
Spanish4.3132

There is therefore no evidence that Spanish failed because of the bilingual corpus composition.

The model learned recognizable structures in both languages despite its small parameter count.

Spanish generations contain vocabulary, grammatical patterns and encyclopedia-like prose, while English generations increasingly resemble natural sentences as training progresses.

Finding The bilingual training mixture worked reasonably well. Spanish was not meaningfully disadvantaged relative to English.

Learned structure versus learned meaning

The strongest capability acquired by AI-001 is the ability to reproduce the statistical form and style of its training sources.

The model learned patterns including:

  • documentation-style prose;
  • Markdown headings;
  • parameter and example sections;
  • code blocks;
  • API and command names;
  • links and reference-like structures;
  • technical vocabulary;
  • encyclopedia-style prose.

This can make generations appear surprisingly convincing.

However, stylistic plausibility does not imply semantic reliability.

Many outputs resemble real MDN, WordPress or command-line documentation while describing nonexistent APIs, incorrect relationships or fabricated technical behaviour.

The per-corpus evaluation reinforces this interpretation: MDN and WordPress are precisely the domains with the lowest loss.

Finding AI-001 learned how technical documentation looks more strongly than it learned what the documentation means.

Domain contamination

A recurring failure mode is leakage between technical domains.

WordPress generations adopt MDN-like structures. CSS prompts can drift into WordPress terminology. Docker generations may resemble unrelated web documentation. Technical prompts frequently converge toward the same general documentation register.

The model often recognizes that a prompt belongs to a technical domain without reliably maintaining the specific domain over a longer continuation.

This suggests that the combination of 22.45M parameters and a relatively diverse corpus provides insufficient capacity to keep all learned technical distributions cleanly separated.

Finding Domain separation is weak. The model learned a shared technical-documentation distribution more clearly than independent representations of each technical domain.

Programming ability remains weak

Programming prompts expose one of the clearest limitations of AI-001.

The prompt:

def fibonacci(n):

was tested across checkpoints.

None of the evaluated models produced a valid Fibonacci implementation.

Even the final checkpoint fails to maintain a useful code-generation trajectory and eventually drifts toward unrelated syntax or documentation-like material.

This is consistent with the dataset composition.

Python documentation teaches terminology, API structure and descriptions of programming concepts, but it is not equivalent to training on substantial amounts of executable source code.

Result AI-001 learned programming-related language, but did not develop reliable functional code-generation ability.

This establishes an important dataset requirement for future experiments: if programming ability is a target, real source code must become a meaningful part of the training corpus.

03 Oct 2026Deterministic generation evaluation

The initial generation tests used temperature = 0.8, which introduces sampling randomness.

A second evaluation was therefore run using greedy decoding with temperature = 0.

In this mode, the model always selects its highest-probability next token.

The results expose a strong tendency toward repetition loops.

For the Spanish prompt:

Barcelona es una ciudad

the final model eventually becomes trapped in repeated variations of:

de la ciudad

For:

A language model learns

the model falls into repeated constructions involving language, eventually producing sequences resembling:

the language is the language language...

The same failure appears in technical prompts.

CSS generations can collapse into repeated CSS sequences.

Git generations produce repeated git- patterns.

Docker generations can become trapped in repeated docker- structures.

The programming test is particularly clear:

def fibonacci(n):

does not produce a valid implementation in any tested checkpoint. Under greedy decoding, the model falls into repeated (n)-like patterns instead.

Finding The final model has a significant deterministic repetition problem.

What the greedy test means

The greedy behaviour does not indicate that training collapsed.

Several independent measurements show that the model learned successfully:

  • validation loss decreased from 9.4776 to 4.2680;
  • every recorded validation checkpoint improved;
  • the final checkpoint was the best checkpoint;
  • per-corpus evaluation shows meaningful differences between learned domains;
  • later checkpoints produce visibly more structured language than early checkpoints.

The problem occurs during autoregressive generation.

Validation loss measures how well the model predicts the next token when the preceding context comes from real text.

Generation is different.

Every generated token becomes part of the context used to predict the next one.

Once the model enters a highly probable repetitive pattern, that generated context can make another repetition even more probable, causing a feedback loop.

Sampling at temperature = 0.8 sometimes escapes these local probability peaks, which explains why sampled generations are more varied than greedy ones.

Finding AI-001 demonstrates successful next-token learning but weak free-running generation dynamics.

Decoding techniques such as temperature, top-p sampling or repetition penalties could reduce the visible repetition, but they would not solve the deeper limitations in capacity, semantic consistency or training data.

Behavioural evaluation summary

CapabilityResult
Learns next-token language structureClearly demonstrated
Validation generalizationClearly demonstrated
English language structureLearned, but shallow
Spanish language structureLearned, but shallow
Bilingual balanceSuccessful
Technical vocabularyClearly learned
Documentation stylesStrongly learned
MDN / WordPress patternsEspecially strong
Technical domain associationsEmerging
Prompt-domain persistenceInconsistent
Long-form coherenceWeak
Reliable factual knowledgeNot demonstrated
Functional code generationNot demonstrated
Greedy generation stabilityPoor / repetitive
Instruction followingNot evaluated / not trained for it
Improvement from early to final checkpointsClearly visible

The final model should therefore not be interpreted as a small general-purpose assistant.

It is better described as an early language model that has successfully learned statistical, linguistic and domain-specific structure from the corpus but lacks the capacity and training signal required for reliable long-form generation, factual consistency or programming.

AI-001 conclusion

AI-001 asked whether a language model could be trained entirely from random initialization on consumer hardware with only 4 GB of VRAM and learn measurable language structure.

The answer is yes.

A 22,451,712-parameter decoder-only Transformer completed 3,795 optimization steps and processed 62,177,280 tokens on an NVIDIA GTX 1650.

Validation loss decreased continuously from 9.4776 to 4.2680, with the final checkpoint producing the best validation result of the entire run.

Peak CUDA memory usage was only 1.54 GiB, demonstrating that the complete training workload was viable within the available hardware constraint.

The main training run also survived a real CUDA interruption by restoring the model, optimizer, AMP scaler and RNG state from checkpoint and continuing normally to completion.

Qualitative evaluation confirms that the numerical improvement corresponds to observable learning.

The model acquired recognizable English and Spanish structure, technical vocabulary, documentation conventions and domain-specific associations.

Per-corpus evaluation shows that this learning was uneven.

Technical text reached a loss of 3.4788, substantially better than English general text at 4.5171 and Spanish general text at 4.3132.

MDN and WordPress were learned particularly strongly, reaching losses of 2.9566 and 2.9995 respectively.

This also explains one of the experiment’s main limitations: the model became particularly good at reproducing the appearance of technical documentation, sometimes without preserving its meaning.

It frequently mixes technical domains, invents plausible-looking facts and APIs, loses coherence over longer generations and fails to produce reliable executable code.

Greedy decoding reveals an additional limitation: the model is strongly susceptible to repetitive autoregressive loops.

These failures do not contradict the original hypothesis.

AI-001 was designed to determine whether meaningful language learning could emerge from random initialization under severe compute constraints, not whether a 22M-parameter model trained for approximately one corpus pass could become a useful general-purpose assistant.

Hypothesis evaluation Supported by evidence.

AI-001 successfully demonstrates the full path:

custom corpus → custom tokenizer → random Transformer → local training → decreasing held-out loss → bilingual language structure → technical-domain learning

It also identifies the next bottlenecks clearly:

model capacity, corpus composition, domain separation, source-code coverage and generation stability.

The experiment therefore provides both a successful proof of training feasibility and a quantitative baseline for the next AI Lab experiment.

Findings

What the evidence says

Results

AI-001 Run 001 completed successfully after 3,795 optimization steps and 62.18M processed tokens. A 22.45M-parameter decoder-only Transformer was trained entirely from random initialization on an NVIDIA GTX 1650 with 4 GB of VRAM.

Validation loss decreased consistently from 9.4776 to 4.2680, with the final checkpoint also achieving the best validation result of the run. Peak CUDA memory usage was 1.54 GiB.

A CUDA runtime interruption after step 3,540 was successfully recovered from the step-3,500 checkpoint, restoring the full training state and completing normally.

Hypothesis evaluation

Supported by evidence

What I learned

The experiment worked: a small transformer trained from scratch on consumer hardware learned measurable linguistic structure within a 4 GB VRAM constraint.

It did not become a useful general-purpose model — that was never the point. The useful result was the training pipeline itself: corpus preparation, tokenisation, checkpointing, evaluation and a repeatable way to test what changes actually improve the model.

The next experiments can now start from that system rather than from zero.

Next steps

  • Run a deterministic generation evaluation using greedy decoding to reduce sampling noise when comparing checkpoints.
  • Measure validation loss independently across corpus domains, including Spanish general text, English general text, Python documentation, WordPress, MDN, Docker and Git.
  • Use per-domain loss to determine which parts of the corpus the model learned most effectively and where the largest weaknesses remain.
  • Analyze corpus proportions and contamination between technical domains.
  • Evaluate whether additional training beyond a single corpus pass continues to improve validation performance or begins to introduce overfitting.

Other experiments

All experiments
  1. Fine-tuning a tiny local model for tools, context and conversation

    AI.004
    Agents · AI · Edge AI
    Planned October 2026
  2. Can domain context improve local speech recognition?

    AI.002
    AI · Experimental computing
    Completed September 2026
  3. Can an AI planner choose the right evidence?

    AI.003
    Agents · AI · Experimental computing
    Completed October 2026

Keep exploring

More questions, tested in the open

Every experiment starts with a question and ends with notes worth keeping. Browse the rest of the lab, or see which ideas grew into full projects.