Training a language model from scratch on 4GB of VRAM
AI.001 · Completed
Training a small language model from random initialization to explore how much linguistic structure can emerge under severe consumer-hardware constraints.
Status
Completed
Started
September 2026
Completed
October 2026
Topics
AILanguage ModelsLocal AIMachine Learning
Tools & Stack
CUDA
Python
PyTorch
WSL2
Research brief
The question
Can I train a language model entirely from random initialization on a consumer laptop with only 4 GB of VRAM, and get it to learn measurable language structure?
01
Why this experiment
Most local AI experiments begin with an existing pretrained model. This experiment goes one layer deeper: starting from random initialization to understand the full training pipeline — dataset construction, tokenization, architecture, optimization, checkpointing, evaluation and generation.
The goal is not to build a competitive general-purpose language model, but to explore how much language structure can emerge under severe compute constraints and to build a reproducible foundation for future AI Lab experiments.
02
Success criteria
The experiment will be considered successful if:
Training remains numerically stable without persistent divergence or out-of-memory failures.
Training and validation loss show meaningful improvement over the initial baseline.
The complete training workflow must remain viable within 4 GB of GPU memory.
AI-001 uses no pretrained model weights, LoRA adapters or external training compute for the core experiment.
04
Technical target
Train a compact decoder-only Transformer entirely from random initialization using a custom corpus and tokenizer.
Current AI-001 configuration
Context length: 256 tokens
Training corpus: 62,171,404 tokens
Validation corpus: 1,194,089 tokens
Microbatch size: 16
Gradient accumulation: 4
Effective tokens per optimization step: 16,384
Mixed-precision GPU training
Periodic validation
Automatic checkpointing
Resumable optimizer and training state
The final architecture, parameter count and training schedule will be recorded from the frozen configuration used for the primary AI-001 run.
Hypothesis
A small decoder-only Transformer trained from scratch on a carefully constructed corpus can learn measurable language structure on a consumer GPU with 4 GB of VRAM, provided that model size, context length, batch strategy and dataset size are adapted to the hardware constraints.
Development
Lab notebook
18 Sep 2026Research environment initialized
AI-001 starts as the first formal experiment inside AI Lab.
Initial constraint The complete training workflow must be viable on an NVIDIA GTX 1650 with 4 GB of VRAM.
The objective is to train a small language model entirely from random initialization using consumer hardware, rather than fine-tuning or adapting an existing pretrained model.
A dedicated ai-lab repository was created to keep research code, datasets, experiment definitions, models and results separated from other local AI projects.
Decision Treat hardware limitations as part of the experiment rather than moving training to external compute.
18 Sep 2026GPU training environment validated
The local training environment was prepared under WSL2 with PyTorch and CUDA support.
PyTorch successfully detected the NVIDIA GTX 1650 and GPU execution was validated.
This established that AI-001 could proceed using the laptop GPU without requiring cloud infrastructure.
Result Local CUDA training is operational.
Sep 2026Dataset strategy defined
The first corpus strategy was designed around a relatively small but varied text dataset that could be processed and trained locally.
The goal was not simply to maximize token count, but to provide enough linguistic variety for a small model to learn observable structure.
Sources were evaluated according to:
text quality;
licensing and accessibility;
language balance;
duplication risk;
preprocessing complexity;
usefulness for a small general-purpose language model.
The corpus design includes general prose alongside educational and technical material.
Decision Build a controlled corpus rather than downloading an unrestricted large-scale web dataset.
Sep–Oct 2026Training pipeline implemented
A dedicated AI-001 training pipeline was built around a compact decoder-only Transformer.
The trainer supports:
GPU training;
gradient accumulation;
periodic validation;
checkpoint creation;
model snapshots;
deterministic experiment configuration;
interruption and resume from saved training state.
The pipeline was designed specifically around the available memory budget rather than using default configurations intended for larger GPUs.
Key constraint 4 GB VRAM.
03 Oct 2026AI-001 training configuration established
The current AI-001 dataset contains:
Metric
Value
Training tokens
62,171,404
Validation tokens
1,194,089
Context length
256
Microbatch size
16
Gradient accumulation
4
Effective tokens / update
16,384
Training runs on the NVIDIA GTX 1650.
This configuration represents the first practical training setup that fits the experiment’s hardware constraints while maintaining a useful effective batch size.
Observation Gradient accumulation allows the effective training batch to be substantially larger than what can physically fit in GPU memory at once.
03 Oct 2026Checkpoint and resume test
Before committing to a long training run, checkpoint recovery was tested independently.
A short AI-001 run was stopped after optimization step 10 and then restarted using the generated latest.pt checkpoint.
At this stage, the experiment has demonstrated that the infrastructure required to train the model from scratch works within the available hardware constraints.
This does not yet validate the main hypothesis.
The remaining question is whether sustained training produces measurable language learning.
Current state AI-001 is ready for longer training and model-quality evaluation.
03 Oct 2026Main training run started
With the training pipeline validated, AI-001 moved from infrastructure testing to the primary learning run.
The model is a compact decoder-only Transformer containing 22,451,712 trainable parameters, initialized entirely from random weights.
Parameter
Value
Parameters
22,451,712
Vocabulary
12,000 tokens
Context length
256
Transformer layers
10
Model dimension
384
Attention heads
6
MLP dimension
1,536
Microbatch size
16
Gradient accumulation
4
Effective tokens / update
16,384
Optimizer
AdamW
Peak learning rate
3e-4
Minimum learning rate
3e-5
Warmup
200 steps
Precision
FP16 AMP
The training target was 3,795 optimization steps, corresponding to 62,177,280 processed tokens.
This is approximately one complete pass through the 62.17M-token training corpus.
Baseline validation loss 9.4776
From this point forward, the experiment measures whether training from random initialization produces consistent improvement on held-out validation data.
03 Oct 2026Validation loss falls throughout training
Validation was performed regularly during the main run.
Step
Validation loss
0
9.4776
250
6.4093
500
5.8517
750
5.5165
1,000
5.2391
1,250
5.0322
1,500
4.8696
1,750
4.7362
2,000
4.6319
2,250
4.5452
2,500
4.4689
2,750
4.4067
3,000
4.3613
3,250
4.3237
3,500
4.2948
3,750
4.2720
3,795
4.2680
Validation loss decreased at every recorded evaluation point, from an initial 9.4776 to a final 4.2680.
This represents a reduction of approximately 55% from the randomly initialized baseline.
There is no sustained validation regression visible during the run. The final checkpoint also produced the lowest validation loss observed during AI-001.
Observation The model continued improving on held-out data throughout the complete training run.
This provides quantitative evidence that training is producing generalizable predictive structure rather than merely reducing loss on the active training batches.
03 Oct 2026Training remains viable within 4 GB VRAM
The complete model and training configuration remained comfortably inside the experiment’s primary hardware constraint.
Peak CUDA memory usage during the run was 1.54 GiB on the NVIDIA GTX 1650 4 GB.
Metric
Result
Peak CUDA memory
1.54 GiB
Available GPU memory
4 GB
Final training loss
4.4347
Final validation loss
4.2680
Tokens processed
62,177,280
Optimization steps
3,795
No NaN divergence or GPU out-of-memory failure was observed during optimization.
Result A 22.45M-parameter Transformer can be trained from scratch using this configuration on the available GTX 1650.
The hardware constraint therefore influenced the design of the experiment, but did not prevent completion of the training run.
03 Oct 2026CUDA interruption and recovery
Shortly after step 3,540, training was interrupted by a CUDA runtime failure reporting:
CUDA_ERROR_UNKNOWN / Sticky error detected
The interruption was not preceded by an out-of-memory condition, NaN loss or sustained validation regression.
A valid checkpoint was already available from step 3,500.
The run was restarted from this checkpoint, restoring:
model weights;
optimizer state;
AMP scaler state;
random number generator state;
training progress.
Training resumed successfully from step 3,500 and continued normally through the planned final step.
Result Checkpoint recovery worked under a real training interruption, not only during the earlier controlled resume test.
The incident did not prevent the resumed run from surpassing the previous best validation loss.
03 Oct 2026Primary training run completed
AI-001 Run 001 completed successfully at step 3,795.
Final metric
Value
Final step
3,795
Tokens processed
62,177,280
Initial validation loss
9.4776
Final validation loss
4.2680
Best validation loss
4.2680
Final training loss
4.4347
Peak CUDA memory
1.54 GiB
Best checkpoint
Final checkpoint
The final checkpoint was also the best validation checkpoint of the complete run.
Validation loss was still decreasing at the end of training, moving from 4.2948 at step 3,500, to 4.2720 at step 3,750, and finally 4.2680 at step 3,795.
No late-run validation deterioration was observed.
Result The optimization phase of AI-001 completed successfully.
The model started from random initialization and processed approximately one complete corpus pass while continuously improving its performance on held-out validation data.
Large runtime artifacts — including checkpoints, intermediate model snapshots, metrics and console logs — are retained locally and intentionally excluded from Git.
03 Oct 2026Quantitative outcome
AI-001 asked whether a small language model could be trained entirely from random initialization on consumer hardware limited to 4 GB of VRAM and learn measurable structure from the prepared corpus.
The completed run establishes several results:
a 22.45M-parameter decoder-only Transformer can be trained locally on the GTX 1650;
the full training workflow remains well within the available GPU memory budget;
checkpoint recovery can survive a real interruption and preserve training continuity;
validation loss decreased consistently from 9.4776 to 4.2680;
the lowest validation loss occurred at the final checkpoint;
no sustained validation regression or optimization divergence was observed.
The quantitative results therefore support the central claim that meaningful learning can occur from random initialization under the experiment’s compute constraints.
They do not, by themselves, establish how coherent or useful the model’s generated language is.
03 Oct 2026Model behaviour evaluated
The final model and selected intermediate checkpoints were evaluated using fixed prompts across general English, Spanish, web development, WordPress and programming topics.
The first evaluation used stochastic sampling with a temperature of 0.8.
A clear progression is visible between early and late checkpoints.
For example, with the prompt A WordPress plugin can, the step-1,000 model produces largely incoherent text that mixes WordPress, MDN-like documentation and Docker-related material.
By the final checkpoint, the model generates structures such as:
A WordPress plugin can be used when the plugin is loaded
and:
A plugin contains a PHP file
These are not reliable technical explanations, but they demonstrate that WordPress-specific vocabulary and relationships emerged during training.
A similar progression appears in web-development prompts.
For HTML is used to, early checkpoints mix HTML with unrelated material, while later checkpoints increasingly associate the prompt with browsers, pages and content.
CSS generations also begin to associate the prompt with layout and HTML structure, although longer continuations eventually lose coherence.
Observation The improvement measured by validation loss is visible in generated text and is not limited to the numerical training objective.
03 Oct 2026Per-corpus evaluation
The final checkpoint was evaluated separately across corpus categories to determine what parts of the dataset the model had learned most effectively.
The resulting global document-aware validation loss was:
4.3015
This is close to the 4.2680 validation loss measured during training.
The values are not expected to be identical: the trainer evaluates sampled windows from the token stream, while the per-corpus evaluation respects individual document boundaries.
The two measurements are therefore consistent.
At the broadest level:
Corpus group
Loss
English general
4.5171
Spanish general
4.3132
Technical
3.4788
The strongest result is immediately visible: technical material was learned substantially better than general-language material.
Within the technical corpus:
Source
Loss
MDN
2.9566
WordPress
2.9995
Docker
3.8273
Python documentation
3.8702
Git
4.1085
MDN and WordPress are the two most predictable technical sources for the final model.
Their perplexities are approximately:
Source
Perplexity
MDN
19.2
WordPress
20.1
This helps explain a behaviour already visible during generation: the model frequently produces MDN- and WordPress-like structures, including headings, examples, APIs, command references and documentation conventions.
The 22.45M-parameter model therefore appears substantially better at modelling constrained, repetitive technical documentation than broad natural-language distributions.
Finding AI-001 learned the technical corpus significantly more strongly than the general-language corpus.
Bilingual learning
The per-corpus results also provide an important result about language balance.
General Spanish performed slightly better than general English:
Language
General loss
English
4.5171
Spanish
4.3132
When measured more broadly, the languages are nearly tied:
Language
Loss
English
4.2942
Spanish
4.3132
There is therefore no evidence that Spanish failed because of the bilingual corpus composition.
The model learned recognizable structures in both languages despite its small parameter count.
Spanish generations contain vocabulary, grammatical patterns and encyclopedia-like prose, while English generations increasingly resemble natural sentences as training progresses.
Finding The bilingual training mixture worked reasonably well. Spanish was not meaningfully disadvantaged relative to English.
Learned structure versus learned meaning
The strongest capability acquired by AI-001 is the ability to reproduce the statistical form and style of its training sources.
The model learned patterns including:
documentation-style prose;
Markdown headings;
parameter and example sections;
code blocks;
API and command names;
links and reference-like structures;
technical vocabulary;
encyclopedia-style prose.
This can make generations appear surprisingly convincing.
However, stylistic plausibility does not imply semantic reliability.
Many outputs resemble real MDN, WordPress or command-line documentation while describing nonexistent APIs, incorrect relationships or fabricated technical behaviour.
The per-corpus evaluation reinforces this interpretation: MDN and WordPress are precisely the domains with the lowest loss.
Finding AI-001 learned how technical documentation looks more strongly than it learned what the documentation means.
Domain contamination
A recurring failure mode is leakage between technical domains.
WordPress generations adopt MDN-like structures. CSS prompts can drift into WordPress terminology. Docker generations may resemble unrelated web documentation. Technical prompts frequently converge toward the same general documentation register.
The model often recognizes that a prompt belongs to a technical domain without reliably maintaining the specific domain over a longer continuation.
This suggests that the combination of 22.45M parameters and a relatively diverse corpus provides insufficient capacity to keep all learned technical distributions cleanly separated.
Finding Domain separation is weak. The model learned a shared technical-documentation distribution more clearly than independent representations of each technical domain.
Programming ability remains weak
Programming prompts expose one of the clearest limitations of AI-001.
The prompt:
def fibonacci(n):
was tested across checkpoints.
None of the evaluated models produced a valid Fibonacci implementation.
Even the final checkpoint fails to maintain a useful code-generation trajectory and eventually drifts toward unrelated syntax or documentation-like material.
This is consistent with the dataset composition.
Python documentation teaches terminology, API structure and descriptions of programming concepts, but it is not equivalent to training on substantial amounts of executable source code.
Result AI-001 learned programming-related language, but did not develop reliable functional code-generation ability.
This establishes an important dataset requirement for future experiments: if programming ability is a target, real source code must become a meaningful part of the training corpus.
03 Oct 2026Deterministic generation evaluation
The initial generation tests used temperature = 0.8, which introduces sampling randomness.
A second evaluation was therefore run using greedy decoding with temperature = 0.
In this mode, the model always selects its highest-probability next token.
The results expose a strong tendency toward repetition loops.
For the Spanish prompt:
Barcelona es una ciudad
the final model eventually becomes trapped in repeated variations of:
de la ciudad
For:
A language model learns
the model falls into repeated constructions involving language, eventually producing sequences resembling:
the language is the language language...
The same failure appears in technical prompts.
CSS generations can collapse into repeated CSS sequences.
Git generations produce repeated git- patterns.
Docker generations can become trapped in repeated docker- structures.
The programming test is particularly clear:
def fibonacci(n):
does not produce a valid implementation in any tested checkpoint. Under greedy decoding, the model falls into repeated (n)-like patterns instead.
Finding The final model has a significant deterministic repetition problem.
What the greedy test means
The greedy behaviour does not indicate that training collapsed.
Several independent measurements show that the model learned successfully:
validation loss decreased from 9.4776 to 4.2680;
every recorded validation checkpoint improved;
the final checkpoint was the best checkpoint;
per-corpus evaluation shows meaningful differences between learned domains;
later checkpoints produce visibly more structured language than early checkpoints.
The problem occurs during autoregressive generation.
Validation loss measures how well the model predicts the next token when the preceding context comes from real text.
Generation is different.
Every generated token becomes part of the context used to predict the next one.
Once the model enters a highly probable repetitive pattern, that generated context can make another repetition even more probable, causing a feedback loop.
Sampling at temperature = 0.8 sometimes escapes these local probability peaks, which explains why sampled generations are more varied than greedy ones.
Decoding techniques such as temperature, top-p sampling or repetition penalties could reduce the visible repetition, but they would not solve the deeper limitations in capacity, semantic consistency or training data.
Behavioural evaluation summary
Capability
Result
Learns next-token language structure
Clearly demonstrated
Validation generalization
Clearly demonstrated
English language structure
Learned, but shallow
Spanish language structure
Learned, but shallow
Bilingual balance
Successful
Technical vocabulary
Clearly learned
Documentation styles
Strongly learned
MDN / WordPress patterns
Especially strong
Technical domain associations
Emerging
Prompt-domain persistence
Inconsistent
Long-form coherence
Weak
Reliable factual knowledge
Not demonstrated
Functional code generation
Not demonstrated
Greedy generation stability
Poor / repetitive
Instruction following
Not evaluated / not trained for it
Improvement from early to final checkpoints
Clearly visible
The final model should therefore not be interpreted as a small general-purpose assistant.
It is better described as an early language model that has successfully learned statistical, linguistic and domain-specific structure from the corpus but lacks the capacity and training signal required for reliable long-form generation, factual consistency or programming.
AI-001 conclusion
AI-001 asked whether a language model could be trained entirely from random initialization on consumer hardware with only 4 GB of VRAM and learn measurable language structure.
The answer is yes.
A 22,451,712-parameter decoder-only Transformer completed 3,795 optimization steps and processed 62,177,280 tokens on an NVIDIA GTX 1650.
Validation loss decreased continuously from 9.4776 to 4.2680, with the final checkpoint producing the best validation result of the entire run.
Peak CUDA memory usage was only 1.54 GiB, demonstrating that the complete training workload was viable within the available hardware constraint.
The main training run also survived a real CUDA interruption by restoring the model, optimizer, AMP scaler and RNG state from checkpoint and continuing normally to completion.
Qualitative evaluation confirms that the numerical improvement corresponds to observable learning.
The model acquired recognizable English and Spanish structure, technical vocabulary, documentation conventions and domain-specific associations.
Per-corpus evaluation shows that this learning was uneven.
Technical text reached a loss of 3.4788, substantially better than English general text at 4.5171 and Spanish general text at 4.3132.
MDN and WordPress were learned particularly strongly, reaching losses of 2.9566 and 2.9995 respectively.
This also explains one of the experiment’s main limitations: the model became particularly good at reproducing the appearance of technical documentation, sometimes without preserving its meaning.
It frequently mixes technical domains, invents plausible-looking facts and APIs, loses coherence over longer generations and fails to produce reliable executable code.
Greedy decoding reveals an additional limitation: the model is strongly susceptible to repetitive autoregressive loops.
These failures do not contradict the original hypothesis.
AI-001 was designed to determine whether meaningful language learning could emerge from random initialization under severe compute constraints, not whether a 22M-parameter model trained for approximately one corpus pass could become a useful general-purpose assistant.
Hypothesis evaluation Supported by evidence.
AI-001 successfully demonstrates the full path:
custom corpus → custom tokenizer → random Transformer → local training → decreasing held-out loss → bilingual language structure → technical-domain learning
It also identifies the next bottlenecks clearly:
model capacity, corpus composition, domain separation, source-code coverage and generation stability.
The experiment therefore provides both a successful proof of training feasibility and a quantitative baseline for the next AI Lab experiment.
Lab photos
Setup and evidence
Findings
What the evidence says
01
Results
AI-001 Run 001 completed successfully after 3,795 optimization steps and 62.18M processed tokens. A 22.45M-parameter decoder-only Transformer was trained entirely from random initialization on an NVIDIA GTX 1650 with 4 GB of VRAM.
Validation loss decreased consistently from 9.4776 to 4.2680, with the final checkpoint also achieving the best validation result of the run. Peak CUDA memory usage was 1.54 GiB.
A CUDA runtime interruption after step 3,540 was successfully recovered from the step-3,500 checkpoint, restoring the full training state and completing normally.
02
Hypothesis evaluation
Supported by evidence
03
What I learned
The experiment worked: a small transformer trained from scratch on consumer hardware learned measurable linguistic structure within a 4 GB VRAM constraint.
It did not become a useful general-purpose model — that was never the point. The useful result was the training pipeline itself: corpus preparation, tokenisation, checkpointing, evaluation and a repeatable way to test what changes actually improve the model.
The next experiments can now start from that system rather than from zero.
04
Next steps
Run a deterministic generation evaluation using greedy decoding to reduce sampling noise when comparing checkpoints.
Measure validation loss independently across corpus domains, including Spanish general text, English general text, Python documentation, WordPress, MDN, Docker and Git.
Use per-domain loss to determine which parts of the corpus the model learned most effectively and where the largest weaknesses remain.
Analyze corpus proportions and contamination between technical domains.
Evaluate whether additional training beyond a single corpus pass continues to improve validation performance or begins to introduce overfitting.