A series · signalandweight.com
Signal & Weight
Where model weights live, how they move, and what they cost: one real model followed from storage to GPU memory to the next token.
Every part follows the same real model, Qwen3.8-Max, so each number connects to the ones before it. Each part comes back to at least one figure you can check yourself from the model's published config.
- Part 1
One token through a 2.4-trillion-parameter model
Every word an LLM writes costs a full trip through its weights. I followed one prompt through Qwen3.8-Max to see what that trip actually looks like, in bytes, GPUs and memory.
- Part 2
What weights actually store: compression, not a database
Coming soon
- Part 3
Quantization: the dial that decides how many GPUs you need
Coming soon
- Part 4
Weights at rest: files, shards and formats
Coming soon
- Part 5
Getting weights onto GPUs: the cold-start problem
Coming soon
- Part 6
The KV cache: the memory that is not weights
Coming soon
- Part 7
Mixture of Experts at scale: why batching touches every expert
Coming soon
- Part 8
Training checkpoints: seven times bigger, and the restore is what hurts
Coming soon