The Thinking Layer

A series · signalandweight.com

Signal & Weight

Where model weights live, how they move, and what they cost: one real model followed from storage to GPU memory to the next token.

Every part follows the same real model, Qwen3.8-Max, so each number connects to the ones before it. Each part comes back to at least one figure you can check yourself from the model's published config.

  • Part 1

    One token through a 2.4-trillion-parameter model

    Every word an LLM writes costs a full trip through its weights. I followed one prompt through Qwen3.8-Max to see what that trip actually looks like, in bytes, GPUs and memory.

  • Part 2

    What weights actually store: compression, not a database

    Coming soon

  • Part 3

    Quantization: the dial that decides how many GPUs you need

    Coming soon

  • Part 4

    Weights at rest: files, shards and formats

    Coming soon

  • Part 5

    Getting weights onto GPUs: the cold-start problem

    Coming soon

  • Part 6

    The KV cache: the memory that is not weights

    Coming soon

  • Part 7

    Mixture of Experts at scale: why batching touches every expert

    Coming soon

  • Part 8

    Training checkpoints: seven times bigger, and the restore is what hurts

    Coming soon