Infrastructure
GPUs, memory, storage and serving: what it takes to run a model.
This topic includes the Signal & Weight series.
One token through a 2.4-trillion-parameter model
Every word an LLM writes costs a full trip through its weights. I followed one prompt through Qwen3.8-Max to see what that trip actually looks like, in bytes, GPUs and memory.