Back to Blogai-architecture-patterns 
HBM Is the Bottleneck, Not FLOPs: Designing Inference Around Memory Bandwidth
hbm memory bandwidth llm inference roofline model gpu architecture nvidia h200 nvidia blackwell b200 amd mi300x hbm4 inference optimization kv cache continuous batching quantization ai infrastructure accelerator selection 2026
