NVIDIA is no longer the only sensible choice for serious AI inference and training in 2026. Groq serves Llama-70B at sub-second time-to-first-token; Cerebras WSE-3 fits an entire 70B model on one wafer; AWS Trainium2 has become the AWS-native cost leader; Google TPU v5p quietly trains models on JAX; AMD MI300X has reached the maturity threshold where ROCm is no longer an active impediment; and Tenstorrent has opened a workstation-class option with a fully open stack. This article is the architect's reference for choosing between them — per-vendor architecture, sweet spots, real benchmarks, tooling maturity, lock-in posture, and a comparison table that surfaces dollars per million output tokens for Llama-70B-class workloads in 2026.