Back to Blogai-architecture 
TPU Inference Architecture: Serving LLMs on Trillium with vLLM
tpu inference trillium tpu tpu v6e vllm tpu backend google cloud tpu llm serving architecture xla compilation jax inference gpu vs tpu ai accelerators self-hosted llm continuous batching mcp infrastructure ops ai inference cost ai architecture patterns 2026
