Back to Blogai-architecture 
Prompt Caching Architecture for LLM Apps & Agents: Prefix Caching, Cost, and Latency
prompt caching prefix caching cached input tokens llm cost optimization llm latency time to first token anthropic prompt caching openai prompt caching gemini context caching agent tool loop rag caching multi-turn chat ai gateway llm inference cost ai architecture patterns 2026
