A way to reuse an already-processed prompt start so repeat requests run faster and cheaper.
Think of a recorded intro. The same opening plays again, so nobody says it twice.
It speeds up repeated long prompts. You may see it in support bots and company Q&A tools.
KV cache
Prompt caching reuses KV cache data for the same prompt start.
Prefill
Prompt caching skips Prefill for repeated prompt starts.
Inference engine
The Inference engine finds matching prompt starts and manages the cache.