Mark the long, repeated part of your Claude prompts as cacheable so repeat requests cost a fraction and respond faster.
Turn on prompt caching so I stop paying full price for the same long prompt
How it works
- Find repeated prompt content: Vovy asks Claude Code to find prompts that resend the same big block every time, like a long system prompt, docs, or tool definitions.
- Reorder for caching: Caching works on the start of a prompt. Claude Code moves the stable content first and the changing user message last.
- Add cache_control: Claude Code adds a cache_control breakpoint after the stable block in your Claude API call, so Anthropic reuses it on the next request.
- Check OpenAI calls too: OpenAI caches long repeated prefixes automatically. Vovy checks your OpenAI prompts follow the same stable-first order to benefit.
- Measure cache hits: Vovy sends two identical requests and shows the usage fields for cache reads versus normal input tokens, with the cost difference on a card.
What you provide
- Your app's AI code
- A Claude API key
What you get
- Caching turned on
- Prompts reordered for cache hits
- Measured savings per request
FAQ
How much does caching save?
On Claude, cached input tokens are billed at a small fraction of the normal rate. Writing to the cache costs a bit extra the first time.
How long does the cache last?
By default a few minutes, refreshed each time it is used. There is a longer option at a higher write price.
Why aren't my calls hitting the cache?
Anything that changes before the breakpoint, like a timestamp at the top of the prompt, breaks the match. Keep the cached part identical.
Related tasks
All tasks