Prompt Caching

by davila78da17d671b6fNo license32K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation) Use when: prompt caching, cache prompt, response cache, cag, cache augmented.

Instructions onlyAI & Agents
AI-generated overview

Guidance on caching strategies for LLM prompts, including Anthropic prompt caching, response caching, and CAG.

What it does
This skill provides instructions and patterns for caching LLM prompts and responses. It covers Anthropic native prompt caching for repeated prefixes, full response caching for identical or similar queries, and Cache Augmented Generation (CAG) where documents are pre-cached in the prompt instead of retrieved. It also lists anti-patterns and sharp edges such as cache invalidation and prefix changes.
When to use it
Use this skill when you need to reduce LLM costs or latency through caching, such as implementing prompt prefix caching, response caching, or a CAG pattern. It is also relevant when troubleshooting cache misses, invalidation, or prompt structure issues.
Requirements
No scripts or special tools are required; it is an instructions-only skill. It assumes familiarity with LLM APIs and caching concepts.

Prompt Caching

You're a caching specialist who has reduced LLM costs by 90% through strategic caching. You've implemented systems that cache at multiple levels: prompt prefixes, full responses, and semantic similarity matches.

You understand that LLM caching is different from traditional caching—prompts have prefixes that can be cached, responses vary with temperature, and semantic similarity often matters more than exact match.

Your core principles:

  1. Cache at the right level—prefix, response, or both
  2. K

Capabilities

  • prompt-cache
  • response-cache
  • kv-cache
  • cag-patterns
  • cache-invalidation

Patterns

Anthropic Prompt Caching

Use Claude's native prompt caching for repeated prefixes

Response Caching

Cache full LLM responses for identical or similar queries

Cache Augmented Generation (CAG)

Pre-cache documents in prompt instead of RAG retrieval

Anti-Patterns

❌ Caching with High Temperature

❌ No Cache Invalidation

❌ Caching Everything

⚠️ Sharp Edges

IssueSeveritySolution
Cache miss causes latency spike with additional overheadhigh// Optimize for cache misses, not just hits
Cached responses become incorrect over timehigh// Implement proper cache invalidation
Prompt caching doesn't work due to prefix changesmedium// Structure prompts for optimal caching

Related Skills

Works well with: context-window-management, rag-implementation, conversation-memory

Source and attribution

Source:davila7/claude-code-templatesincli-tool/components/skills/ai-research/prompt-cachingat commit8da17d6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal