How to Build a Versioned LLM Cache in TypeScript for Reliable AI Apps

Neptune Infotech Team
Neptune Infotech Team
|
October 10, 2026
How to Build a Versioned LLM Cache in TypeScript for Reliable AI Apps

Introduction

When you cache large language model (LLM) responses, a tiny oversight in the cache key can cause the wrong answer to be served to the wrong user. In practice, identical prompt text from two different users often lands in the same cache entry, leaking personalized or context‑specific information.

Cache key versioning is defined as the practice of embedding model identifiers, prompt fingerprints, and contextual metadata into a single, immutable key so that each distinct request maps to its own cache entry.

Why a Simple Text Key Fails

Most developers start with the raw prompt string as the cache key. This approach ignores three critical dimensions:

  • Model version: Different model releases (e.g., "gpt‑4‑v1" vs "gpt‑4‑v2") produce different outputs for the same text.
  • User context: Two users asking "What is my refund limit?" expect answers based on their own accounts, not each other's.
  • Additional parameters: Temperature, system messages, or function calls change the result even if the visible prompt stays the same.

When any of these pieces are omitted, the cache returns a stale or wrong answer, as illustrated by Alice and Bob receiving each other's refund limits.

Designing a Versioned Cache Key in TypeScript

Below is a concise pattern you can drop into any Node.js service:

```typescript interface LLMRequest { model: string; // e.g. "gpt‑4‑v2" prompt: string; // user‑provided text temperature?: number; userId: string; // unique identifier for the caller // any other fields that affect the response } function hash(value: string): string { // use a fast, non‑cryptographic hash like xxhash64 return xxhash64(value).toString(16); } function buildCacheKey(req: LLMRequest): string { const modelPrefix = `model:${req.model}`; const userPrefix = `user:${req.userId}`; const promptHash = `prompt:${hash(req.prompt)}`; const tempSuffix = req.temperature !== undefined ? `temp:${req.temperature}` : ''; return [modelPrefix, userPrefix, promptHash, tempSuffix].filter(Boolean).join('|'); } ```

This key concatenates the model name, user identifier, a deterministic prompt hash, and any optional parameters. Changing the model version automatically creates a new namespace, allowing old entries to expire naturally.

Managing TTL and Invalidation

Even with a perfect key, cache entries must be refreshed. Common strategies include:

  1. Time‑based TTL: Set a short TTL (e.g., 5 minutes) for dynamic queries and a longer TTL (e.g., 24 hours) for static knowledge.
  2. Version‑driven expiration: When you upgrade a model, increment the version prefix; old keys become unreachable and are eventually evicted.
  3. Explicit invalidation: On user‑profile updates (e.g., refund limit change), purge keys that contain the affected userId.

Leveraging Semantic Caching with Vectors

Exact‑text keys miss opportunities when paraphrases convey the same intent. Redis 8 introduces native vector operations (VADD/VSIM) that let you store an embedding alongside the response. On a new request you compute its embedding, perform a similarity search, and treat a high‑score match as a cache hit.

This approach reduces API spend while preserving answer relevance, especially for FAQ‑style workloads where wording varies but meaning stays constant.

Best Practices for Production Deployment

  • Always include the model name and version in the key.
  • Hash the prompt with a fast, deterministic algorithm.
  • Scope the key by user or tenant ID for multi‑tenant services.
  • Combine exact‑text caching with a fallback semantic cache for paraphrase hits.
  • Monitor cache hit ratios and adjust TTLs based on observed drift in model outputs.

Frequently Asked Questions

What happens if I forget to include the user ID in the cache key?

The cache will treat requests from different users as identical, potentially leaking private information and violating data‑privacy regulations.

Is hashing the prompt enough to avoid collisions?

Using a high‑quality, low‑collision hash (e.g., xxhash64) makes collisions extremely unlikely. For ultra‑critical systems you can add a secondary checksum.

Can I use the same cache for both text and embeddings?

Yes. Store a JSON object that contains the raw response and the embedding vector. Redis allows you to index the vector separately for similarity queries.

How often should I rotate the model version prefix?

Whenever you upgrade to a new model release that changes output semantics—typically after each major release from the provider.

Do I need a separate cache for each microservice?

Not necessarily. A shared cache with well‑scoped keys (including service name or namespace) can serve multiple services while keeping entries isolated.

Neptune Infotech can help you implement robust, versioned LLM caching in your next AI‑powered product—reach out to turn these best practices into production‑ready code.

You Might Also Like

Explore more articles related to "AI/ML"

How Anthropic’s Free OSS Scanner Elevates Open‑Source Security

How Anthropic’s Free OSS Scanner Elevates Open‑Source Security

Anthropic’s recent launch of a free security scanning service for open‑source projects has sparked c...

How Atlassian‑OpenAI Partnership is Shaping Enterprise AI Workflows

How Atlassian‑OpenAI Partnership is Shaping Enterprise AI Workflows

Atlassian and OpenAI have announced an expanded partnership that embeds the latest frontier AI model...

How VS Code Extensions Like Lodestar Transform Codebase Navigation

How VS Code Extensions Like Lodestar Transform Codebase Navigation

Modern development teams often inherit large, complex codebases that lack up‑to‑date documentation,...