Every repeat query hitting the LLM at full price adds up fast, and it’s usually finance who notices first, after the AI line item has already tripled. Most teams treat this as a budget problem to fix later, but this is actually a design decision that belongs in the architecture from the start.
This session adds Redis semantic caching, so repeat and near-duplicate queries get served from cache instead of hitting the model again.
We’ll cover how semantic matching decides what counts as a repeat query, where the savings show up first, and the reduction in LLM spend.
What you’ll learn:
Redis LangCache added to the app for semantic caching.
How semantic matching decides what counts as a cache hit.
The Mangoes.ai benchmark: 70% lower LLM spend, and what made that possible.
Speakers
Redis
Samuel Agbede
Developer Advocate
Register now
Get started with Redis today
Speak to a Redis expert and learn more about enterprise-grade Redis today.