← 回到 Reading
Daily Dose of DS 2026-09-10

Why Multi-turn Agents Need More Than a Task Graph

Redis has introduced Redis LangCache, a managed service designed to reduce LLM operational costs and latency through external response caching. By using embeddings to identify semantically similar questions, the system can return pre-generated answers without invoking the LLM. This method bypasses token processing and decoding time, offering significant performance gains over traditional prefix caching. Redis reports that this approach can cut API costs by up to 90% and improve response speeds by up to 15x. Standard agent frameworks often fail in multi-turn conversations because they treat each run as a bounded graph, leading to leaked execution states where previous outputs are repeated. CrewAI addresses this by introducing conversational flows that separate the lifecycles of application state, such as conversation history, from execution state, such as node progress. This architecture ensures that while domain data persists across turns, the graph's internal execution records reset for every new message. Key components of this system include session-based restoration, a router for efficient task handling, and session-level tracing for debugging long-term interactions.

閱讀原文 ↗
目錄 2 段
  1. 01Redis built a cache that cuts LLM costs by 90%!
  2. 02Why multi-turn agents need more than a task graph
LLMs

Redis built a cache that cuts LLM costs by 90%!

Redis has introduced Redis LangCache, a managed service designed to reduce LLM operational costs and latency through external response caching. By using embeddings to identify semantically similar questions, the system can return pre-generated answers without invoking the LLM. This method bypasses token processing and decoding time, offering significant performance gains over traditional prefix caching. Redis reports that this approach can cut API costs by up to 90% and improve response speeds by up to 15x.

  • Redis LangCache caches generated LLM responses externally to prevent redundant model calls for similar questions.
  • The system utilizes embeddings to match new queries against a database of previously answered questions.
  • A successful cache hit results in zero input or output token consumption for that specific request.
  • Redis LangCache provides enterprise features like TTL, eviction controls, and data isolation via Redis Cloud.
  • In internal tests, the tool reduced response time from 2.232 seconds to 0.373 seconds for paraphrased queries.
  • The service can potentially reduce LLM API costs by up to 90% depending on workload repetition.
HANDS-ON

Why multi-turn agents need more than a task graph

Standard agent frameworks often fail in multi-turn conversations because they treat each run as a bounded graph, leading to leaked execution states where previous outputs are repeated. CrewAI addresses this by introducing conversational flows that separate the lifecycles of application state, such as conversation history, from execution state, such as node progress. This architecture ensures that while domain data persists across turns, the graph's internal execution records reset for every new message. Key components of this system include session-based restoration, a router for efficient task handling, and session-level tracing for debugging long-term interactions.

  • Multi-turn agents require separate lifecycles for conversation history, which must persist, and execution state, which must reset per turn.
  • Reusing the same flow instance without resetting execution records causes agents to repeat previous answers instead of processing new messages.
  • CrewAI conversational flows use session IDs to restore history and domain data while clearing per-run execution records.
  • Routing allows agents to bypass expensive research paths for simple clarifications, reducing latency and operational cost.
  • Effective output management separates internal agent scratch work from the final user-visible conversation history to prevent context bloat.
  • Session-level tracing is necessary to debug interactions across multiple independent graph runs that form a single conversation.