← 回到 Reading
Daily Dose of DS 2026-09-30

Jev for RAG, clearly explained!

Large language models are inherently stateless and rely on external systems to retain context across interactions. Agent memory solves this by dividing state into short-term working memory for active sessions and long-term memory for cross-session persistence. Long-term memory is further categorized into semantic, episodic, and procedural types, each governed by specialized storage and retrieval policies rather than model weight updates. Oracle AI Agent Memory implements this structured approach, significantly reducing per-turn token consumption and outperforming flat conversation history in benchmark evaluations. Hybrid search often retrieves passages that share terms or semantic similarity without actually providing valid evidence to answer a query. Jev introduces an explicit evaluation stage between retrieval and generation to score candidates via typed probability estimates in a single batch request. Application logic can then filter passages based on these probabilities and assess overall answerability before passing context to the LLM. This decouples subjective relevance judgments from deterministic business policies and enables improved auditability. Jev is a tool designed to handle semantic decisions that standard code cannot reliably express, outputting typed answers and probabilities while code controls the workflow. It supports twelve practical agentic use cases including bounded DOM action selection, trace pruning, dynamic skill loading, function argument extraction, semantic CI evaluation, and trace lesson extraction. The Beacon open-source project by Asymptote Labs uses Jev to analyze sessions across developer agents like Claude Code, Cursor, and Codex to extract reusable lessons.

閱讀原文 ↗
目錄 3 段
  1. 01Agents without memory aren’t agents at all
  2. 02Jev for RAG, clearly explained!
  3. 03Top 12 agentic use cases for Jev:
AGENTS

Agents without memory aren’t agents at all

Large language models are inherently stateless and rely on external systems to retain context across interactions. Agent memory solves this by dividing state into short-term working memory for active sessions and long-term memory for cross-session persistence. Long-term memory is further categorized into semantic, episodic, and procedural types, each governed by specialized storage and retrieval policies rather than model weight updates. Oracle AI Agent Memory implements this structured approach, significantly reducing per-turn token consumption and outperforming flat conversation history in benchmark evaluations.

  • Large language models do not store state in weights across sessions; persistence must be managed externally by the application.
  • Agent memory operates across two primary scopes: short-term working state for current sessions and long-term state across sessions.
  • Long-term memory comprises semantic memory for facts, episodic memory for past experiences, and procedural memory for instructions and workflows.
  • Different memory categories require distinct write and retrieval policies based on persistence needs.
  • In an 80-turn evaluation, Oracle AI Agent Memory capped inputs at around 1,300 tokens per request, compared to over 13,900 tokens for flat history.
  • The managed-memory agent outperformed flat history in 48 out of 80 evaluated turns, with 13 losses and 19 ties.
LLMs

Jev for RAG, clearly explained!

Hybrid search often retrieves passages that share terms or semantic similarity without actually providing valid evidence to answer a query. Jev introduces an explicit evaluation stage between retrieval and generation to score candidates via typed probability estimates in a single batch request. Application logic can then filter passages based on these probabilities and assess overall answerability before passing context to the LLM. This decouples subjective relevance judgments from deterministic business policies and enables improved auditability.

  • Hybrid search can achieve high recall via dense search, BM25, and reciprocal rank fusion, but retrieval quality bounds the system's overall performance.
  • Jev evaluates whether passages provide true evidence by returning typed probability scores across all candidates in a single packed request.
  • Decoupling probability estimation in Jev from policy thresholds in application code allows teams to tune precision-recall tradeoffs independently.
  • Jev can perform secondary checks in the same request, such as evaluating whether the collective retained passages are sufficient to answer the query or flagging potential prompt injections.
  • An open-source example integrates Jev as an evaluation judge alongside Comet Opik for AI observability and trace auditing.
JEV

Top 12 agentic use cases for Jev:

Jev is a tool designed to handle semantic decisions that standard code cannot reliably express, outputting typed answers and probabilities while code controls the workflow. It supports twelve practical agentic use cases including bounded DOM action selection, trace pruning, dynamic skill loading, function argument extraction, semantic CI evaluation, and trace lesson extraction. The Beacon open-source project by Asymptote Labs uses Jev to analyze sessions across developer agents like Claude Code, Cursor, and Codex to extract reusable lessons.

  • Jev provides typed answers and probabilities to make semantic decisions within programmatic workflows.
  • Key agentic use cases for Jev span DOM interaction, verbatim trace pruning, skill selection, argument validation, trajectory labeling, and CI semantic regression testing.
  • Jev can act as a secondary verification or reranking layer behind cheaper extractors or embedding-based retrieval pipelines.
  • The Beacon open-source project uses Jev to extract reusable lessons and corrections across more than 20 agent harnesses, including Claude Code, Codex, Cursor, and OpenCode.