← 回到 Reading
Daily Dose of DS 2026-09-24

Build a Jev Judge

DeepLearning.AI and Oracle have launched a free course titled 'Building Adaptive AI Agents' to help developers prevent coding agents from repeating errors. The curriculum focuses on transforming noisy execution traces into structured, reusable procedures and code knowledge graphs. It also covers the specific conditions under which model fine-tuning is necessary to resolve persistent failures. Students learn a multi-layered approach to agent adaptation that combines procedural learning, structural context retrieval, and weight adjustment. Jev is a specialized model from TypeSafe AI designed for structured decision-making and agent evaluation, offering an alternative to traditional generative LLM judges. Unlike generative models that produce text token-by-token, Jev returns typed values from a predefined answer space, which reduces latency and cost. It works by evaluating multiple independent questions against a shared state in parallel. When paired with observability tools like Opik, it allows teams to scale their evaluation workflows and catch failures more efficiently.

閱讀原文 ↗
目錄 2 段
  1. 01Another brilliant course by Andrew Ng!
  2. 02MoE inference engineering, clearly explained
LLMs

Another brilliant course by Andrew Ng!

DeepLearning.AI and Oracle have launched a free course titled 'Building Adaptive AI Agents' to help developers prevent coding agents from repeating errors. The curriculum focuses on transforming noisy execution traces into structured, reusable procedures and code knowledge graphs. It also covers the specific conditions under which model fine-tuning is necessary to resolve persistent failures. Students learn a multi-layered approach to agent adaptation that combines procedural learning, structural context retrieval, and weight adjustment.

  • Raw execution traces often contain redundant or misleading information like abandoned hypotheses.
  • The course introduces a three-layer adaptation framework: procedure extraction, context retrieval, and fine-tuning.
  • Agents can extract reusable procedures from traces, which are then validated by humans.
  • Code knowledge graphs utilize Git history and structural imports to improve retrieval beyond simple keyword search.
  • Fine-tuning is specifically reserved for failures that persist even when the agent has access to correct procedures and context.
LLMs

MoE inference engineering, clearly explained

Jev is a specialized model from TypeSafe AI designed for structured decision-making and agent evaluation, offering an alternative to traditional generative LLM judges. Unlike generative models that produce text token-by-token, Jev returns typed values from a predefined answer space, which reduces latency and cost. It works by evaluating multiple independent questions against a shared state in parallel. When paired with observability tools like Opik, it allows teams to scale their evaluation workflows and catch failures more efficiently.

  • Jev uses a typed interface to return structured values instead of generating free-form text, reducing token usage and latency.
  • The model evaluates multiple independent questions against the same context (state) in a single request to improve efficiency.
  • Opik serves as the observability and experimentation layer, recording traces and evaluation results produced by Jev.
  • Jev's Noul primitive returns a probability for a proposition, while the Score primitive provides a probability-weighted average across rubric levels.
  • Jev is optimized for focused, bounded judgments rather than evaluations requiring detailed written explanations or complex hidden reasoning.
  • Evaluation systems using Jev should still incorporate deterministic checks for exact rules and human review for uncertain cases.