← 回到 Reading
Daily Dose of DS 2026-09-22

Build your own Jev (100% local)

HarnessRouter is a new open-source infrastructure layer that provides a unified interface for multiple agent harnesses. It simplifies development by handling sessions, streaming, and failure management across different tools like Codex and Claude code. The system utilizes the Unified Harness Protocol (UHP) to maintain an OpenAI-compatible API while running harnesses locally. This section describes how to implement a local version of the Jev inference pattern, which prioritizes scoring over text generation for classification tasks. By providing a finite set of possible answers, the system extracts the model's logits for specific tokens at the next position to determine a probability distribution. This approach is significantly faster than standard generation because it avoids the autoregressive decoding loop and eliminates the need for parsing text or JSON. The implementation utilizes SGLang's scoring endpoint to efficiently retrieve these scores from open models like Qwen and DeepSeek.

閱讀原文 ↗
目錄 2 段
  1. 01Finally, an OpenRouter for agent harnesses
  2. 02Build your own Jev (100% local)
OPEN-SOURCE

Finally, an OpenRouter for agent harnesses

HarnessRouter is a new open-source infrastructure layer that provides a unified interface for multiple agent harnesses. It simplifies development by handling sessions, streaming, and failure management across different tools like Codex and Claude code. The system utilizes the Unified Harness Protocol (UHP) to maintain an OpenAI-compatible API while running harnesses locally.

  • HarnessRouter acts as a plug-and-play infrastructure layer for agent harnesses.
  • Supported harnesses include Codex, Hermes, Claude code, DeepSeek Harness, and System One.
  • The tool manages sessions, streaming, file handling, and cancellation logic automatically.
  • The Unified Harness Protocol (UHP) defines a common task interface for all supported harnesses.
  • The system provides an API compatible with OpenAI Responses and executes harnesses locally.
HANDS-ON

Build your own Jev (100% local)

This section describes how to implement a local version of the Jev inference pattern, which prioritizes scoring over text generation for classification tasks. By providing a finite set of possible answers, the system extracts the model's logits for specific tokens at the next position to determine a probability distribution. This approach is significantly faster than standard generation because it avoids the autoregressive decoding loop and eliminates the need for parsing text or JSON. The implementation utilizes SGLang's scoring endpoint to efficiently retrieve these scores from open models like Qwen and DeepSeek.

  • Jev-style inference treats LLM calls as decisions by scoring a predefined list of valid answers rather than generating new text.
  • Scoring is more efficient than structured output because it retrieves probabilities from the first next-token vector without an autoregressive loop.
  • The SGLang /v1/score endpoint allows users to extract logits for specific token IDs and apply a restricted softmax to get a probability distribution.
  • To ensure accuracy, labels such as A, B, or C must be verified as single tokens using the model's specific tokenizer.
  • Scoring provides a confidence measure that can be used in application code to set thresholds for automated routing versus manual review.
  • This method is best suited for workloads with a finite set of labels where the caller does not require a textual explanation.