Contrastive Language Model, clearly explained
TrueFoundry's Auto Routing feature optimizes LLM production traffic by classifying requests based on complexity and routing them to appropriate models. The system uses zero-latency heuristics to categorize tasks into simple, medium, or complex tiers. This approach maintains conversation context and provides fallback mechanisms if a lower-tier model fails. Benchmarks show significant cost savings with minimal impact on output quality compared to using high-end models for all tasks. NVIDIA and Stanford researchers have introduced the Contrastive Language Model (CLM), a System 1 architecture designed for rapid decision-making in AI agents. Unlike traditional LLMs that generate text token-by-token, CLM treats decision-making as a retrieval problem by encoding states and actions into a shared vector space. This approach allows for action embeddings to be cached, resulting in up to 9x lower latency compared to the Jev model while maintaining comparable performance in gaming and tool-calling tasks. The model uses a frozen Qwen3-8B backbone and is trained using contrastive learning techniques.
閱讀原文 ↗目錄
Your best model shouldn’t be answering every request
TrueFoundry's Auto Routing feature optimizes LLM production traffic by classifying requests based on complexity and routing them to appropriate models. The system uses zero-latency heuristics to categorize tasks into simple, medium, or complex tiers. This approach maintains conversation context and provides fallback mechanisms if a lower-tier model fails. Benchmarks show significant cost savings with minimal impact on output quality compared to using high-end models for all tasks.
- TrueFoundry Auto Routing classifies incoming LLM requests into three tiers: simple, medium, and complex.
- The classification process uses in-process heuristics based on signals like code or reasoning phrases to avoid added latency.
- Conversation pinning ensures that follow-up requests in a complex thread remain with the higher-tier model.
- The system includes an automatic fallback that moves requests to a higher tier if the assigned model fails.
- In testing against an all-Opus baseline, the routing system reduced costs by 69% while maintaining 98% quality.
Contrastive Language Model, clearly explained
NVIDIA and Stanford researchers have introduced the Contrastive Language Model (CLM), a System 1 architecture designed for rapid decision-making in AI agents. Unlike traditional LLMs that generate text token-by-token, CLM treats decision-making as a retrieval problem by encoding states and actions into a shared vector space. This approach allows for action embeddings to be cached, resulting in up to 9x lower latency compared to the Jev model while maintaining comparable performance in gaming and tool-calling tasks. The model uses a frozen Qwen3-8B backbone and is trained using contrastive learning techniques.
- CLM treats AI decision-making as a retrieval task rather than a generative one, improving speed for repeated actions.
- The architecture achieves up to 9x faster performance than the Jev model by caching action embeddings and using cheap dot products.
- The model utilizes a frozen Qwen3-8B backbone with a small trainable state projection head.
- Training follows a three-stage process using InfoNCE, starting with QA pairs and moving to real agent trajectories.
- CLM is optimized for System 1 tasks such as tool selection, request routing, and game actions.
- While highly efficient, CLM cannot generate new actions and is limited to evaluating supplied candidates.
- The research and code for CLM are open-source.