6 LLM Deployment Formats in Production
Traditional AI agent web search is inefficient because agents must manually fetch and clean pages for every search hop, leading to high token usage. Seltz addresses this by providing a pre-processed index of web content across people, news, and wiki scopes. By using an owned index, the token count for a three-hop search can be reduced from 28,000 to under 7,000. This approach streamlines the reasoning process by delivering finished documents directly to the agent. The text outlines six primary LLM deployment formats, highlighting the trade-off between inference speed and hardware portability. Formats like Pickle and Safetensors offer broad compatibility but require Python environments, while specialized formats like MLX and TensorRT provide significant performance gains by optimizing for specific hardware like Apple Silicon or NVIDIA GPUs. GGUF and ONNX serve as intermediate solutions, with GGUF enabling standalone execution and ONNX facilitating cross-framework interoperability. Data leakage occurs when a model inadvertently accesses information during training that will not be available during inference, leading to over-optimistic performance. Common sources include train-test contamination, target leakage, and fitting preprocessing steps on the entire dataset rather than just the training split. To prevent these issues, practitioners should maintain temporal integrity in time-series data and fit all transformations exclusively on training data. Detection methods include monitoring for suspicious feature importance and evaluating models on independent datasets.
閱讀原文 ↗目錄
Your agents are doing web search wrong
Traditional AI agent web search is inefficient because agents must manually fetch and clean pages for every search hop, leading to high token usage. Seltz addresses this by providing a pre-processed index of web content across people, news, and wiki scopes. By using an owned index, the token count for a three-hop search can be reduced from 28,000 to under 7,000. This approach streamlines the reasoning process by delivering finished documents directly to the agent.
- Standard search calls return limited snippets, forcing agents to fetch and clean full pages repeatedly.
- Multi-hop search loops using raw web fetching are token-intensive, consuming up to 28,000 tokens for three hops.
- Seltz is a specialized index that crawls and processes web pages before a query is made.
- Using a pre-processed index like Seltz can reduce token consumption by approximately 75% for complex searches.
- Seltz covers specific scopes including people, news, and wiki content.
6 LLM deployment formats in production
The text outlines six primary LLM deployment formats, highlighting the trade-off between inference speed and hardware portability. Formats like Pickle and Safetensors offer broad compatibility but require Python environments, while specialized formats like MLX and TensorRT provide significant performance gains by optimizing for specific hardware like Apple Silicon or NVIDIA GPUs. GGUF and ONNX serve as intermediate solutions, with GGUF enabling standalone execution and ONNX facilitating cross-framework interoperability.
- LLM inference speed is increased by locking hardware-specific optimizations into the model file, which reduces portability.
- Pickle (.pt) files are potentially insecure because they can execute arbitrary code during the loading process.
- Safetensors improves upon Pickle by storing plain numbers in a memory-ready layout, eliminating the need for rebuilding during loading.
- GGUF bundles weights, tokenizers, and chat templates into a single file to allow execution without an ML framework.
- ONNX acts as an interchange format that allows models to run across different hardware by deferring specific optimizations to the runtime.
- NVIDIA's TensorRT compiles models into machine code tailored for a specific GPU architecture, making the files non-portable to other architectures.
- Apple's MLX framework optimizes performance by allowing the CPU and GPU to share a single memory pool on Apple Silicon.
Prevent data leakage in ML pipelines
Data leakage occurs when a model inadvertently accesses information during training that will not be available during inference, leading to over-optimistic performance. Common sources include train-test contamination, target leakage, and fitting preprocessing steps on the entire dataset rather than just the training split. To prevent these issues, practitioners should maintain temporal integrity in time-series data and fit all transformations exclusively on training data. Detection methods include monitoring for suspicious feature importance and evaluating models on independent datasets.
- Data leakage results in models that perform well on historical data but fail in real-world deployment.
- Train-test contamination often occurs when random shuffling is applied to time-series datasets.
- Preprocessing transformations like scaling must be fitted only on the training set to avoid leaking test distribution statistics.
- Target leakage happens when features include information about the outcome that is unavailable at prediction time.
- High feature importance can be a diagnostic signal for identifying leaked features.
- The MLOps crash course provides educational content on managing data leakage and production ML systems.