11 LLM Evaluation Methods
Kimi K3 utilizes a novel mechanism called delta attention to manage context windows of up to a million tokens without the memory overhead of a standard KV cache. Unlike traditional attention that stores every key-value pair in a growing list, delta attention compresses history into a fixed-size matrix using a delta rule to update associations. This approach reduces the computational cost from quadratic to linear by writing only the difference between new data and existing memory. To balance efficiency with accuracy, production models interleave these linear delta attention layers with standard full-attention layers. LLM evaluation is a fragmented field because different metrics encode varying assumptions about correctness, ranging from n-gram overlap to semantic similarity. Traditional metrics like BLEU and ROUGE are being supplemented by model-based approaches such as G-Eval and LLM-as-Judge to handle paraphrasing and subjective criteria. Specialized evaluation methods also exist for agentic trajectories, multi-turn conversations, and safety gating. Tools like Comet Opik provide an integrated platform to implement these diverse metrics for production observability. Quantile regression is a technique used to estimate specific percentiles of a target variable's distribution, providing more context than traditional point estimates like the mean. By applying a weighted loss function, models can be biased to predict different quantiles, such as the 25th or 75th percentiles, to capture best-case and worst-case scenarios. This approach is particularly effective for understanding the spread of data and can be implemented using optimization tools or specialized machine learning libraries.
閱讀原文 ↗目錄
Delta attention in Kimi K3 to fix growing KV cache
Kimi K3 utilizes a novel mechanism called delta attention to manage context windows of up to a million tokens without the memory overhead of a standard KV cache. Unlike traditional attention that stores every key-value pair in a growing list, delta attention compresses history into a fixed-size matrix using a delta rule to update associations. This approach reduces the computational cost from quadratic to linear by writing only the difference between new data and existing memory. To balance efficiency with accuracy, production models interleave these linear delta attention layers with standard full-attention layers.
- Delta attention replaces the growing KV cache with a fixed-size matrix to prevent memory exhaustion in long-context sequences.
- The delta rule updates memory by reading the current state and writing only the 'delta' or difference between the guess and the target value.
- Standard attention scales quadratically with sequence length, while delta attention scales linearly.
- Fixed-size matrices in delta attention allow old entries to fade over time to accommodate new information.
- Because compressed matrices result in approximate recall, Kimi K3 interleaves delta attention with standard attention layers for exact lookup where needed.
11 LLM evaluation methods
LLM evaluation is a fragmented field because different metrics encode varying assumptions about correctness, ranging from n-gram overlap to semantic similarity. Traditional metrics like BLEU and ROUGE are being supplemented by model-based approaches such as G-Eval and LLM-as-Judge to handle paraphrasing and subjective criteria. Specialized evaluation methods also exist for agentic trajectories, multi-turn conversations, and safety gating. Tools like Comet Opik provide an integrated platform to implement these diverse metrics for production observability.
- BLEU and ROUGE are n-gram based metrics that often fail to account for valid paraphrasing, making them unsuitable for creative generation.
- BERTScore addresses the paraphrase problem by using token embeddings to measure semantic similarity rather than exact word matches.
- Model-based evaluation (LLM-as-Judge) requires controls for length bias and position bias to ensure objective results.
- LLM juries utilize multiple model families to mitigate the inherent biases of a single judge model.
- Safety evaluations should function as binary gates rather than being averaged into a general quality score to prevent shipping violations.
- Trajectory accuracy is essential for agentic systems to ensure the model reaches the correct answer through a valid logical path.
- Comet Opik is an open-source platform that implements many of these evaluation metrics for production data.
Quantile regression
Quantile regression is a technique used to estimate specific percentiles of a target variable's distribution, providing more context than traditional point estimates like the mean. By applying a weighted loss function, models can be biased to predict different quantiles, such as the 25th or 75th percentiles, to capture best-case and worst-case scenarios. This approach is particularly effective for understanding the spread of data and can be implemented using optimization tools or specialized machine learning libraries.
- Traditional regression models like OLS typically estimate the mean value, which may not represent the full distribution of the outcome.
- Quantile regression estimates the conditional percentiles of the response variable based on input features.
- The method utilizes a parameterized loss function where a weight parameter 'w' determines the specific quantile being targeted.
- A weight of w=0.5 in quantile regression corresponds to predicting the median of the distribution.
- Scipy's minimize method can be used to find optimal parameters for the quantile loss objective function.
- LightGBM natively supports quantile objective functions, making it suitable for tree-based quantile modeling.