← 回到 Reading
Daily Dose of DS 2026-07-21

5 LLM Quantization Techniques

AI dependencies such as open-source models and MCP servers often bypass traditional dependency scanners and lockfiles, leading to a lack of governance in production. Checkmarx AI Inventory, a component of the Checkmarx One platform, addresses this by deterministically cataloging these components and tracing them to specific lines of code. This visibility allows organizations to generate AI-BOMs for compliance audits and security oversight. Gartner has recognized Checkmarx's efforts in this space by naming them a Leader in the Software Supply Chain Security category. Quantization is a critical technique for reducing the memory footprint of large language models, enabling them to run on hardware with limited VRAM. The process involves mapping high-precision weights (FP16) to lower-bit integers (INT4 or INT8), but naive rounding is often disrupted by outlier features present in models larger than 6.7B parameters. Five primary methods—RTN, GPTQ, AWQ, LLM.int8(), and QAT—offer different strategies for managing these outliers during rounding, inference, or training to maintain model accuracy. AI systems require observability practices similar to traditional software, utilizing traces and spans to monitor internal pipeline steps. In a RAG pipeline, individual spans track specific operations like embedding, retrieval, and generation to capture metrics such as latency, token counts, and costs. A unique Trace ID links these spans together, providing a complete view of a single request's journey. This granular visibility is essential for debugging failures, tracking expenses, and identifying performance drift in production LLM applications.

閱讀原文 ↗
目錄 3 段
  1. 01Trace AI components in production to their exact line
  2. 025 LLM Quantization Techniques
  3. 03Layers of observability in AI systems
TOGETHER WITH CHECKMARX

Trace AI components in production to their exact line

AI dependencies such as open-source models and MCP servers often bypass traditional dependency scanners and lockfiles, leading to a lack of governance in production. Checkmarx AI Inventory, a component of the Checkmarx One platform, addresses this by deterministically cataloging these components and tracing them to specific lines of code. This visibility allows organizations to generate AI-BOMs for compliance audits and security oversight. Gartner has recognized Checkmarx's efforts in this space by naming them a Leader in the Software Supply Chain Security category.

  • AI components like LLM APIs and agents often lack representation in standard manifest files and lockfiles.
  • A Checkmarx study found that 43% of teams have zero governance over AI components in production.
  • Checkmarx AI Inventory provides deterministic cataloging of models, SDKs, and MCP servers.
  • The tool can trace AI dependencies to the exact file and line of code where they are implemented.
  • Checkmarx AI Inventory generates an AI-BOM (AI Bill of Materials) to assist with compliance audits.
  • Gartner named Checkmarx a Leader in its inaugural Magic Quadrant for Software Supply Chain Security.
LLMs

5 LLM Quantization Techniques

Quantization is a critical technique for reducing the memory footprint of large language models, enabling them to run on hardware with limited VRAM. The process involves mapping high-precision weights (FP16) to lower-bit integers (INT4 or INT8), but naive rounding is often disrupted by outlier features present in models larger than 6.7B parameters. Five primary methods—RTN, GPTQ, AWQ, LLM.int8(), and QAT—offer different strategies for managing these outliers during rounding, inference, or training to maintain model accuracy.

  • A 70B parameter model in FP16 requires 140GB of VRAM, which exceeds the capacity of a single H100 GPU.
  • Models larger than 6.7B parameters develop outlier dimensions with values 20x to 100x larger than average, which can break standard quantization.
  • GPTQ reduces rounding error by using calibration data to adjust remaining weights during the quantization process.
  • AWQ protects important weights by scaling them before rounding and is widely used in serving engines like vLLM.
  • LLM.int8() handles outliers by running them in FP16 while processing 99.9% of other values in INT8.
  • Quantization-Aware Training (QAT) simulates quantization damage during fine-tuning to let the model adapt its weights accordingly.
  • Google's Gemma 3 QAT checkpoints demonstrate that 4-bit models can maintain near full-precision quality with 3x lower memory.
LLM OBSERVABILITY

Layers of observability in AI systems

AI systems require observability practices similar to traditional software, utilizing traces and spans to monitor internal pipeline steps. In a RAG pipeline, individual spans track specific operations like embedding, retrieval, and generation to capture metrics such as latency, token counts, and costs. A unique Trace ID links these spans together, providing a complete view of a single request's journey. This granular visibility is essential for debugging failures, tracking expenses, and identifying performance drift in production LLM applications.

  • Observability in AI systems is modeled using Traces for the entire request journey and Spans for individual operations.
  • A standard RAG pipeline consists of Query, Embedding, Retrieval, Context, and Generation spans.
  • Retrieval spans are critical for identifying issues like bad chunks, low relevance scores, or incorrect top-k values.
  • Generation spans are typically the most expensive and time-consuming, requiring tracking of input/output tokens and latency.
  • Trace IDs are used to link multiple spans to a single user request for end-to-end debugging.
  • Span-level metrics allow developers to catch model drift and tune individual components independently.
  • Opik is an open-source tool that implements these LLM observability and tracing standards.