← 回到 Reading
Daily Dose of DS 2026-08-20

What is (was?) GIL in Python?

Retrieval-Augmented Generation (RAG) can be inefficient when repeatedly querying static data from a vector database. Cache-Augmented Generation (CAG) optimizes this by storing static, 'cold' information in the model's internal key-value (KV) memory. By combining both approaches, systems can achieve faster inference and lower costs while maintaining access to dynamic, 'hot' data. Advanced techniques like CacheBlend, implemented in the open-source LMCache, allow for more flexible caching that overcomes the strict prefix-matching limitations of standard prompt caching. The Global Interpreter Lock (GIL) has historically limited Python's multi-threading performance by restricting execution to a single thread per process. While multi-processing offers a workaround by using isolated memory spaces, it introduces complexity compared to shared-memory threading. Python 3.14 introduces the ability to disable the GIL, enabling true parallel execution across multiple CPU cores to improve performance for compute-intensive tasks. Contrastive learning is a self-supervised technique that trains models to learn representations by comparing samples rather than using traditional classification. In applications like face unlock systems, Siamese networks map inputs into a shared embedding space where similar items are positioned close together and dissimilar items are far apart. This approach avoids the need for retraining on-device when new users are added, as it relies on comparing generated embeddings against stored ones. The training process is driven by a contrastive loss function that optimizes the distance between these embeddings based on a defined margin.

閱讀原文 ↗
目錄 3 段
  1. 01RAG vs. CAG, explained visually!
  2. 02What is (was?) GIL in Python?
  3. 03What is Contrastive Learning?
OPEN-SOURCE

RAG vs. CAG, explained visually!

Retrieval-Augmented Generation (RAG) can be inefficient when repeatedly querying static data from a vector database. Cache-Augmented Generation (CAG) optimizes this by storing static, 'cold' information in the model's internal key-value (KV) memory. By combining both approaches, systems can achieve faster inference and lower costs while maintaining access to dynamic, 'hot' data. Advanced techniques like CacheBlend, implemented in the open-source LMCache, allow for more flexible caching that overcomes the strict prefix-matching limitations of standard prompt caching.

  • RAG is often redundant for static information that does not change frequently.
  • CAG stores static data in the model's KV memory to reduce repeated computation and database hits.
  • Standard prompt caching in APIs like OpenAI and Anthropic requires exact byte-for-byte prefix matching to be effective.
  • CacheBlend is a technique that allows for reusing cached document blocks even when their order or combination changes.
  • LMCache is an open-source implementation of these caching optimizations that can speed up multi-document queries by 2x to 4x.
  • Effective caching strategy involves separating 'cold' static data from 'hot' dynamic data to avoid hitting context limits.
PYTHON

What is (was?) GIL in Python?

The Global Interpreter Lock (GIL) has historically limited Python's multi-threading performance by restricting execution to a single thread per process. While multi-processing offers a workaround by using isolated memory spaces, it introduces complexity compared to shared-memory threading. Python 3.14 introduces the ability to disable the GIL, enabling true parallel execution across multiple CPU cores to improve performance for compute-intensive tasks.

  • The GIL prevents Python processes from executing multiple threads simultaneously on different CPU cores.
  • Multi-threading in Python with the GIL results in performance similar to single-threading for CPU-bound tasks.
  • Multi-processing achieves parallelism by using separate memory spaces for each process, avoiding the GIL's limitations.
  • The primary historical reason for the GIL was to ensure thread safety and prevent race conditions in shared memory.
  • Python 3.14 introduces a feature to run without the GIL, allowing threads to utilize multiple CPU cores fully.
  • Inter-process communication (IPC) mechanisms like pipes and queues are necessary for data sharing in multi-processing but add complexity.
machine learning

What is Contrastive Learning?

Contrastive learning is a self-supervised technique that trains models to learn representations by comparing samples rather than using traditional classification. In applications like face unlock systems, Siamese networks map inputs into a shared embedding space where similar items are positioned close together and dissimilar items are far apart. This approach avoids the need for retraining on-device when new users are added, as it relies on comparing generated embeddings against stored ones. The training process is driven by a contrastive loss function that optimizes the distance between these embeddings based on a defined margin.

  • Contrastive learning teaches models to learn useful representations by comparing similar and dissimilar samples.
  • Siamese networks use a shared network to map inputs to a common embedding space for distance-based comparison.
  • Binary classification is often unsuitable for face unlock due to the difficulty of obtaining diverse negative samples and the risk of catastrophic forgetting.
  • The contrastive loss function minimizes the distance between embeddings of the same class and maximizes the distance for different classes up to a specific margin.
  • On-device face unlock systems can support multiple users by storing unique embeddings in memory without requiring model retraining.
  • Model compression and federated learning are complementary techniques for building efficient, privacy-preserving on-device ML applications.