← 回到 Reading
Daily Dose of DS 2026-09-09

Your Agent Harness Needs Runtime Security

Agent Beacon is an open-source telemetry layer designed to provide runtime security for AI agents by recording and normalizing their activities into a structured system of record. It addresses a critical visibility gap identified in security incidents at major AI labs where agents bypassed traditional log scanners using packed payloads. The tool integrates with over 23 agent harnesses, enabling real-time monitoring, local detection rules, and data forwarding to existing security platforms like Splunk and Datadog. KV cache management at scale requires treating cached tensors as a persistent storage system rather than temporary data. LMCache is an open-source standalone service that manages these caches across GPU, CPU, and remote storage tiers, decoupling them from the inference engine. By utilizing CUDA IPC for zero-copy memory access, LMCache allows multiple inference workers to share cached prompt prefixes efficiently. This architecture significantly improves throughput and ensures cache persistence even during engine restarts.

閱讀原文 ↗
目錄 2 段
  1. 01Your Agent harness needs runtime security
  2. 02How do LLMs handle GBs of KV cache in production?
HANDS-ON

Your Agent harness needs runtime security

Agent Beacon is an open-source telemetry layer designed to provide runtime security for AI agents by recording and normalizing their activities into a structured system of record. It addresses a critical visibility gap identified in security incidents at major AI labs where agents bypassed traditional log scanners using packed payloads. The tool integrates with over 23 agent harnesses, enabling real-time monitoring, local detection rules, and data forwarding to existing security platforms like Splunk and Datadog.

  • Agent Beacon provides a 100% open-source telemetry layer that records tool calls, shell commands, and file changes in real-time.
  • The tool normalizes activity from 23+ different agent harnesses into a single, consistent schema for unified investigation.
  • It introduces an 'event.fidelity' field to distinguish between directly observed runtime actions and those inferred from log patterns.
  • Beacon includes a local scanner that evaluates YAML-based detection rules against agent behavior as it occurs.
  • Security incidents at OpenAI, Anthropic, and xAI demonstrate that traditional scanners often fail to interpret encoded payloads executed by model runtimes.
  • The system is local-by-default, requiring no external account, but supports forwarding events to SIEMs like Microsoft Sentinel and CrowdStrike.
LLMs

How do LLMs handle GBs of KV cache in production?

KV cache management at scale requires treating cached tensors as a persistent storage system rather than temporary data. LMCache is an open-source standalone service that manages these caches across GPU, CPU, and remote storage tiers, decoupling them from the inference engine. By utilizing CUDA IPC for zero-copy memory access, LMCache allows multiple inference workers to share cached prompt prefixes efficiently. This architecture significantly improves throughput and ensures cache persistence even during engine restarts.

  • KV cache becomes a complex storage system at scale, requiring management across multiple memory and storage tiers.
  • LMCache is a standalone open-source service that handles KV cache storage, movement, and cleanup.
  • Using CUDA IPC allows LMCache and inference engines to share GPU memory without expensive data copying.
  • Multiple vLLM instances can connect to a single LMCache service to share cached data.
  • LMCache can achieve up to 15x higher throughput by reusing prompt prefixes across different requests.
  • Decoupling the cache from the inference engine allows cached data to survive engine restarts.