Karpathy's Agentic Engineering Finally Has Proper Tooling
Andrej Karpathy defines agentic engineering as a disciplined approach to production-grade AI agents, moving beyond informal vibe coding. Google's new Agents CLI addresses the tooling gap by providing a unified workflow for scaffolding, evaluating, and deploying agents using the Agent Development Kit (ADK). The tool integrates with popular coding agents like Claude Code and Cursor, allowing developers to manage the entire lifecycle through natural language prompts. Key features include automated evaluation suites with LLM-as-judge scoring and seamless deployment to Google Cloud's Agent Runtime. KV cache compression often fails to reduce memory usage in production environments like vLLM because PagedAttention manages memory in fixed blocks that are only freed when entirely empty. Standard eviction methods leave scattered tokens that keep these blocks occupied, and fast kernels like FlashAttention do not provide the attention scores needed for eviction decisions. NVIDIA's TriAttention addresses these challenges by using vector geometry for token scoring and a periodic compaction pass to consolidate the cache. This approach allows for significant memory savings and faster decoding without sacrificing model accuracy.
閱讀原文 ↗目錄
Karpathy's agentic engineering finally has proper tooling
Andrej Karpathy defines agentic engineering as a disciplined approach to production-grade AI agents, moving beyond informal vibe coding. Google's new Agents CLI addresses the tooling gap by providing a unified workflow for scaffolding, evaluating, and deploying agents using the Agent Development Kit (ADK). The tool integrates with popular coding agents like Claude Code and Cursor, allowing developers to manage the entire lifecycle through natural language prompts. Key features include automated evaluation suites with LLM-as-judge scoring and seamless deployment to Google Cloud's Agent Runtime.
- Agentic engineering separates production-grade agent development from informal vibe coding through spec design and eval loops.
- Google’s Agents CLI automates the agent lifecycle by injecting seven specific skills into existing coding agents like Cursor and Claude Code.
- The CLI supports the Agent Development Kit (ADK) patterns for building RAG agents with built-in citation and vector search capabilities.
- Automated evaluation tools in the CLI use LLM-as-judge scoring to detect hallucinations and verify citation accuracy.
- Agents can be deployed to Google Cloud's Agent Runtime and registered with Gemini Enterprise for organization-wide discovery.
- Observability is integrated by default through Cloud Trace for all deployed requests.
A tricky LLM interview question on KV cache compression
KV cache compression often fails to reduce memory usage in production environments like vLLM because PagedAttention manages memory in fixed blocks that are only freed when entirely empty. Standard eviction methods leave scattered tokens that keep these blocks occupied, and fast kernels like FlashAttention do not provide the attention scores needed for eviction decisions. NVIDIA's TriAttention addresses these challenges by using vector geometry for token scoring and a periodic compaction pass to consolidate the cache. This approach allows for significant memory savings and faster decoding without sacrificing model accuracy.
- PagedAttention requires an entire block of tokens to be empty before memory is returned to the GPU allocator.
- Importance-based eviction typically leaves survivor tokens scattered across blocks, preventing the memory allocator from reclaiming space.
- FlashAttention kernels do not store attention scores, making it difficult for eviction algorithms to identify which tokens to drop without performance loss.
- NVIDIA's TriAttention method identifies important tokens using the geometry of key and query vectors before RoPE is applied.
- TriAttention uses a compaction pass every 128 tokens to consolidate survivors and free up physical memory blocks.
- TriAttention achieves a 10.7x reduction in KV memory and 2.5x faster decoding while maintaining accuracy.