Continuous Batching in LLMs
Google's Agents CLI streamlines the agentic engineering lifecycle by consolidating scaffolding, deployment, security, and evaluation into a natural-language interface. It allows developers to manage the transition from a local idea to a governed enterprise asset using simple prompts. The tool integrates with existing coding agents like Claude Code and Cursor to automate complex tasks such as identity provisioning and prompt injection screening. By leveraging Agent Runtime and Gemini Enterprise, it ensures agents are stateful, secure, and accessible across an organization. Continuous batching, also known as iteration-level scheduling, optimizes LLM inference by rebuilding the batch at every forward pass rather than waiting for all requests in a static batch to complete. This approach allows finished requests to exit and waiting requests to enter immediately, preventing GPU resources from being wasted on empty slots caused by varying output lengths. Serving engines like vLLM and TensorRT-LLM use this method to significantly increase throughput, often managing complex tasks like chunked prefill and KV cache allocation within a unified scheduling loop.
閱讀原文 ↗Karpathy’s agentic engineering lifecycle, clearly explained
Google's Agents CLI streamlines the agentic engineering lifecycle by consolidating scaffolding, deployment, security, and evaluation into a natural-language interface. It allows developers to manage the transition from a local idea to a governed enterprise asset using simple prompts. The tool integrates with existing coding agents like Claude Code and Cursor to automate complex tasks such as identity provisioning and prompt injection screening. By leveraging Agent Runtime and Gemini Enterprise, it ensures agents are stateful, secure, and accessible across an organization.
- Google's Agents CLI automates the end-to-end lifecycle of AI agents using natural language prompts.
- The CLI integrates with popular coding tools including Claude Code, Cursor, Codex, and Antigravity.
- Security features include Model Armor for prompt injection screening and the provisioning of least-privilege identities.
- Deployment to Agent Runtime enables state persistence through Sessions and Memory Bank features.
- The evaluation stage specifically targets grounding, hallucination prevention, and prompt optimization.
- Finalized agents can be registered and published directly into Gemini Enterprise for organizational use.
Continuous batching in LLMs
Continuous batching, also known as iteration-level scheduling, optimizes LLM inference by rebuilding the batch at every forward pass rather than waiting for all requests in a static batch to complete. This approach allows finished requests to exit and waiting requests to enter immediately, preventing GPU resources from being wasted on empty slots caused by varying output lengths. Serving engines like vLLM and TensorRT-LLM use this method to significantly increase throughput, often managing complex tasks like chunked prefill and KV cache allocation within a unified scheduling loop.
- Static batching is inefficient for LLMs because output lengths are unpredictable, causing the entire batch to run at the speed of the slowest request.
- Continuous batching updates batch membership at every iteration, allowing for immediate replacement of completed sequences.
- Selective batching flattens tokens into a single sequence for linear operations while splitting them for attention to handle unique KV caches.
- The vLLM scheduler manages a token budget and sequence cap to balance throughput against inter-token latency.
- Preemption occurs when KV cache memory is exhausted, forcing requests to be recomputed from scratch in default vLLM configurations.
- Anyscale benchmarks demonstrated that vLLM can achieve 23x the throughput of naive Hugging Face serving on OPT-13B models.
- Key metrics for monitoring scheduler health include the total_cumulative_preemption_cnt in Prometheus.