Where does all the VRAM go during LLM inference?
The Nebius AI Builder Program is a newly launched initiative designed to support developers by providing over $400 in free credits and discounts. Participants gain access to a variety of tools in the open AI ecosystem, including Nebius Token Factory, Tavily, Toloka, and LangSmith. The program also offers educational resources such as runnable cookbooks and training from industry leaders to help developers build production-ready AI applications. By offering these resources for free, the program aims to facilitate the transition for developers into AI engineering roles. LLM inference VRAM usage is divided into four main categories: model weights, KV cache, temporary activations/workspace, and runtime overhead. While weights are relatively fixed based on parameter count and precision, the KV cache grows dynamically with context length and batch size, often exceeding weight memory in long-context scenarios. Temporary activations peak during the prefill phase, and runtime overhead includes CUDA contexts and memory fragmentation. Successful deployment requires planning for the total workload memory, not just the model size, including a safety margin for fragmentation. The Model Context Protocol (MCP) has introduced a standardized method for agents to discover and load "Skills," which represent reusable workflows or playbooks. These skills are served through MCP's existing Resources primitive, allowing agents to fetch specific SKILL.md files and supporting assets only when needed. This approach improves context window management by avoiding the need to load all workflow instructions upfront. By combining capabilities with the playbooks for using them, MCP serves as both a connection layer and a distribution mechanism for agent know-how.
閱讀原文 ↗目錄
Get started with all the best tools in the open AI ecosystem for free
The Nebius AI Builder Program is a newly launched initiative designed to support developers by providing over $400 in free credits and discounts. Participants gain access to a variety of tools in the open AI ecosystem, including Nebius Token Factory, Tavily, Toloka, and LangSmith. The program also offers educational resources such as runnable cookbooks and training from industry leaders to help developers build production-ready AI applications. By offering these resources for free, the program aims to facilitate the transition for developers into AI engineering roles.
- The Nebius AI Builder Program provides $400+ in free credits and discounts for AI development.
- Participating tools include Nebius Token Factory, Tavily, Toloka, and LangSmith.
- Educational resources include runnable cookbooks and training from top industry companies.
- Developers can learn to build deep research agents and search skills using models like Kimi K3.
- The program aims to lower the barrier to entry for building production-grade AI applications.
Where does all the VRAM go during LLM inference?
LLM inference VRAM usage is divided into four main categories: model weights, KV cache, temporary activations/workspace, and runtime overhead. While weights are relatively fixed based on parameter count and precision, the KV cache grows dynamically with context length and batch size, often exceeding weight memory in long-context scenarios. Temporary activations peak during the prefill phase, and runtime overhead includes CUDA contexts and memory fragmentation. Successful deployment requires planning for the total workload memory, not just the model size, including a safety margin for fragmentation.
- VRAM for LLM inference consists of four buckets: weights, KV cache, temporary activations/workspace, and runtime overhead.
- Model weights are fixed based on parameter count and numerical precision, requiring approximately 2 bytes per parameter for FP16/BF16.
- The KV cache scales linearly with context length, batch size, and the number of active sequences, often catching deployments by surprise.
- Grouped-query attention (GQA) and KV-cache quantization are specific architectural and optimization methods used to reduce the KV cache footprint.
- Temporary memory for activations peaks during the prefill phase and is reused across layers, unlike the KV cache which accumulates per token.
- Runtime overhead includes CUDA contexts and fragmentation, which can cause 'nvidia-smi' to report higher usage than actual live tensors.
- Workload capacity planning must account for the sum of all memory buckets plus a safety margin to avoid out-of-memory errors.
MCP meets agent skills
The Model Context Protocol (MCP) has introduced a standardized method for agents to discover and load "Skills," which represent reusable workflows or playbooks. These skills are served through MCP's existing Resources primitive, allowing agents to fetch specific SKILL.md files and supporting assets only when needed. This approach improves context window management by avoiding the need to load all workflow instructions upfront. By combining capabilities with the playbooks for using them, MCP serves as both a connection layer and a distribution mechanism for agent know-how.
- MCP now defines a standard for discovering and loading Agent Skills directly from servers.
- Skills are implemented using the MCP Resources primitive, typically involving a SKILL.md file and supporting scripts or examples.
- On-demand loading of skills helps manage the agent's context window by only pulling in relevant workflow instructions for the current task.
- The protocol distinguishes between tools (actions), resources (data access), and skills (reusable workflows).
- MCP integrates with major agentic frameworks including LangGraph, LlamaIndex, CrewAI, and PydanticAI.
- The new standard allows workflow knowledge to be distributed and versioned alongside the server that provides the underlying tools and resources.