The Four Types of Agent Loops
A new fine-tuning studio has been developed that allows users to fine-tune Large Language Models directly through the Claude interface. The application integrates with the Hugging Face Hub for model and dataset discovery and utilizes Hugging Face's AutoTrain infrastructure for GPU-based training. It is built using the mcp-use SDK, an open-source framework that enables developers to create Model Context Protocol (MCP) applications with integrated React-based user interfaces. This setup allows for direct configuration of training parameters like LoRA rank and learning rate, as well as post-training interaction with the models. Loop engineering is the practice of designing systems to steer AI agents by defining their triggers and completion criteria. The text outlines four specific loop structures: turn-based, goal-based, time-based, and proactive. Each structure represents a different level of automation, shifting responsibility from humans to the system based on the task's nature. Choosing the correct loop depends on whether the work is exploratory, measurable, recurring, or standing. NVIDIA and MIT researchers introduced SparDA, a transformer architecture modification that adds a fourth projection called 'Forecast' to each layer. This projection predicts the KV blocks required by the subsequent layer, enabling asynchronous prefetching from CPU memory to GPU while the current layer computes. This approach mitigates the latency of KV cache offloading and simplifies block selection in sparse attention, leading to significant speedups in decoding and improved long-context accuracy.
閱讀原文 ↗目錄
Fine-tune any LLM directly from Claude!
A new fine-tuning studio has been developed that allows users to fine-tune Large Language Models directly through the Claude interface. The application integrates with the Hugging Face Hub for model and dataset discovery and utilizes Hugging Face's AutoTrain infrastructure for GPU-based training. It is built using the mcp-use SDK, an open-source framework that enables developers to create Model Context Protocol (MCP) applications with integrated React-based user interfaces. This setup allows for direct configuration of training parameters like LoRA rank and learning rate, as well as post-training interaction with the models.
- The fine-tuning studio enables LLM training and configuration directly from within Claude.
- Training execution is handled by Hugging Face's AutoTrain on their GPU infrastructure.
- The application is built on the mcp-use SDK, a framework for building full-stack MCP Apps for Agents.
- mcp-use allows for the association of MCP tools with custom React UI components.
- The framework handles complex tasks like tool registration, prop mapping, and hot reloading.
- The studio follows the MCP Apps standard, which draws inspiration from OpenAI’s Apps SDK.
The four types of agent loops
Loop engineering is the practice of designing systems to steer AI agents by defining their triggers and completion criteria. The text outlines four specific loop structures: turn-based, goal-based, time-based, and proactive. Each structure represents a different level of automation, shifting responsibility from humans to the system based on the task's nature. Choosing the correct loop depends on whether the work is exploratory, measurable, recurring, or standing.
- Loop engineering automates the start and end conditions of an agent's execution.
- Turn-based loops are designed for exploratory tasks where human feedback is needed after every step.
- Goal-based loops utilize an evaluator model to check success criteria against a defined budget.
- Time-based loops automate recurring tasks using clock-based triggers and can be scheduled in the cloud.
- Proactive loops handle standing responsibilities by triggering multi-agent workflows in response to events.
- The progression of loop types reduces the need for human monitoring as more control is handed to the system.
NVIDIA researchers built a new transformer variant
NVIDIA and MIT researchers introduced SparDA, a transformer architecture modification that adds a fourth projection called 'Forecast' to each layer. This projection predicts the KV blocks required by the subsequent layer, enabling asynchronous prefetching from CPU memory to GPU while the current layer computes. This approach mitigates the latency of KV cache offloading and simplifies block selection in sparse attention, leading to significant speedups in decoding and improved long-context accuracy.
- SparDA adds a Forecast projection to the standard Q, K, and V projections to predict future KV block needs.
- The architecture enables prefetching KV blocks from CPU RAM on a separate CUDA stream, hiding transfer latency.
- Block selection costs are reduced by using one Forecast head per GQA group and removing the per-query-head scoring loop.
- SparDA achieves up to 1.7x faster decoding and 5.3x higher throughput compared to sparse offload baselines.
- The modification adds only 0.41% more parameters (33.5M) to an 8B model.
- Long-reasoning accuracy improved by 6.5 points on the NOSA-8B model.