← 回到 Reading
Daily Dose of DS 2026-07-13

Agentic RL: Environments, Trajectories, and the Training Loop

Lightning AI has launched its AI Cloud, a platform that integrates GPUs, datacenter fabric, and hypervisors into a single system to avoid the limitations of traditional VM abstractions. By controlling the entire stack, the service provides deterministic multi-node placement and direct visibility into hardware topology for guest systems. This architecture ensures that performance-critical tools like NCCL and PyTorch function correctly without the need for manual tuning. The AI Cloud is available for both on-demand and spot capacity and is built on the same infrastructure as PyTorch Lightning. Part 12 of the Reinforcement Learning course focuses on establishing a complete training loop for LLM agents that utilize multi-step tool actions. It moves beyond prompt engineering by using trajectories as the primary training unit and implementing RL to reward or penalize specific behaviors. The curriculum introduces RULER for scoring and addresses technical challenges like the credit assignment problem in long-horizon episodes. A practical component demonstrates training a 3B parameter SQL agent on Google Colab using ART and RULER. Hermes agent introduces skill bundles to group multiple related skills into a single command using YAML configuration. This feature streamlines complex workflows, such as coding and testing, by loading all necessary tools and instructions simultaneously. Bundles are designed for resilience, cross-platform compatibility, and easy team sharing via Git repositories. They provide a mechanism for small teams to standardize agent behavior without performance overhead.

閱讀原文 ↗
目錄 3 段
  1. 01Instant H100 access. Fully on demand for AI teams.
  2. 02Agentic RL: Environments, trajectories, and training loop
  3. 03Skill bundles in Hermes agent
TOGETHER WITH LIGHTNING

Instant H100 access. Fully on demand for AI teams.

Lightning AI has launched its AI Cloud, a platform that integrates GPUs, datacenter fabric, and hypervisors into a single system to avoid the limitations of traditional VM abstractions. By controlling the entire stack, the service provides deterministic multi-node placement and direct visibility into hardware topology for guest systems. This architecture ensures that performance-critical tools like NCCL and PyTorch function correctly without the need for manual tuning. The AI Cloud is available for both on-demand and spot capacity and is built on the same infrastructure as PyTorch Lightning.

  • Lightning AI's AI Cloud integrates hardware and software layers, including the hypervisor and scheduler, into a unified system.
  • The platform allows guest systems to see the actual hardware topology rather than a flattened abstraction.
  • Deterministic multi-node placement is achieved through the scheduler's direct knowledge of the physical switch fabric.
  • The service eliminates intermediate leasing layers, providing a single point of accountability for hardware provisioning.
  • Users can access H100 GPUs through either guaranteed on-demand or spot capacity.
  • The infrastructure is designed to support NCCL and PyTorch without requiring manual performance tuning.
AI ENGINEERING

Agentic RL: Environments, trajectories, and training loop

Part 12 of the Reinforcement Learning course focuses on establishing a complete training loop for LLM agents that utilize multi-step tool actions. It moves beyond prompt engineering by using trajectories as the primary training unit and implementing RL to reward or penalize specific behaviors. The curriculum introduces RULER for scoring and addresses technical challenges like the credit assignment problem in long-horizon episodes. A practical component demonstrates training a 3B parameter SQL agent on Google Colab using ART and RULER.

  • Reinforcement Learning provides a mechanism to reward or penalize agent behaviors across thousands of rollouts, surpassing the limitations of prompt engineering.
  • Trajectories are defined as the fundamental unit of training for multi-step agentic workflows.
  • RULER is a specialized tool for scoring agent outputs that incorporates prefix deduplication.
  • The course addresses the credit assignment problem which occurs during long-horizon episodes in agentic tasks.
  • A 3B parameter SQL agent can be trained on a free Colab GPU using tool calls and a simple correctness-based reward signal.
HERMES

Skill bundles in Hermes agent

Hermes agent introduces skill bundles to group multiple related skills into a single command using YAML configuration. This feature streamlines complex workflows, such as coding and testing, by loading all necessary tools and instructions simultaneously. Bundles are designed for resilience, cross-platform compatibility, and easy team sharing via Git repositories. They provide a mechanism for small teams to standardize agent behavior without performance overhead.

  • Skill bundles group multiple skills and custom instructions under a single slash command using a YAML file.
  • The system is resilient to missing components; if one skill in a bundle is uninstalled, the others still load.
  • Skill bundles take precedence over individual skills in the event of a name collision.
  • Bundles are platform-agnostic, working across CLI, TUI, dashboards, Telegram, Discord, and Slack.
  • Team standardization is achieved by sharing bundle YAML files through Git and symlinking them to the local Hermes directory.
  • Invoking a bundle generates a fresh user message with no additional performance cost or cache invalidation issues.