How to Fine-Tune LLMs in 2026
Rowboat Spaces is an open-source platform that allows personal AI assistants to collaborate within a shared team environment. Each team member brings an assistant with access to their individual files, notes, and meetings into a common channel. Assistants can contribute specific knowledge to the group, with answers attributed to the respective user. The system also facilitates drafting and versioning shared documents based on the conversation history. Modern LLM fine-tuning is evolving from imitation-based Supervised Fine-Tuning (SFT) to Reinforcement Fine-Tuning (RFT) using algorithms like GRPO. This shift allows small open-source models to outperform much larger proprietary models by learning through trial and error rather than static datasets. Frameworks like ART and RULER automate the complex parts of this process, such as reward function engineering and trajectory evaluation. These tools enable developers to build agents that improve their reasoning and tool-use capabilities autonomously. The text explains how Grouped-query attention (GQA) optimizes KV cache size by sharing key and value heads among groups of query heads. While standard Multi-head attention (MHA) requires a unique KV head for every query head, GQA offers a middle ground between MHA and Multi-query attention (MQA). Using Llama 3 70B as an example, the article demonstrates how GQA can reduce the KV cache size and data read during decoding by 8x. This optimization is an architectural feature rather than a configuration setting that can be applied to existing MHA models.
閱讀原文 ↗目錄
A shared space for personal AI assistants
Rowboat Spaces is an open-source platform that allows personal AI assistants to collaborate within a shared team environment. Each team member brings an assistant with access to their individual files, notes, and meetings into a common channel. Assistants can contribute specific knowledge to the group, with answers attributed to the respective user. The system also facilitates drafting and versioning shared documents based on the conversation history.
- Rowboat Spaces is an open-source tool for collaborative AI assistant usage.
- Individual assistants maintain access to a user's private context like notes and meeting records.
- Team members can prompt their assistants to share information in a collective channel.
- The platform tracks the history of conversations to help draft and update shared files.
- Changes to shared documents are versioned and linked back to the messages that triggered them.
How to fine-tune LLMs in 2026
Modern LLM fine-tuning is evolving from imitation-based Supervised Fine-Tuning (SFT) to Reinforcement Fine-Tuning (RFT) using algorithms like GRPO. This shift allows small open-source models to outperform much larger proprietary models by learning through trial and error rather than static datasets. Frameworks like ART and RULER automate the complex parts of this process, such as reward function engineering and trajectory evaluation. These tools enable developers to build agents that improve their reasoning and tool-use capabilities autonomously.
- Reinforcement Fine-Tuning (RFT) is more effective than Supervised Fine-Tuning (SFT) for training agents that require multi-step reasoning and tool use.
- GRPO (Group Relative Policy Optimization) eliminates the need for a separate reward model by using relative rankings within groups of completions.
- ART (Agent Reinforcement Trainer) is an open-source framework that manages the training loop between agent environments and GRPO-based fine-tuning.
- RULER (Relative Universal LLM-Elicited Rewards) uses an LLM-as-judge to provide relative reward signals, removing the bottleneck of manual labeling.
- The ART backend leverages vLLM for high-speed inference and Unsloth for efficient training of LoRA checkpoints.
- Modern fine-tuning stacks allow models to master complex protocols like MCP (Model Context Protocol) through automated reinforcement cycles.
One architectural change can cut KV cache by 8x!
The text explains how Grouped-query attention (GQA) optimizes KV cache size by sharing key and value heads among groups of query heads. While standard Multi-head attention (MHA) requires a unique KV head for every query head, GQA offers a middle ground between MHA and Multi-query attention (MQA). Using Llama 3 70B as an example, the article demonstrates how GQA can reduce the KV cache size and data read during decoding by 8x. This optimization is an architectural feature rather than a configuration setting that can be applied to existing MHA models.
- Multi-head attention (MHA) stores separate key and value vectors for every query head at every layer.
- Multi-query attention (MQA) shares a single KV head across all query heads to minimize cache size, potentially at the cost of model quality.
- Grouped-query attention (GQA) divides query heads into groups that share a single KV head.
- Llama 3 70B uses GQA with eight query heads per KV head, resulting in an 8x smaller KV cache compared to MHA.
- GQA reduces the amount of KV data read during decoding by the same factor as the cache reduction.
- GQA is an architectural design choice and cannot be enabled on standard MHA models via serving flags.