Verifiable Rewards and GRPO in RL
Part 10 of the Reinforcement Learning course introduces the GRPO algorithm, which is the primary method behind the DeepSeek-R1 reasoning model. The section explores using verifiable rewards instead of learned reward models for tasks with checkable correctness, such as math and code. This approach simplifies the standard RLHF architecture from four models to two, significantly reducing memory overhead during training. It also demonstrates how complex reasoning behaviors like chain-of-thought can emerge purely through reinforcement learning without supervised demonstrations. Standard knowledge graphs used for agent memory often suffer from generic labeling and temporal staleness, leading to undifferentiated nodes and outdated context. Zep Graphiti addresses these issues by allowing developers to define structured memory schemas using Pydantic. The system utilizes contradiction detection and temporal annotations to invalidate old facts and track the validity of information over time. This ensures that retrieval processes filter for current, typed, and structured data before presenting it to the agent. Context engineering for Model Context Protocol (MCP) servers addresses the inefficiency of loading excessive tool definitions into an LLM's context. By using the open-source Bright Data MCP server, developers can scope tool access to specific groups or individual tools, reducing context token usage by up to 95%. Furthermore, the server optimizes tool outputs by stripping unnecessary markdown formatting, which saves additional tokens and improves agent focus. These optimizations enable agents to successfully interact with protected web data on platforms like LinkedIn and YouTube while avoiding common scraping obstacles.
閱讀原文 ↗目錄
Verifiable rewards and GRPO in RL
Part 10 of the Reinforcement Learning course introduces the GRPO algorithm, which is the primary method behind the DeepSeek-R1 reasoning model. The section explores using verifiable rewards instead of learned reward models for tasks with checkable correctness, such as math and code. This approach simplifies the standard RLHF architecture from four models to two, significantly reducing memory overhead during training. It also demonstrates how complex reasoning behaviors like chain-of-thought can emerge purely through reinforcement learning without supervised demonstrations.
- GRPO (Group Relative Policy Optimization) is the algorithm used by DeepSeek-R1 to achieve advanced reasoning capabilities.
- Verifiable rewards provide an exact, unhackable signal for tasks like math and formal logic, replacing model-based approximations of human judgment.
- The use of GRPO and verifiable rewards reduces the RLHF setup from four models (policy, reference, reward, critic) to just two.
- Chain-of-thought reasoning can emerge as a result of the RL process alone, without the need for supervised demonstrations.
- GRPO utilizes group-relative advantages, which makes the critic model optional when rewards are computationally cheap to calculate.
- The course includes practical training examples for math problems using the Unsloth tool.
Standard KG vs Zep’s temporal KG
Standard knowledge graphs used for agent memory often suffer from generic labeling and temporal staleness, leading to undifferentiated nodes and outdated context. Zep Graphiti addresses these issues by allowing developers to define structured memory schemas using Pydantic. The system utilizes contradiction detection and temporal annotations to invalidate old facts and track the validity of information over time. This ensures that retrieval processes filter for current, typed, and structured data before presenting it to the agent.
- Standard LLM extraction often defaults to generic labels like 'Object' and 'RELATES_TO' without specific schema guidance.
- Naive retrieval in standard knowledge graphs cannot distinguish between current and outdated facts, resulting in stale context.
- Zep Graphiti uses Pydantic to define entity types, edge types, and attributes for structured extraction.
- Contradiction detection in Zep Graphiti automatically invalidates outdated facts when new information is presented.
- Temporal annotations allow the system to track when facts were true and filter by validity during query time.
- Zep Graphiti is an open-source tool with over 28,000 stars on GitHub.
Context engineering for MCP servers
Context engineering for Model Context Protocol (MCP) servers addresses the inefficiency of loading excessive tool definitions into an LLM's context. By using the open-source Bright Data MCP server, developers can scope tool access to specific groups or individual tools, reducing context token usage by up to 95%. Furthermore, the server optimizes tool outputs by stripping unnecessary markdown formatting, which saves additional tokens and improves agent focus. These optimizations enable agents to successfully interact with protected web data on platforms like LinkedIn and YouTube while avoiding common scraping obstacles.
- Loading all tool definitions into an MCP client can consume over 10,000 tokens, causing wasted costs and reduced LLM accuracy.
- The Bright Data MCP server allows for scoping tools into logical groups like ECOMMERCE or SOCIAL_MEDIA to minimize context overhead.
- Dynamic construction of tool lists can reduce MCP tool context tokens by 78-95%.
- Optimizing tool outputs by stripping markdown formatting (headings, bolding, image syntax) can reduce token usage by 40-80%.
- The Bright Data MCP server utilizes the remark and strip-markdown libraries for automated content processing.
- Effective context engineering helps agents bypass anti-scraping measures, IP blocks, and captchas on complex websites.