Loop Engineering, Clearly Explained!
Doc Holiday is an automation tool designed to prevent knowledge degeneration by keeping engineering documentation in sync with code releases. It integrates directly into CI/CD pipelines and connects with various upstream sources like Jira, Slack, and Notion. When a pull request is merged, the tool analyzes commit history and linked tickets to automatically generate changelogs and release notes. This automation ensures that documentation updates do not require additional manual steps in the release process. AI agent development has shifted from basic prompting to loop engineering, which focuses on the systems wrapping the model's execution cycle. While the core loop of sending context and running tools is standardized across frameworks like LangGraph and Claude Code, the challenge lies in managing context rot and defining objective completion criteria. Effective loop engineering requires implementing external brakes like budget caps and no-progress detection to prevent doom loops. Ultimately, the goal is to move from steering an agent move-by-move to building a system that autonomously pursues a verifiable success criterion. MIT researchers have introduced Recursive Language Models (RLMs) to address "context rot," a phenomenon where LLM performance declines as conversation length increases. RLMs store context externally in a Python REPL and use tools like regex and recursive sub-calls to programmatically explore and process data in smaller chunks. This approach allows smaller models like GPT-5-mini to outperform larger ones on long-context tasks while maintaining efficiency at scales exceeding 10 million tokens.
閱讀原文 ↗目錄
Automated release docs for engineering teams
Doc Holiday is an automation tool designed to prevent knowledge degeneration by keeping engineering documentation in sync with code releases. It integrates directly into CI/CD pipelines and connects with various upstream sources like Jira, Slack, and Notion. When a pull request is merged, the tool analyzes commit history and linked tickets to automatically generate changelogs and release notes. This automation ensures that documentation updates do not require additional manual steps in the release process.
- Doc Holiday automates the creation of changelogs, release notes, and documentation updates.
- The tool integrates into existing CI/CD pipelines to trigger actions upon PR merges.
- It connects to multiple upstream sources including Notion, Jira, Slack, Zendesk, Confluence, and Google Docs.
- The system analyzes commit history, linked tickets, and connected specs to generate its output.
- The primary goal is to eliminate the gap between product functionality and company knowledge without adding manual overhead.
Loop engineering, clearly explained!
AI agent development has shifted from basic prompting to loop engineering, which focuses on the systems wrapping the model's execution cycle. While the core loop of sending context and running tools is standardized across frameworks like LangGraph and Claude Code, the challenge lies in managing context rot and defining objective completion criteria. Effective loop engineering requires implementing external brakes like budget caps and no-progress detection to prevent doom loops. Ultimately, the goal is to move from steering an agent move-by-move to building a system that autonomously pursues a verifiable success criterion.
- The basic agent loop is now a commodity, standardized across major frameworks like OpenAI Agents SDK and LangGraph.
- Models are poor judges of their own completion, often exiting loops before tasks are truly finished or verified.
- Context rot occurs as loops accumulate irrelevant data, leading to a doom loop where model performance degrades over time.
- Effective context management involves compaction, offloading large outputs, and delegating subtasks to separate agents.
- Reducing the number of available tools and ensuring they are non-overlapping significantly improves agent success rates.
- A separate verifier or model should be used to grade output, separating the maker from the checker to ensure objective quality.
Recursive language models
MIT researchers have introduced Recursive Language Models (RLMs) to address "context rot," a phenomenon where LLM performance declines as conversation length increases. RLMs store context externally in a Python REPL and use tools like regex and recursive sub-calls to programmatically explore and process data in smaller chunks. This approach allows smaller models like GPT-5-mini to outperform larger ones on long-context tasks while maintaining efficiency at scales exceeding 10 million tokens.
- Context rot occurs when LLMs lose reasoning and recall capabilities as the context window fills up.
- RLMs store context as a variable in a Python REPL environment rather than feeding it directly into the model's prompt.
- The RLM architecture uses tools like grep, partitioning, and recursive self-calls to decompose complex queries into manageable sub-tasks.
- GPT-5-mini using the RLM method outperformed the standard GPT-5 model on long-context benchmarks.
- RLMs maintain performance stability even at context lengths exceeding 10 million tokens.
- The RLM methodology mirrors the behavior of agentic coding tools like Claude Code, which selectively pull relevant snippets into context.