Kimi K3's Sandbox Problem Finally Has an Open-Source Fix
The cost of running AI agents is primarily determined by the runtime harness rather than the model itself, as the harness controls prompt assembly and call frequency. Redundant context, such as re-reading tool outputs in conversation history, significantly inflates token bills. To mitigate this, developers can use strategies like on-demand tool schema loading, result offloading, and subagent delegation. TrueFoundry's TrueForge harness demonstrates these efficiencies, achieving comparable performance to Claude Managed Agents at a fraction of the cost on the Enterprise-Bench benchmark. Training the Kimi K3 model required a massive infrastructure of 51 million sandboxes to handle workload isolation, native GPU access, and fast environment forking. While the Kimi team originally built a proprietary tool called AgentENV using Firecracker, smolvm has emerged as an open-source alternative that provides these three capabilities in a single binary. smolvm utilizes microVMs to ensure security boundaries while maintaining sub-200ms boot times and supporting portable environment artifacts. Active learning is a strategy for building supervised machine learning models when starting with unlabeled data. The process involves training an initial model on a small manual sample and then iteratively labeling the examples the model is least confident about. This approach optimizes human labeling effort by focusing on the most informative data points. A variation known as cooperative learning further enhances this by using high-confidence predictions as labels for additional training data.
閱讀原文 ↗目錄
The harness decides your token bill, not the model
The cost of running AI agents is primarily determined by the runtime harness rather than the model itself, as the harness controls prompt assembly and call frequency. Redundant context, such as re-reading tool outputs in conversation history, significantly inflates token bills. To mitigate this, developers can use strategies like on-demand tool schema loading, result offloading, and subagent delegation. TrueFoundry's TrueForge harness demonstrates these efficiencies, achieving comparable performance to Claude Managed Agents at a fraction of the cost on the Enterprise-Bench benchmark.
- Agent token costs are largely driven by the harness re-sending context in every turn.
- Efficiency can be improved by loading tool schemas only when needed and offloading large tool results to disk.
- Delegating tasks to subagents or running toolchains in code reduces the context burden on the root agent.
- TrueForge reduced token usage by nearly 70% and model calls by 40% compared to Claude Managed Agents.
- Using open models like GLM-5.2 with an efficient harness can further reduce costs while maintaining performance.
- TrueForge is an open-source, MIT-licensed agent harness that allows for model swapping and local execution.
Kimi K3’s sandbox problem finally has an open-source fix
Training the Kimi K3 model required a massive infrastructure of 51 million sandboxes to handle workload isolation, native GPU access, and fast environment forking. While the Kimi team originally built a proprietary tool called AgentENV using Firecracker, smolvm has emerged as an open-source alternative that provides these three capabilities in a single binary. smolvm utilizes microVMs to ensure security boundaries while maintaining sub-200ms boot times and supporting portable environment artifacts.
- Kimi K3 training utilized 51 million sandboxes across 1.5 million images to manage parallel agent trajectories.
- The three critical requirements for agent sandboxing are strong workload isolation, native GPU access, and fast environment forking.
- smolvm is an open-source VM runtime that provides hardware-level isolation, unlike containers which share the host kernel.
- smolvm supports native GPU access within the sandbox and allows forking a running environment's exact state.
- The runtime achieves boot times under 200ms and supports OCI images from Docker Hub without a Docker installation.
- Environments in smolvm can be packed into portable .smolmachine artifacts for consistent execution across different host systems.
Build ML models with little labeled data
Active learning is a strategy for building supervised machine learning models when starting with unlabeled data. The process involves training an initial model on a small manual sample and then iteratively labeling the examples the model is least confident about. This approach optimizes human labeling effort by focusing on the most informative data points. A variation known as cooperative learning further enhances this by using high-confidence predictions as labels for additional training data.
- Active learning reduces the time and cost of building supervised models by focusing human labeling on difficult examples.
- The iterative cycle involves training, predicting, measuring confidence, and human labeling of low-confidence samples.
- Probabilistic models are preferred in this context because they provide natural proxies for prediction confidence.
- Cooperative learning is a specific variant that utilizes high-confidence model predictions as labels to augment the training set.
- The success of active learning depends heavily on the accuracy of the confidence measures used to select samples for labeling.