← 回到 Reading
Berkeley RDI 2026-08-26

Agentic AI Weekly | Berkeley RDI | August 26, 2026

The focus of artificial intelligence is shifting from merely scaling model size toward creating systems capable of acting, learning, and adapting over long horizons in complex environments. Discussions spanning agentic foundational capabilities, robotics, and world models examine how autonomous systems can reliably operate on real-world data and act within the physical world. This evolution highlights a broader paradigm shift where scaling encompasses complete system capabilities rather than solely the underlying model. Recent AI progress has relied primarily on scaling compute, data, and model parameters, but researchers are increasingly shifting attention to optimizing the broader system surrounding the model. Industry experts argue that the next major advancements will stem from architectural components such as memory, context compression, evaluation, and tool integration. Additionally, enterprise deployment highlights the importance of proprietary, dynamic data that cannot be stored directly within model weights. Consequently, the focus is transitioning from standalone model scaling toward full-system optimization and agent-level engineering. As AI agents handle longer-horizon tasks, expanding their flexibility and tool access significantly increases their attack surface and operational risk. Addressing this requires an ecosystem-level approach to AI resilience, integrating automated red teaming, least-privilege architecture, and defense-in-depth from the start. Furthermore, rigorous evaluation is a prerequisite for recursive self-improvement, as autonomous systems cannot effectively improve without accurately measuring progress. Consequently, evaluation and resilience must be treated as foundational architecture rather than afterthoughts as autonomy scales.

閱讀原文 ↗
目錄 7 段
  1. 01From Longer-Horizon Agents to Physical Intelligence
  2. 021. Scaling Is Moving Beyond the Model
  3. 032. Longer Autonomy Requires Stronger Evaluation and Resilience
  4. 043. Physical AI Needs Its Own Scaling Laws
  5. 054. The Real Frontier May Be the Learning Loop
  6. 06The Bigger Picture
  7. 07Trends This Week

From Longer-Horizon Agents to Physical Intelligence

The focus of artificial intelligence is shifting from merely scaling model size toward creating systems capable of acting, learning, and adapting over long horizons in complex environments. Discussions spanning agentic foundational capabilities, robotics, and world models examine how autonomous systems can reliably operate on real-world data and act within the physical world. This evolution highlights a broader paradigm shift where scaling encompasses complete system capabilities rather than solely the underlying model.

  • AI development is moving beyond scaling model parameters toward systems that learn and adapt across long time horizons.
  • Autonomous agents require capabilities to improve through experience and handle real-world data reliably.
  • A key objective is enabling agents to reason and perform physical actions through robotics and world models.
  • The next scaling paradigm focuses on moving beyond the individual model to entire systemic workflows.

1. Scaling Is Moving Beyond the Model

Recent AI progress has relied primarily on scaling compute, data, and model parameters, but researchers are increasingly shifting attention to optimizing the broader system surrounding the model. Industry experts argue that the next major advancements will stem from architectural components such as memory, context compression, evaluation, and tool integration. Additionally, enterprise deployment highlights the importance of proprietary, dynamic data that cannot be stored directly within model weights. Consequently, the focus is transitioning from standalone model scaling toward full-system optimization and agent-level engineering.

  • Jerry Tworek predicts that architectural breakthroughs will drive AI gains over the next two years rather than raw model scaling alone.
  • Oriol Vinyals advocates for expanding model architecture to include memory, context compression, evaluation frameworks, tools, and agent scaffolding.
  • Weizhu Chen emphasizes optimizing surrounding infrastructure and systems before focusing directly on model optimization.
  • Dan Roth notes that real-world deployment in healthcare, finance, and government depends on dynamic, proprietary data that cannot reside solely in model weights.
  • Data governance, quality, and structure are identified as primary determinants of agentic AI capabilities.

2. Longer Autonomy Requires Stronger Evaluation and Resilience

As AI agents handle longer-horizon tasks, expanding their flexibility and tool access significantly increases their attack surface and operational risk. Addressing this requires an ecosystem-level approach to AI resilience, integrating automated red teaming, least-privilege architecture, and defense-in-depth from the start. Furthermore, rigorous evaluation is a prerequisite for recursive self-improvement, as autonomous systems cannot effectively improve without accurately measuring progress. Consequently, evaluation and resilience must be treated as foundational architecture rather than afterthoughts as autonomy scales.

  • Granting AI agents greater flexibility and tool access expands their attack surface, requiring defenses like least-privilege design from the start.
  • According to Wojciech Zaremba, AI resilience has no single silver bullet and instead requires an ecosystem of standards, detection, and infrastructure analogous to fire safety.
  • Evaluation and ideation are core bottlenecks for recursive self-improvement because agents cannot improve without reliable performance metrics.
  • Scaling autonomy across longer horizons necessitates scaling evaluation frameworks and resilience mechanisms in tandem.

3. Physical AI Needs Its Own Scaling Laws

Physical AI requires scaling dynamics distinct from language models, as current robotics remains in an early, unpredictable phase. Leading researchers suggest that physical intelligence must scale through world-action models, multimodal reasoning, and egocentric video rather than narrow task specialization. Additionally, the field is transitioning from open-loop imitation learning toward closed-loop reinforcement learning and interactive world simulators.

  • Jim Fan argues robotics lacks predictable scaling dynamics and remains stuck in the age of alchemy.
  • Proposed scaling axes for physical AI include world-action models, long-context memory, and large-scale egocentric video.
  • Sergey Levine advocates for robotic reasoning using high-level goal decomposition and images as intermediate thoughts.
  • Anastasis Germanidis highlighted interactive world models acting as learned simulators for counterfactual exploration.
  • Autonomous driving and robotics policies are shifting from open-loop imitation to closed-loop reinforcement learning and self-play.

4. The Real Frontier May Be the Learning Loop

The AI field is not converging on a single technical recipe, with unresolved debates spanning robotics, world models, and digital agents. Across both physical and digital domains, progress relies increasingly on iterative loops that connect action, observation, evaluation, and feedback. Ultimately, the next frontier of AI scaling may depend less on single model breakthroughs and more on building end-to-end systems capable of continuous learning from experience and environments.

  • The AI field is not converging on a single unified methodology across robotics and digital agents.
  • Robotics research remains divided between video-based versus structured world representations, imitation learning versus reinforcement learning, and simulation versus real-world experience.
  • Effective digital agents require integrations of external memory, tool harnesses, evaluation frameworks, security guardrails, and automated experimentation loops.
  • Robotics extends the agent loop into the physical domain via: perceive, reason, act, observe outcome, learn, and try again.
  • Future AI scaling is shifting focus toward closed-loop continuous learning systems rather than individual model breakthroughs.

The Bigger Picture

The AI frontier is moving beyond standalone models toward comprehensive systems, feedback loops, and environments that support autonomous agents and real-world interaction. In a featured discussion, Jeff Dean and Prof. Dawn Song reviewed milestones in AI development, such as mixture-of-experts and the natively multimodal design of Gemini. Looking ahead, Dean addressed the dual-use cybersecurity challenges posed by agentic systems and introduced Discovery Loop, a venture focused on accelerating scientific discovery through coordinated AI agents.

  • System design—including learning environments, feedback loops, and safeguards—is becoming as crucial as underlying model capabilities for autonomous agents.
  • Jeff Dean reflected on foundational advancements in modern AI, notably mixture-of-experts architectures and designing Gemini as natively multimodal.
  • Increasing model autonomy and agency introduces dual-use risks, particularly within cybersecurity.
  • Jeff Dean's venture, Discovery Loop, aims to accelerate scientific discovery by decomposing problems, coordinating agents, and rapidly running experiments.

Trends This Week

Enterprises are shifting focus from selecting single frontier models toward dynamic model routing to lower costs and maintain performance across tasks. Companies such as AT&T, Ramp, and Callosum are investing heavily in routing infrastructure, while Stripe has moved into the layer by agreeing to acquire model gateway OpenRouter. Meanwhile, AI capabilities and developer tooling continue expanding into adjacent workflows, exemplified by DeepSeek's new experimental visual-agent model and Cursor's release of its Origin code-hosting platform.

  • AT&T routes approximately 40% of its internal AI queries to open models, achieving up to 56% cost reductions with only a ~2% performance drop.
  • Ramp routes over 2.75 trillion tokens per month via its internal model router, reducing LLM costs by roughly 30%.
  • Callosum raised $100 million to build infrastructure that dynamically matches workloads across different models and hardware chips.
  • Stripe agreed to acquire OpenRouter, an AI gateway routing over 10 trillion tokens daily across 400+ models from 80+ providers.
  • DeepSeek introduced V4-Flash-Vision-Exp, an experimental multimodal model tailored for visual-agent tasks at low inference costs.
  • Cursor launched Origin, an agent-native code-hosting platform featuring repositories, pull requests, and GitHub synchronization.