主題聚類
具身智能與物理世界模型(約 15 篇) — 本週具身智能領域呈現爆發式增長,研究重點從單純的視覺語言模型轉向具備物理常識的「世界模型」。開發者正致力於讓 AI 理解 3D 空間一致性、物體恆常性(Object Permanence)以及預測物理變量(Deltas),以實現更精準的機器人操控。
[2609.27656](https://arxiv.org/abs/2609.27656) InternW0: A Foundational Physical World Model for Efficient Real-World Interactions[2609.23863](https://arxiv.org/abs/2609.23863) Grounded Action Model: 3D Grounding as a Foundation for Robotics[2609.24815](https://arxiv.org/abs/2609.24815) Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI[2609.28654](https://arxiv.org/abs/2609.28654) Training Object Permanence in World Models
Agent 架構演進與長效記憶(約 12 篇) — AI Agent 正從簡單的 ReAct 模式轉向更複雜的系統架構,引入了類似人類「系統一與系統二」的雙重處理機制。新技術如 Just-in-Time Memory (JitMem) 允許 Agent 在讀取時才根據任務過濾記憶,而 Agensh 則展示了將協作規模擴展至 1,024 個 Agent 的可能性。
[2609.29429](https://arxiv.org/abs/2609.29429) Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures[2609.27334](https://arxiv.org/abs/2609.27334) Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents[2609.26781](https://arxiv.org/abs/2609.26781) Agensh: Scaling Organizational Intelligence to 1,024 Agents[2609.29444](https://arxiv.org/abs/2609.29444) IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
影片生成與物理規律校準(約 10 篇) — 影片生成技術正從追求視覺華麗轉向追求「物理真實」。研究者開始診斷為何 Diffusion 模型會違反物理定律,並透過強化學習(Reward Modeling)與 Prompt 增強技術來提升影片的連貫性與電影感,特別是針對長達 30 秒的多鏡頭序列。
[2609.30221](https://arxiv.org/abs/2609.30221) WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation[2609.23658](https://arxiv.org/abs/2609.23658) Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms[2609.22947](https://arxiv.org/abs/2609.22947) RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling[2609.28923](https://arxiv.org/abs/2609.28923) ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
安全評測與數據洩漏防禦(約 8 篇) — 隨著模型能力增強,評測的真實性成為焦點。Schrödinger’s Code Repository 提出動態變換程式碼庫以防止模型依賴記憶(Data Leakage)來解題;同時,研究者開發出針對 LLM 的「測謊技術」,探測模型是否在刻意隱藏已知知識(Sandbagging)。
[2609.27891](https://arxiv.org/abs/2609.27891) Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?[2609.21996](https://arxiv.org/abs/2609.21996) A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal[2609.29647](https://arxiv.org/abs/2609.29647) AgentKernel: The Trust-Native Agentic Operating System[2609.22076](https://arxiv.org/abs/2609.22076) APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
產業動態:Agent 硬體化與模型更新(約 15 篇) — 本週產業焦點在於 Meta Connect 大會,Meta 透過 Muse AI Agent 與智慧眼鏡的深度整合,成功在熱度上蓋過了 OpenAI GPT-6 與 Anthropic Opus 5.5 的發佈。硬體端也有重大突破,Qualcomm 新晶片已支持在裝置端運行 30B 規模的 MoE 模型,預示著 Edge AI 時代的加速到來。
techcrunch-com-2026-09-25-at-meta-connect-the-companys-smart-glasses-were-everywhere At Meta Connect, the company’s smart glasses were everywheretechcrunch-com-2026-09-22-openai-launches-gpt-6-sol-and-luna OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakestechcrunch-com-2026-09-22-qualcomm-launches-two-new-smartphone-chips-with-emphasis-on-ai Qualcomm launches two new smartphone chips with emphasis on AItechcrunch-com-2026-09-21-googles-899-googlebook-is-a-bet-that-youll-buy-a-new-laptop-for-gemini Google’s $899 Googlebook is a bet that you’ll buy a new laptop for Gemini
本週趨勢觀察
本週 AI 領域呈現出明顯的「從雲端走向物理世界」與「從模型走向系統」的雙重趨勢。Meta 的 Muse Agent 策略顯示,頂尖科技公司正試圖將 AI 封裝進智慧眼鏡、穿戴式裝置(如 Muse Charm)等硬體中,使其成為全天候的個人助理,這延續了上週關於「具身化」的討論。技術層面上,研究界開始正視模型在物理常識上的缺陷,並提出多個針對影片生成與機器人操控的校準框架。值得注意的是「Vibe Coding」與 AI 生成應用的興起,雖然帶動了開發效率與營收,但也引發了如 Supabase 數據洩漏等嚴重的安全性隱憂,顯示出快速開發與安全審計之間的失衡。
最值得深讀
[2609.27891](https://arxiv.org/abs/2609.27891) Schrödinger's Code Repository — 揭示了當前程式碼評測基準可能因數據洩漏而失效,並提出動態混淆的創新解決方案。[2609.27656](https://arxiv.org/abs/2609.27656) InternW0 — 上海 AI Lab 推出的首個物理世界模型,為機器人實時環境適應提供了不對稱影片-動作架構的新範式。[2609.21996](https://arxiv.org/abs/2609.21996) A Lie Detector Test for Language Models — 透過分析內部狀態而非外部輸出,成功識別模型是否在「裝傻」或隱藏知識,對 AI 審計至關重要。
📈 本週上升中的名字
- JEV — 1→9 次提及
- Muse — 2→9 次提及
- VLA — 6→17 次提及
- GSM8K — 2→5 次提及
- Transformer — 8→14 次提及
- LLM — 22→37 次提及
- Meta — 7→12 次提及
- TechCrunch Disrupt 2026 — 13→20 次提及
- RAG — 7→11 次提及
- LoRA — 3→5 次提及