🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
We talked to Google’s Oscar winning “Giganerd” about automating science, solving climate change, and how future generations can contribute to science in the age of superintelligent AI
一句話總結
來賓
Google 資深科學家、奧斯卡技術獎得主與機器學習演算法先驅,長期致力於以計算與 AI 解決氣候變遷等重大科學挑戰。
主要討論話題
將科學問題轉為可評分任務與 Auto-Kaggle
John Platt 與團隊發現許多科學難題本質上可簡化為「可評分任務」(scoreable task),核心挑戰在於制定評分函數,接著尋找能將分數最大化的程式碼。團隊藉由 Google 旗下的 Kaggle 數據,提出了「Auto-Kaggle」構想,目標是自動化解決各類能以指標量化的競賽與科研挑戰。
Google ERA 架構:MCTS 與 LLM 的演化結合
Google 的經驗研究輔助系統(ERA)概念上結合了蒙地卡羅樹狀搜尋(MCTS)與大型語言模型。系統維護過去實驗 Notebook 的運作樹,使用 UCB 規則樂觀選擇最有潛力的節點(甚至第 5 佳的節點)並共享分支歷史,再由 Gemini 提出約 10 個變異方向。該系統在模型從 Gemini 2.0 升級到 2.5 時出現質的躍升,使自動化科研從不可行變為實用。
避免過擬合與 Goodhart 定律的陷阱
John 強調 ERA 產出的是預測模型,科學家必須嚴格驗證其描述真實世界的能力,否則這項強大工具容易「切斷你的手指」。他以 Google 飛機凝結尾偵測競賽為例,獲勝者利用了標註偏差中的半像素誤差獲勝,但這無助於真正解決問題。這印證了 Goodhart 定律,當指標變成唯一目標時,人類和 LLM 都會出現 reward hacking 行為,因此研究起步應先嘗試線性回歸或 SVM。
AI 解決氣候變遷:飛機凝結尾與 FireSat
飛機產生的冰晶凝結尾佔人為全球暖化達 1%,單一克廢氣在冰超飽和區域可引發十公斤冰晶並長時間阻擋熱量散逸。解法是讓飛機微調巡航高度,但評估預防暖化幅度的反事實模型卡關兩年,ERA 成功找出包含先前未考慮的混淆因子之簡潔模型。此外團隊亦發展 FireSat 衛星群技術,在火勢僅房間大小時及早發現,防止蔓延成數十萬英畝的大火。
扎根領域知識與親手實作的價值
回顧 1982 年受教於 Richard Feynman 的經驗,John 認為在超智慧 AI 時代,科學家最關鍵的技能仍是建立深厚的領域專業知識與品味。他比喻做研究可以開車上山,但偶爾自己「徒步登山」親手實作基礎細節非常必要。過度追求最佳化會導致如同機器學習般的過擬合,研究者必須牢記 Feynman 的忠告:絕對不要欺騙自己,而自己正是最容易被欺騙的人。
對你的啟發
樂觀探索優於貪婪搜尋的 Agent 探索架構
在設計反覆嘗試與程式碼生成的 Agent 時,不要只挑選目前分數最高的解法;導入類似 MCTS 的 UCB 機制與分支歷史共享,能讓 Agent 保留次佳路徑的變異潛力,避免陷入局部最佳解。
基礎模型的推理門檻決定複雜工作流的可行性
複雜的程式碼演化與自我除錯工作流對底層 LLM 的推理品質高度敏感;如同 ERA 在 Gemini 2.0 到 2.5 之間經歷質變,當 Agent 表現不佳時,升級底層模型能力往往比過度調校 Prompt 更具決定性。
自動化評估指標需防範 Reward Hacking
構建以評分函數驅動的評估與最佳化流程時,必須警惕 Goodhart 定律;過度依賴單一代理指標(proxy metric)容易讓 Agent 鑽標籤格式瑕疵等漏洞,需搭配簡潔基準與反事實驗證。
堅持以最簡單基準線(Baseline)作為系統起點
在搭建複雜的 RAG 或 Agent 工作流前,應始終優先建立最樸素的 Baseline(如傳統線性回歸或簡易規則);這能快速界定邊界效應,防止團隊過早陷入超最佳化(Hyper-optimization)的泥淖。
節目筆記
How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar , two textbook algorithms , two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date. And the result was an epic and fun chat covering Google’s Empirical Research Assistance (ERA), how AI can help battle climate change, and tons of great stories about the co-evolution of science and AI.
John’s colleague Dave Bacon likes to tease John that his career has been defined by being twenty years early to the next big thing. This may be convolutional neural networks (some credit him with coining the term), fusion research, quantum computing. John and Google have been working on solving some of humanity’s hardest problems with AI and computation for well over a decade now. Recently John and his team set their sights on using AI to solve any scientific problem that can be written down as a score.
John’s team has taken on many hard scientific problems over the years. In solving these, they noticed a pattern, many scientific problems can be reduced to what John calls a “scoreable task”. Once you have the score function, the goal is to find some code that maximizes the score. The hard part is in formulating the score, but once you have the score finding the maximizer can still be quite a lot of effort.
John’s team set out to automate solutions to this general problem. This came out of the idea of an “auto-Kaggle” AI, which can solve any Kaggle problem you can throw at it. Kaggle is owned by Google, so all the data was ready and easily available to them!
The result is Google’s Empirical Research Assistance or ERA ( paper , github , blog ). 1 ERA is surprisingly simple conceptually. Gemini (or your LLM of choice) keeps a running tree of past experiments (notebooks) and where they’re going. It’s a close cousin of Monte Carlo Tree Search : at each iteration the Upper Confidence Bound rule picks which notebooks are most promising to mutate. This is optimistic, not greedy, so sometimes even the fifth-best notebook gets chosen. Gemini then proposes mutations for each one, about ten at a time. The history of each branch is shared, so different leaves can learn from each other.
“It’s almost like having a hyper-eager grad student who doesn’t sleep.”
Evolutionary algorithms have been around since the 70s, but this works because Gemini actually knows where to look! What’s even more interesting is that there was a step change between Gemini 2.0 and 2.5, and this went from just not working to working great.
ERA is so powerful that John and his team solved many outstanding problems with it, resulting in at least ten papers . Some of these were climate change related, which we talk about in the next section.
So, we had to ask: if you have an optimization god how do you avoid fooling yourself? John’s answer is that ERA provides predictive models. It’s up to the scientist to make sure they’re truly descriptive. Some of this just involves good old-fashioned careful machine learning science. “It’s a power tool. It can slice your fingers off.” This led to some fun discussion about Kaggle competitions, and the fun ways people can overfit to datasets without meaningfully solving the problem you actually care about: Google’s contrail-detection competition was won by entrants who noticed a half-pixel error in the labels (is the origin at the corner of the pixel or the center?) and this turned out to be a part of the winning special sauce . Great for winning $15,000, not so helpful if you actually want to solve contrails.
“People themselves will act like these LLMs and try to reward hack. It goes back to Goodhart’s law : any metric that becomes a target is no longer good as a metric.”
His advice for where to start instead?
“Always just fit linear regression. Just do it. Just do it. Just do it. Or SVM.”
John and his team have worked extensively to mitigate the effects of climate change. We talked about several of their initiatives.
Perhaps the most interesting result we talked about was reducing the effects of condensation trails (contrails) from airplanes. Those little streaks you see running behind planes somehow account for 1% of all human-induced global warming ?!? Some of these trails of ice crystals can hang out for days. These crystals are black in the infrared, acting like a thermal blanket that traps heat day and night.
It’s easy to understand what’s happening here, a region of atmosphere becomes “ice supersaturated”, 2 and a tiny bit of exhaust seeds water vapor that instantly crystallizes. The scale here is astounding, with a single gram of exhaust resulting in ten kilograms of ice crystals.
The solution to all of this is quite simple, in principle! We know what parts of the atmosphere are most likely for the trails to form. Just have the planes drop a flight level or two. Problem solved, right? Well, the hard part is accounting for how much warming was prevented. This is a counterfactual problem, parts of which stumped John’s team for over two years. They had a working model for the heat-trapping half , but not for the reflected sunlight. ERA was able to find a simple model with some confounders they hadn’t considered. Cracked it!
Modeling climate generally is a hard problem. Climate is best thought of an attractor of many different possible weather outcomes. 3 This makes it much harder to model.
“Weather is where you are on the attractor , and climate is the statistics of the attractor. The problem with climate is that we’re altering it. The attractor itself is changing, it’s moving.”
John and his team have worked on treating both the symptoms and the disease of climate change, with several other works in the area. Another fun example we briefly cover is FireSat, a way of using a constellation of satellites to rapidly identify fires before they grow too big to put out. For anyone living in California, you understand the problem. In dry years a small fire can result in hundreds of thousands of acres. If you could find this fire when it’s the size of a room, it could be put out. By the time it hits an acre we have a much harder problem.
By now it should be clear John has an incredible and unique view over the intersection of science, computation, and AI. John talked about a class on physics of computation 4 he took with Richard Feynman back in 1982. This was when quantum computing was an ill-defined concept with no theory or experimental backing. John recalls every Tuesday was a guest lecture, and every Thursday was Feynman explaining why the Tuesday guest was wrong. John also recalls doing science back when there was essentially no compute, a million operations per second was cutting edge.
What is John’s recommendation: the most important skill is developing deep domain expertise. There’s no other way to develop taste than to tackle hard problems. One surprising part of this is that John recommends spending time doing things the old fashioned way. Play with tools, and just implement things yourself.
“You could drive up the mountain, or you could hike up the mountain, and maybe it’s okay, even fun, to occasionally hike.”
Summing it up, John’s message to the audience is that there will still be a place for scientists, and that if anything it will just open up more opportunities for “the creative stuff, the rigorous stuff, the philosophy stuff.” But don’t forget to spend time doing the grunt work.
“There just seems to be this strong impetus in the world to optimize and squeeze everything out. But you do lose something when you hyper-optimize. It’s overfit.”
And whatever tools you end up using, John’s advice is the same one Feynman gave him forty years ago: you must not fool yourself, and you are the easiest person to fool.
We had a great time talking with John. We hope you enjoy!
Fusion is three years away, not thirty, if you ask John. And why the Lawson criterion means every fusion approach has an Achilles heel. Why superconducting qubits are still finicky. The asteroid he named after his mom, which turned out to have a moon. The looming helium shortage nobody talks about. How NeurIPS started as people crashing a private workshop at Snowbird, and why Hopfield networks are all you need . Being Carver Mead’s sysadmin on a VAX with an 80 MB disk the size of a dishwasher. Finding asteroids in 1985 with film, a stereoscope, and a letter to Brian Marsden. The Vera Rubin Observatory found 11,000 in six weeks. The Feynman effect: total clarity in the room, none once you leave. Quantum echoes, the NISQ era , and why he thinks quantum is neither thirty years away nor tomorrow. A startup that wants to inject mercury into a fusion reactor and sell the transmuted gold. “It might not work.” John’s 20% time rule for his own group: do stuff for learning, and you don’t even have to tell him what.
The ERA GitHub repo features an open source implementation that ran Gemini but can be used with any LLM. ERA is not currently available as a Google product.
“Ice-supersaturated” is about water vapor, not liquid water. Cold air can hold a given amount of vapor, and there are two different limits: the amount in equilibrium with liquid water, and the smaller amount in equilibrium with ice. Below freezing, a pocket of air can sit between those two limits. It has more vapor than ice can tolerate, but not enough to condense into droplets, and ice won’t form directly from vapor without a seed. So the vapor just hangs there, metastable, sometimes for days, until something seeds it.
We recently covered the weather-climate crossover in our episode with Anima Anandkumar , and we plan on covering both weather and climate more in future episodes.
This was really about quantum computing, but in the early days before anyone really knew what this meant and it was just a vague idea Feynman and a few others were kicking around.
節目內容來自 Latent Space,版權屬原作者。 閱讀原文 ↗