← 回到 Reading
ByteByteGo 2026-08-17

Waymo vs Tesla: Two Ways to Build Self-Driving Cars

Autonomous vehicle perception fundamentally contrasts direct distance measurements from sensors like lidar against derived depth calculations from camera arrays. Waymo incorporates a redundant multi-sensor suite consisting of lidar, radar, and cameras to ensure reliability across adverse environmental conditions like rain and ice. In contrast, Tesla relies strictly on a pure vision strategy powered by exterior cameras and neural networks. Even as Waymo trims its camera count in sixth-generation vehicles through higher-resolution imagers, it retains sensor redundancy, highlighting the industry tradeoff between expensive direct physical measurements and cost-effective but error-prone derived estimates. Autonomous vehicle architectures hinge on how raw sensor inputs are converted into an operational representation. Waymo relies on the Waymo Foundation Model, combining a Sensor Fusion Encoder and a Gemini-trained Driving VLM to populate a World Decoder with compact, structured representations and pre-surveyed maps. In contrast, Tesla utilizes an extensive suite of per-camera neural networks to generate birds-eye-view representations on the fly without prior mapping surveys. This highlights an essential trade-off between inspectable, easily verifiable structured representations and expressive, nuance-capturing learned representations. Autonomous vehicle prediction systems must account for multiple possible future paths for each road user rather than committing to a single trajectory. In June 2025, Waymo published findings showing that motion forecasting quality scales predictably according to a power law with training compute, analogous to scaling laws in language models. This power-law relationship also held for closed-loop simulation performance, demonstrating that increasing data and compute directly enhances real-world driving safety. Meanwhile, Tesla introduced an updated reinforcement learning stage in version 14.3 to address long-tail edge cases.

閱讀原文 ↗
目錄 7 段
  1. 01Sensing
  2. 02Representation
  3. 03Prediction
  4. 04Planning
  5. 05Validation
  6. 06Training
  7. 07Conclusion

Sensing

Autonomous vehicle perception fundamentally contrasts direct distance measurements from sensors like lidar against derived depth calculations from camera arrays. Waymo incorporates a redundant multi-sensor suite consisting of lidar, radar, and cameras to ensure reliability across adverse environmental conditions like rain and ice. In contrast, Tesla relies strictly on a pure vision strategy powered by exterior cameras and neural networks. Even as Waymo trims its camera count in sixth-generation vehicles through higher-resolution imagers, it retains sensor redundancy, highlighting the industry tradeoff between expensive direct physical measurements and cost-effective but error-prone derived estimates.

  • Cameras require depth to be computed from pixel arrangements, creating potential ambiguity between object size and distance.
  • Lidar directly measures physical distance via laser pulse return times, producing a 3D point cloud.
  • Waymo's sixth-generation system includes 13 cameras, four lidar units, six radar units, and audio receivers, reaching up to 500 meters of overlapping coverage.
  • Waymo reduced its camera count from 29 on the fifth-generation Jaguar I-PACE to 13 on the sixth-generation system by using 17-megapixel imagers.
  • Tesla utilizes a 'pure vision' approach with Tesla Vision on Model 3 and Model Y, omitting radar in favor of camera input and neural network processing.

Representation

Autonomous vehicle architectures hinge on how raw sensor inputs are converted into an operational representation. Waymo relies on the Waymo Foundation Model, combining a Sensor Fusion Encoder and a Gemini-trained Driving VLM to populate a World Decoder with compact, structured representations and pre-surveyed maps. In contrast, Tesla utilizes an extensive suite of per-camera neural networks to generate birds-eye-view representations on the fly without prior mapping surveys. This highlights an essential trade-off between inspectable, easily verifiable structured representations and expressive, nuance-capturing learned representations.

  • Waymo's core architecture centers on the Waymo Foundation Model, composed of a Sensor Fusion Encoder and a Driving VLM.
  • Waymo's Driving VLM is trained using Gemini and fine-tuned on Waymo driving data to manage rare edge cases.
  • A World Decoder consumes outputs from both Waymo components to forecast behaviors, generate trajectories, and produce high-definition maps.
  • Waymo's compact structured representations facilitate inference-time safety verification, large-scale simulation, and measurable training feedback.
  • Tesla utilizes 48 networks requiring nearly 70,000 GPU hours of training to output 1,000 distinct tensors per time step across semantic segmentation, object detection, depth estimation, and birds-eye-view modeling.
  • Waymo relies on pre-surveyed prior mapping matched against live sensor feeds, whereas Tesla skips prior surveys and derives equivalent environmental context in real time.

Prediction

Autonomous vehicle prediction systems must account for multiple possible future paths for each road user rather than committing to a single trajectory. In June 2025, Waymo published findings showing that motion forecasting quality scales predictably according to a power law with training compute, analogous to scaling laws in language models. This power-law relationship also held for closed-loop simulation performance, demonstrating that increasing data and compute directly enhances real-world driving safety. Meanwhile, Tesla introduced an updated reinforcement learning stage in version 14.3 to address long-tail edge cases.

  • Predictive driving systems maintain several weighted future trajectories for road users to avoid fragility in critical scenarios.
  • Waymo's research showed that motion forecasting quality follows a power law relative to training compute, similar to language models.
  • Waymo evaluated motion forecasting using an internal dataset of 500,000 driving hours.
  • Increasing compute at inference time improved autonomous driving performance on more difficult driving scenarios.
  • The power-law scaling trend observed by Waymo translated to closed-loop simulation performance where agent actions affect the environment.
  • Tesla introduced an upgraded reinforcement learning stage in version 14.3 targeted at handling rare, long-tail edge cases.

Planning

Autonomous vehicle planning systems select trajectories that require verification before execution. Waymo generates action sequences using large Teacher models distilled into onboard Student models, requiring trajectory confirmation from an independent validation layer before the vehicle acts. Conversely, Tesla relies on fleet-scale planning algorithms and human verification via Full Self-Driving (Supervised), using driver-monitoring strikeouts to enforce safety. While automated validation layers prevent defined failure modes, their protective scope is inherently limited by their explicit criteria.

  • A trajectory is defined as a specific path paired with assigned speeds that must be verified before execution.
  • Waymo distills large Teacher models into onboard Student models to run action-sequence generation in real time.
  • Waymo requires agreement between two independent components—the generative student model and a validation layer—prior to vehicle movement.
  • Tesla relies on human drivers to verify trajectories in Full Self-Driving (Supervised), enforcing engagement via a strikeout suspension mechanism.
  • Validation layers evaluate trajectories strictly against pre-defined criteria, functioning similarly to assertions in production code.

Validation

Waymo and Tesla take fundamentally different approaches to validating and reporting autonomous vehicle safety metrics. Waymo evaluates fully driverless operations without human intervention against adjusted local human baselines, reporting significant reductions in injury-causing crashes backed by peer-reviewed methodology and public data. Conversely, Tesla assesses its Full Self-Driving (Supervised) feature against manual Tesla driving using internal telemetry, measuring collision frequency where a human driver remains accountable. Because these reports measure distinct populations and operational paradigms, their safety figures cannot be compared directly side-by-side.

  • Waymo recorded 220.6 million rider-only miles through March 2026, reporting a 94% reduction in serious injury crashes and an 82% reduction in injury-causing crashes relative to human benchmarks.
  • Waymo makes its validation methodology reproducible by publishing in peer-reviewed journals and releasing raw data.
  • Tesla evaluates Full Self-Driving (Supervised) by comparing it to manually driven Teslas using telemetry, reporting 7 times fewer major/minor collisions.
  • Tesla counts any crash where the system was active within five seconds before impact as an engaged collision, but omits fault attribution.
  • The metrics are not directly comparable because Waymo measures autonomous systems with no human backup, while Tesla measures a driver-assist system where the human remains responsible.
  • Waymo separates pre-deployment validation (governed by its Safety Framework and Safety Case) from post-deployment safety impact analysis.

Training

Waymo and Tesla utilize fundamentally different training mechanisms to improve their autonomous driving systems between releases. Waymo operates a single foundation model powering three distilled components—the Driver, the Simulator, and the Critic—connected via simulation reinforcement learning and outer real-world safety loops. In contrast, Tesla relies on fleet telemetry and event packets collected from consumer vehicles to build automated test suites and simulation data. Tesla trains its models at scale on dedicated supercomputers, including Cortex 1 and Cortex 2, utilizing hundreds of thousands of H100-equivalent GPUs.

  • Waymo powers three core components (Driver, Simulator, and Critic) off the same foundation model and distills them for production volume.
  • Waymo optimizes driving using an inner reinforcement learning loop in simulation and an outer validation loop driven by Critic evaluations of real driving.
  • Waymo reports that its fully autonomous mileage exceeds its manually driven data and encounters situations that simulation cannot fully reproduce.
  • Tesla collects consumer fleet data via two telemetry paths: trip-end anonymized mileage by road/control type and collision-triggered packets tied to specific vehicles.
  • Tesla reported receiving 2.5 billion telemetry packages in the third quarter of 2025.
  • Tesla conducts model training on Cortex 1 with over 100,000 H100-equivalent GPUs and Cortex 2 with over 130,000 in early ramp.

Conclusion

Autonomous driving systems navigate a fundamental tradeoff across every operational stage: explicit, pre-determined knowledge versus real-time computation by internal model states. This dynamic shapes sensing, representation, prediction, planning, validation, and training. Waymo favors explicit, pre-written knowledge requiring city-specific preparation and bespoke hardware, whereas Tesla relies heavily on computed knowledge. It remains an open question whether one philosophy will dominate or if an intermediate hybrid approach will prevail.

  • Autonomous driving continuously balances pre-determined, inspectable information against real-time model computation across all stages.
  • Waymo emphasizes written-down knowledge, which relies on per-city preparation and purpose-built hardware.
  • Tesla focuses more on computed knowledge derived on-the-fly during the drive.
  • Key architectural questions span distance sensing, inspectable world representations, multi-future prediction, trajectory verification, safety evidence, and training mileage selection.