The cleaner comparison is also a warning
A new Insurance Institute for Highway Safety study makes the automated-vehicle safety comparison more useful by cleaning the underlying record before calculating a rate. It reviewed 736 Level 4 public-road crash involvements from 2021 through 2024, removed duplicate or out-of-scope reports, and manually coded whether each incident was likely to be police-reportable. Only 22% survived that common severity screen as likely or possibly police-reportable.[1,2]
On that narrower comparison, Waymo's studied driverless vehicles had 68% fewer police-reportable crash involvements per vehicle mile than human drivers in San Francisco, Phoenix, Los Angeles and Austin. The study reports roughly 50 million driverless Waymo miles against about 222 billion human-driver miles in the same cities and period. Phoenix, San Francisco and Los Angeles showed lower Waymo involvement rates; Austin was 4% higher in a small sample, a result the study says needs more exposure before it can carry much weight.[1,2,4]
The control point is the reporting system
The result changes the operating question from whether a raw crash count looks large to whether every operator can supply a comparable denominator and a comparable severity classification. IIHS says the federal reporting system collects crash records but not consistent operator mileage, and that Waymo is the only major operator in the study providing the driverless mileage needed for this comparison. That makes exposure reporting part of the safety infrastructure, not just an analyst's preference.[1,2]
NHTSA's Standing General Order requires named manufacturers and operators to report certain crashes when an automated driving system was in use within 30 seconds of the event. The agency says its public data can include multiple reports for one crash, different information from different reporting entities and missing or inaccurate matching fields. The current dashboard is therefore a regulatory notification system first; it is not yet a clean, exposure-normalized league table.[3]
What the study does not establish
The study is a bounded result for the vehicles, cities, operating conditions and 2021–2024 records it examined. It measures crash involvement after a severity and duplication screen; it does not assign fault to an operator, establish a causal mechanism for each event or prove that every automated-driving system is safer than a human driver. The manual narrative coding is also difficult to scale, and human crashes may be underreported under the same comparison threshold.[1,2,3]
That caveat matters for insurers, regulators and road operators deciding whether a fleet can expand. A lower rate in the studied Waymo operating design domain is evidence for one safety case, not a transferable guarantee. The next comparison can become stronger only if operators publish exposure, reporting boundaries and event classifications that can be reconciled with the federal record.[1,3,4]
The next safety requirement is comparable exposure
The immediate checkpoint is not another headline crash count. It is whether NHTSA's monthly files and operator disclosures converge on deduplicated incidents, consistent severity thresholds and vehicle miles traveled for every large fleet. If they do, the industry can test whether the Waymo result persists as service expands and operating conditions diversify. If they do not, public confidence and insurance pricing will continue to rest on datasets that are timely enough to report incidents but too uneven to compare risk.[1,3,4]