A storm forms over the Caribbean. Weather models disagree on where it is going and how strong it will become. Five days before landfall, an AI model predicts, with 80 percent confidence, that the system will strike Jamaica as a Category 5 hurricane. That prediction turns out to be correct. The storm, Hurricane Melissa, causes widespread flooding and landslides, but communities receive earlier warnings than would have been possible with traditional forecasting tools. This is not a hypothetical scenario. It is what happened in October 2025, and it marks a turning point in how artificial intelligence is being applied to one of the most consequential prediction problems in science.
The One-Day Advantage That Took a Decade to Earn
The WeatherNext model, developed by Google DeepMind and Google Research, delivers cyclone predictions with a lead time that is, on average, one full day longer than existing models. To put that in perspective: its three-day forecasts are as accurate as previous models’ two-day forecasts. That single day of additional warning time is not a minor technical improvement. Historically, advancing forecast accuracy by a day would require roughly a decade of incremental scientific work, according to the researchers behind the model.
Mike Brennan, director of the US National Hurricane Center, frames the stakes clearly. Evacuations need to be organized. Supplies need to be staged. Resources need to be positioned. Every one of those tasks is time-sensitive, and acting on a wrong forecast carries serious consequences. “Time is really golden when it comes to those types of decisions,” Brennan has said. An extra day does not just improve a number on a performance chart. It changes what is operationally possible for the people responsible for protecting lives.
The Intensity Problem That Stumped Earlier Models
Track prediction, meaning the direction a storm travels, has been a relative strength of AI weather models for some time. Intensity prediction is a different matter entirely. Hurricanes operate across multiple spatial scales simultaneously. Forecasting a storm’s path requires global-scale atmospheric data: cold fronts, prevailing winds, large-scale pressure systems. Forecasting how strong a storm will become requires fine-grained, local data about atmospheric and ocean conditions at the storm’s core.
Earlier AI models handled track reasonably well. Intensity, as Kate Musgrave, tropical cyclone group lead at the Cooperative Institute for Research in the Atmosphere and a co-author on the paper, puts it, “they could not do well at all.” This distinction matters enormously in practice. A storm that appears manageable at Category 1 can intensify rapidly overnight into a Category 5 emergency. Hurricane Melissa illustrated exactly this scenario, and it marked the first time the National Hurricane Center was able to predict a Category 5 hurricane while the storm was still at Category 1 intensity.
What makes WeatherNext’s performance on intensity particularly puzzling to researchers is that it achieves this accuracy using lower-resolution atmospheric data than traditional models require. When the broader scientific community learned this, the reaction was surprise. The assumption had been that fine-grained intensity prediction demanded fine-grained input data. WeatherNext appears to extract meaningful signal from coarser inputs in ways that are not yet understood. “It’s a black box at the end of the day,” says Ferran Alet, a research scientist at Google DeepMind and one of the paper’s lead authors, “but that gives physicists a signal that something is happening that was not previously understood.”
What a Black Box Teaches Science
This is where the story becomes interesting beyond meteorology. The WeatherNext model does not produce a single forecast. It generates a range of potential scenarios for each developing storm, capturing the sensitivity of complex systems to small initial variations. Last year it produced 50 scenarios per storm. It now generates 1,000. That volume of probabilistic output is simply not achievable with conventional numerical modeling given current computing constraints.
The model’s opacity is both a limitation and, unexpectedly, a scientific instrument. Researchers do not fully understand why it works as well as it does. But the fact that it works, using inputs that the scientific community considered insufficient, is itself informative. It suggests that lower-resolution data contains more predictive signal about storm intensity than previously believed. That is a hypothesis worth investigating, and it is one that would not have emerged without the AI model’s anomalous performance.
Google DeepMind has announced that it is open-sourcing the WeatherNext models used during hurricane season. The intent is to allow the broader research community to examine, use, and build on them. Alet has expressed hope that this openness could generate new scientific insights into how cyclones behave, not just better forecasts.
Brennan’s framing is worth holding onto here. The model is a powerful new tool, but it is one tool among many. No single model guarantees superior performance across every storm or every season. Human expertise remains essential, not because AI cannot process data, but because translating a track and intensity forecast into actionable impact assessments requires judgment that goes beyond pattern recognition. “It’s the impacts that kill people,” Brennan notes. That translation work belongs to experts.
In Short
WeatherNext demonstrates that AI can close a long-standing gap in hurricane forecasting, specifically the difficulty of predicting rapid intensity changes, while operating on data inputs that conventional science considered inadequate. The one-day improvement in lead time is meaningful in human terms: it expands the window for evacuation, logistics, and preparation. The model’s unexplained accuracy also points toward something broader. When an AI system outperforms expectations using methods that scientists do not yet understand, the black box becomes a research question. The goal is not just better forecasts. It is new knowledge about how storms work.
Based on reporting from Ars Technica.