Dean Shabi
All work

Renewcast · 2026

Weather forecasts can't see what a plant did an hour ago. This model can.

Renewcast's forecasts come from weather models, so they miss what a plant is doing right now. Soiling, a tripped inverter or a cloud bank 20 km off course all show up in the readings first. I trained a small recurrent network that reads the latest readings next to the forecast customers already received and corrects the next four hours. One model covers the solar fleet and another covers the wind fleet.

Role
Model design, training and evaluation
Stack
PyTorch, Python, MLflow, Databricks
Wind turbines on green hills under a cloudy sky

How it works

  1. 1

    Delivered forecast

  2. 2

    Latest plant readings

  3. 3

    Fleet model

    solar or wind

  4. 4

    Four-hour correction

The model reads the forecast customers already got and the latest readings. Anything beyond four hours is left as it was.

Results

Solar forecast error over the next four hours

Average error at each step ahead. Persistence is slightly better for the first 30 minutes and the fleet model is better from 45 minutes. Both hand back to the delivered forecast at four hours.

  • Fleet GRU
  • Persistence
  • Delivered forecast

Backtest of an earlier version of the fleet model on Renewcast's solar fleet.

Show data
Delivered forecastPersistenceFleet GRU
15 min6.1%2.7%3.1%
30 min6.1%3.7%3.8%
45 min6.0%4.3%4.1%
60 min6.1%4.6%4.4%
75 min6.1%5.0%4.7%
90 min6.1%5.2%4.9%
105 min6.0%5.3%5.1%
120 min6.1%5.6%5.3%
135 min6.1%5.5%5.3%
150 min6.1%5.7%5.5%
165 min6.1%5.8%5.6%
180 min6.2%5.8%5.7%
195 min6.2%6.0%5.7%
210 min6.2%6.0%5.7%
225 min6.1%6.0%6.0%
240 min6.2%6.2%6.2%

In a backtest that held out each client, the median plant-month's error over the first two hours fell 21.9% for solar and 26.6% for wind. Carrying the latest error forward managed 16.3% and 25.3%.

Key decisions

  1. 01

    One model for the whole fleet

    A single GRU learns how forecast errors develop across plants. It beat per-plant models, and it means maintaining two models instead of 224.

  2. 02

    Test it the way it will run

    Training uses forecasts that were actually delivered and readings with realistic delays. I held out each client in turn and scored each month with a model trained only on earlier months. A check rejects any feature that leaks the future.

  3. 03

    Never correct its own output

    The model always corrects the original forecast, never one it already corrected. Output stays within plant capacity. If the newest reading is more than two hours old, the original forecast goes out unchanged.

Have a machine learning system that has to hold up in production? I'd like to hear about it.