Model & Methodology
Historical cancellation data, prediction model performance, and how FerryForecast works.
Historical Data
Historical Cancellation Data
Based on 578 historical cancellation events currently sourced from SSA cancellation-email records and Hy-Line #ALERT feed records.
Steamship Authority
223 cancellations (2023-01-16 to 2026-10-08)Hy-Line Cruises
355 cancellations (2023-08-25 to 2026-09-26)Cancellations by Month
Key Takeaways
- Weather is the dominant cause: 44% of SSA and 92% of Hy-Line cancellations
- SSA has more mechanical (33%) and crew (12%) cancellations, consistent with running year-round car/passenger service
- Winter months (Dec-Mar) see the most cancellations from both operators
- Hy-Line runs year-round but with reduced winter service (~6 departures/day vs. 10+ in summer)
Source note: This card currently uses alert-derived labels. We are migrating to operator trip-level logs (primary) plus AIS run/no-run verification (secondary), with monthly reconciliation against SSA’s official performance reports. Until that transition is complete, category percentages here remain alert-derived.
Model Performance
Model Performance
Trained on 33,416 trips from Jan 2023 to Dec 2025, 1,026 of them cancelled. Last retrained Oct 4, 2026. Chances are calibrated, so a 20% chance of running means about 1 in 5 comparable trips ran.
46% Brier skill
less error than always predicting the usual 2.8% cancellation rate
0.91 AUC on windy trips
how well it ranks cancelled above running trips when wind is 20+ mph
0.98 AUC, all trips
most trips are in calm weather, which is easy to call
Tested on 11,943 trips from Jan 2025 on, using a model trained only on the 21,473 trips before then.
CalibrationEach trip is scored by a model that never saw its date, then grouped by the chance shown
When we show an X% chance of running, how often did those trips run? Bars with fewer than 30 trips are faded: too few to judge.
Feature effectsHow much each weather variable moves the odds of a cancellation in the served model
Odds ratio per 1 standard deviation increase. Above 1 raises the odds of a cancellation; below 1 lowers them.
Detailed statistics
| Check | AUC | Windy AUC | Brier skill | Calibration error |
|---|---|---|---|---|
| Unseen trips from Jan 2025 on | 0.98 | 0.91 | 46% | 0.3% |
| 5-fold, whole days held out | 0.98 | 0.90 | 46% | 0.3% |
| Feature | Coefficient (per SD) | Odds ratio |
|---|---|---|
| Wave Height | 0.8482 | 2.336 |
| Wind Gusts | 0.6241 | 1.867 |
| Wind Speed | 0.5906 | 1.805 |
| Wave Period | 0.1854 | 1.204 |
| Season (winter vs summer) | 0.1389 | 1.149 |
| Season (spring vs autumn) | -0.0139 | 0.986 |
| Wave Period Missing | -0.0002 | 1.000 |
Wind speed and gusts move together, so their separate effects are less certain than their combined one.
Methodology
Methodology
How this site predicts whether Nantucket ferries will run, start to finish.
What we predict
For each scheduled ferry departure between Hyannis and Nantucket, we estimate the probability that the trip will operate as planned, based on forecast weather conditions. A “95%” means that historically, trips operating under similar weather conditions ran about 95% of the time. A “40%” means conditions resemble those when roughly 6 in 10 comparable trips were cancelled.
We estimate weather-related cancellation risk only. Mechanical breakdowns, crew shortages, and operational decisions are outside the model’s scope. SSA’s official performance data attributes about 33% of Nantucket trip cancellations to weather and 10% to mechanical causes, with the remainder in a broad “other” category that likely includes some weather-adjacent decisions. Hy-Line cancellations skew more heavily toward weather (~85% in our dataset).
When the estimate is above 95%, weather-related cancellation risk appears low based on current forecasts. When it falls below 40%, conditions resemble historical cancellation days. Non-weather cancellations are not modeled and will not appear in these estimates.
The prediction model
We use logistic regression, a well-understood statistical model that estimates the probability of a binary outcome (cancelled or not cancelled) given a set of weather measurements.
Training data: Trained on 33,416 individual trip records from Jan 2023 to Dec 2025, each labeled as cancelled or normal. Of those, 1,026 were weather cancellations and 32,390 ran as scheduled. The model is retrained weekly; last retrained Oct 4, 2026.
Performance on trips it never saw: a copy of the model trained only on trips before Jan 2025 was tested on the 11,943 trips from then on. It removed 46% of the error of always predicting the usual 2.8% cancellation rate (Brier skill score). On windy trips (20+ mph, where most cancellations happen) its AUC is 0.91; across all trips it is 0.98. We don’t quote accuracy: with about 3% of trips cancelled, always saying “runs” is right almost every time and says nothing about storm days.
Probabilities are calibrated with isotonic regression, fitted on predictions for trips each model never saw, so when the site shows 80%, roughly 80% of comparable trips ran as scheduled.
Why logistic regression? It’s interpretable (you can see exactly which weather variables matter and by how much), stable with limited data, and hard to overfit. More complex models (random forest, gradient boosting) showed marginal accuracy gains but lost interpretability.
Variables and what they mean
The model uses wind, gusts, wave height, wave period and the time of year. Wind is read an hour before each departure and waves three hours before: operators decide before the boat leaves, and those earlier readings predict the outcome better than the conditions at departure. The number after each is its odds ratio in the model currently served.
- Wave Height (2.34x)
- Significant wave height — the average of the tallest third of waves. This is what mariners mean by “seas are 5 feet.” In our training data, cancelled trips average 3.1 ft waves vs. 1.3 ft on normal trips. Measured in feet. If the buoy reports no wave height at all, the trip is scored by a second model trained without waves, never as if the sea were flat.
- Wind Speed (1.81x)
- Sustained wind speed at 10m height. The baseline wind condition — less dramatic than gusts but more persistent. Cancelled trips in our dataset average 24.4 mph sustained wind vs. 13.1 mph on normal trips. Measured in mph.
- Wind Gusts (1.87x)
- Maximum wind speed recorded during the observation period. Gusts cause the sudden lurches that make docking dangerous. Measured in mph.
- Wave Period (1.20x)
- Time between wave crests. Shorter periods (3–5 seconds) mean choppy wind waves; longer periods (8–12 seconds) mean smoother ocean swells. Its effect on its own is small. The buoy often goes hours without reporting it; the model then uses a "wave period missing" marker it was trained with (1.00x), rather than a period of zero.
- Time of year (1.15x)
- The month, encoded as a position on a yearly cycle so December and January sit next to each other. At the same wind speed, winter trips are cancelled more often than summer ones. (The odds ratio shown is the winter-versus-summer term.)
Odds ratios are per 1 standard deviation increase. A ratio of 1.5x means one standard deviation more of that variable multiplies the odds of a cancellation by 1.5; below 1x lowers them. These come from the served model and change when it is retrained.
Fast ferries vs. traditional ferries
Fast ferries (Hy-Line’s Grey Lady, SSA’s Iyanough) are high-speed catamarans that make the crossing in ~1 hour. They’re more sensitive to rough seas due to their hull design and higher operating speed.
Traditional ferries (SSA’s Eagle, Woods Hole) are larger, heavier car/passenger vessels. They take ~2.25 hours but tolerate significantly worse conditions.
How we adjust: The base model produces a single cancellation probability. For fast ferries, we shift the internal score toward cancellation (+0.3 logit). For traditional ferries, we shift it toward running (−0.8 logit). These shifts are set by hand. We tested learning them from the most recent year of trips instead: shifts learned on 2024 made the forecasts for 2025 worse, because SSA cancelled far less often in 2025 than in 2024 in the same weather. We will switch once recent trip records feed the weekly retrain, so the learned shifts keep up.
With the model served now, a calm March day shows 99% for a fast ferry and 99% for a traditional one. In marginal March weather (20 mph wind, 26 mph gusts, 3 ft waves) a fast ferry shows 96% and a traditional ferry 99%.
Classification: All Hy-Line Nantucket service is fast ferry (Grey Lady catamaran). SSA trips are classified by their published vessel type — “Fast Ferry” (Iyanough) vs. “Vehicle/Passenger” (traditional).
Data sources
- Historical cancellation labels
- Trip-level ground truth: SSA cancellation emails, Hy-Line #ALERT feed records via Flockler API, and 33,416 scheduled trips from the ACK Schedules API used as the baseline schedule. Cross-referenced against SSA annual performance reports for QA.
- Current weather: NOAA Buoy 44020
- A physical weather buoy operated by the National Oceanic and Atmospheric Administration, moored directly in Nantucket Sound on the ferry route. Reports wind speed, gusts, direction, wave height, wave period, temperature, and pressure every 10 minutes. Wave height and dominant wave period (DPD) are reported less frequently than wind data and are found independently — wave height is available on roughly a third of readings, while DPD may be unavailable for longer stretches. When the latest reading lacks either value, we backfill from the most recent reading that includes it (scanning up to 12 hours of data). The model was trained exclusively on DPD; we do not substitute the average wave period (APD) as a fallback.
- Forecast weather: Open-Meteo
- Free, open-source weather API that blends multiple models (ECMWF IFS, GFS, ICON, and others) to select the best forecast for a given location. For marine data (wave height, period), it uses ECMWF WAM and NCEP GFS Wave models. No API key required. Each trip is scored on the forecast for its departure hour (a 7:30 PM boat reads the 7 PM hour), including today’s boats more than two hours away; the day summaries sample every three hours (6 AM to 9 PM ET).
- Fallback weather: Open-Meteo current
- If the NOAA buoy is offline (maintenance, communication failure), we fall back to Open-Meteo’s model-derived current conditions for the same coordinates. This is a forecast model’s estimate, not a direct observation, so it’s less accurate. The data source is shown on the Current Conditions card.
- SSA schedule & status
- Scraped from the Steamship Authority website. Includes departure times, vessel types, and real-time status (On Time, Delayed, Cancelled).
- Hy-Line schedule
- Retrieved from ACK Schedules API, a third-party service that publishes Hy-Line’s departure times and route information.
- Hy-Line alerts
- Fetched from Hy-Line’s social media feed via the Flockler API. We check for posts tagged #alert that mention cancellation-related terms (cancel, suspend, weather, wind, sea, storm).
How training data was built
The model is trained on trip-level data: for each individual scheduled departure from Jan 2023 to Dec 2025, was this specific trip cancelled due to weather?
Label sources (33,416 trips total): SSA weather cancellations are from their automated email alert system. Hy-Line cancellations come from their #ALERT social media feed via the Flockler API. The remaining 32,390 scheduled trips without a corresponding cancellation alert were assumed to have operated as scheduled — the standard approach for modeling rare events from alert-based data sources.
SSA cancellations were extracted from the full history of trip-cancellation email alerts. Reasons include weather, mechanical, crew shortage, and other causes — only weather-attributed trips were labeled cancelled.
Hy-Line cancellations were matched from Flockler alert posts to specific ACK Schedules API trip records.
Weather data for each trip was pulled from NOAA Buoy 44020’s historical records. Each departure's Eastern time is converted to the buoy's UTC clock; wind and gusts come from the reading nearest one hour before departure, and wave height and period from the nearest readings to three hours before (within three hours, since the buoy reports waves less often than wind).
SSA annual performance statistics (downloaded from steamshipauthority.com/performance) were cross-referenced with our email data. Note: SSA’s official breakdown categorizes only ~33% of Nantucket cancellations as “weather,” while our email alerts show 64%. The discrepancy likely reflects SSA’s broad “other” category (56% of their total) capturing some weather-adjacent operational decisions that their email alerts explicitly label as weather.
Forecast accuracy by day
Our 7-Day Outlook uses Open-Meteo’s forecast models, which blend ECMWF, GFS, and regional models. Forecast accuracy degrades with lead time — here’s what we know:
| Horizon | Wind accuracy | Wave accuracy | Our confidence |
|---|---|---|---|
| Day 1–2 | RMSE ~2.1 m/s (~4.7 mph) | High skill, RMSE <0.3 m | High |
| Day 3–4 | RMSE ~2.3 m/s (~5.1 mph) | Good skill, RMSE ~0.3–0.5 m | Moderate |
| Day 5–7 | RMSE ~2.4+ m/s (~5.4+ mph) | Useful for large signals only | Lower |
Wind RMSE figures are from ECMWF verification reports. Open-Meteo does not publish its own accuracy statistics. Wave forecast verification from the WMO Lead Centre indicates wave forecasts are “useful to day 5 in the Northern Hemisphere.”
What this means for our predictions: Cancellation days in our dataset average 24.4 mph wind and 3.1 ft waves. A 5 mph RMSE at day 5 could swing a 15 mph forecast to 20 mph (still below typical cancellation conditions) or a 20 mph forecast to 25 mph (now in the danger zone). Days 1–3 are reliable. Days 4–5 are useful. Days 6–7 should be treated as directional only — good for spotting obvious storm patterns, not for making travel decisions.
How forecast days are scored: the forecast for each trip’s departure hour goes through the same model, then a correction fitted separately for today, tomorrow, days 2–3 and days 4–7 on archived forecasts and the trips that actually ran or were cancelled. Very low chances show as “<10%” rather than a single-digit number: a forecast days ahead cannot be that precise. Since September 2026 we also keep a copy of each day’s forecast, because Open-Meteo keeps no archive of its wave forecasts; in time that will show how far ahead the wave forecasts can be trusted.
Assumptions and limitations
- Weather only. We cannot predict mechanical failures, crew shortages, or operational decisions. SSA attributes about 10% of cancellations to mechanical causes, with a large “other” bucket covering the rest. Hy-Line’s non-weather cancellations are ~15%.
- Single buoy. Conditions at Buoy 44020 represent mid-Sound conditions. Harbor conditions at Hyannis or Nantucket may differ, especially in fog.
- Captain’s discretion. Cancellation decisions are ultimately made by the vessel captain and operator dispatch. Two identical weather days can have different outcomes depending on the captain, vessel condition, passenger load, and tidal state.
- Schedule approximation. Tomorrow’s trips use today’s published schedule. If the operator adjusts the schedule overnight (e.g., adding or removing trips in response to weather), our departure list may be stale.
- Forecasts are not buoy readings. The model learned from the buoy, and Open-Meteo’s waves run higher than the buoy’s. Fed in as they are, forecasts overstate the chance of cancellation, so forecast-based chances are recalibrated for each lead time on two years of archived forecasts. On the stormiest forecast days in that archive, about 3 in 4 of the trips given the lowest chances were cancelled; chances below 10% show as “<10%”. Archived wave forecasts do not exist, so the wave part of that recalibration rests on wave analyses rather than true forecasts.
- Coastal accuracy. Open-Meteo’s marine API notes that “accuracy at coastal areas is limited.” Nantucket Sound is semi-enclosed, which may cause local wave patterns that global models underestimate.
Why weather-only predictions
There are two reasonable approaches to predicting ferry cancellations: model the base rate (what percentage of trips get cancelled on any given day?) or model weather-conditional probability (given today’s specific weather, what’s the chance of cancellation?).
Base rate approach: About 11.6% of days in our dataset had at least one weather cancellation. You could simply say “there’s an 88% chance boats will run today” every single day. This would be correct 88% of the time — but it tells you nothing useful. It can’t distinguish a calm July morning from a nor’easter.
Weather-conditional approach (what we use): We feed current wind speed, gusts, wave height, and other conditions into a logistic regression model. On a calm day, the model returns 99%. When a storm approaches, it drops to 60% or lower. This is genuinely useful — it answers the question people actually ask: “Should I worry about my ferry today?”
What we intentionally leave out: Mechanical failures, crew shortages, and operational decisions. SSA officially attributes about 10% of cancellations to mechanical causes and 56% to a broad “other” category (which likely includes weather-adjacent operational decisions). Hy-Line’s non-weather cancellations are ~15%. We don’t model these because they’re effectively random from a passenger’s perspective — there’s no public data that would let us predict them. Including them would add noise without adding predictive value.
Transparency note: When the model says 95%, it means “weather conditions are fine for running.” There’s still a small residual risk from non-weather factors that the model doesn’t capture. We think this is more honest than inflating uncertainty to cover unknowable events.
Continuous improvement
The model retrains every Sunday on the trip-level dataset. Each retrain fits a fresh calibration and is tested on trips from Jan 2025 on with a copy trained only on earlier trips. It replaces the served model only if that test is no worse than a fixed baseline by more than 0.01 AUC, 0.02 windy-trip AUC and 3 points of Brier skill. The bar stays fixed, so small weekly slips add up and fail rather than drifting through.
Daily weather snapshots are recorded but not trained on yet: they are day-level, and the model learns from individual trips. They will be added once each trip’s operator status is recorded.