Skip to content
AI Atlas
PaperActive

A Station-Based Evaluation of Machine Learning-based Weather Forecasting Models in Northern Norway

arxiv.org/abs/2609.10564

quality89

Updated 1 h ago · first seen 11 Sept 2026

paper_01M294FPMNFQ55YRVYVFFEATEN

Published
11 Sept 2026
T1 · 1 h ago
arXiv
2609.10564
T1 · 1 h ago
Category
physics.ao-ph
T1 · 1 h ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history · Abstract

1 claims · 1 propertiesShow all properties

Abstractabstract1

Claim history for Abstract
ValueValid from → toStatusSourceConfidenceExtractor
Recent machine learning weather prediction (MLWP) models have demonstrated remarkable forecasting skill on global reanalysis-based benchmarks. However, their performance remains unclear in challenging environments such as Northern Norway, where narrow fjords and rapidly changing weather result in highly variable local wind conditions. In this case study, we evaluate FourCastNet3 (FCN3), GraphCast, and ECMWF High Resolution Forecast (HRES) for wind speed forecasting using multi-year station observations from Northern Norway, focusing on their relative performance, generalization beyond the training period, and performance under high-wind conditions. Our results show that HRES slightly outperforms FCN3 and GraphCast, with an overall RMSE of 2.89 $\mathrm{m\,s^{-1}}$, compared to 2.96 $\mathrm{m\,s^{-1}}$ for FCN3 and 2.94 $\mathrm{m\,s^{-1}}$ for GraphCast. Notably, the MLWP models maintain comparable performance beyond their respective training periods, with no clear evidence of noticeable degradation. FCN3 performs best under high-wind conditions, although all models substantially underestimate strong winds. Our findings suggest that MLWP has become competitive with NWP for local wind, but further refinements are still needed to capture complex terrain better.currentcurrentarXiv (Atom API + RSS)T1highdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →