A Scalable Mega-Data Management Framework for Global Offshore Wind Resource Assessment: From ERA5 Reanalysis to Machine-Learning-Ready Datasets
Contenido principal del artículo
Resumen
This paper presents a scalable mega-data management framework that transforms raw ERA5 reanalysis data into machine-learning-ready feature matrices for global offshore wind resource assessment. The framework processes 49.4 billion individual data values derived from 134,475 offshore grid points across 240 geographic tiles (10°×10°) spanning 11 sampled years (2004–2024, biennial intervals) yielding 12,672 timesteps per point. A five-stage feature engineering pipeline converts 5 raw ERA5 variables into 29 physics-informed features through physical derivation, cyclic sinusoidal encoding of temporal and directional variables, thermodynamic interaction terms, multi-scale rolling statistics, and per-tile Z-score normalization. The pipeline addresses three gaps in the offshore wind literature: the absence of documented end-to-end data management at global scale, the lack of reproducible feature engineering for wind energy machine learning, and missing computational benchmarks for mega-data processing. Preprocessing fidelity validation confirms that normalization preserves physical statistics within acceptable thresholds (ΔMean < 10%, ΔStd < 10%, RMSE < 0.72 m/s) across representative pilot tiles spanning tropical, temperate, and polar zones. Unsupervised clustering validation using six algorithms demonstrates that the processed dataset produces well-separated offshore wind regimes (silhouette coefficients 0.54–0.63 for the physics-aware method), confirming machine-learning readiness. The framework establishes a reproducible, modular pipeline for downstream energy, exergy, and predictive modelling applications.
Detalles del artículo
Sección
© SEECMAR | All rights reserved