Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models
Updated 46 min ago · first seen 11 Sept 2026
paper_01M294FP1BFBN7F4KJP96T50FN
- Published
- 11 Sept 2026
- T1 · 46 min ago
- arXiv
- 2609.11135
- T1 · 46 min ago
- Category
- cs.LG
- T1 · 46 min ago
Abstract
Short-term photovoltaic (PV) power and global horizontal irradiance (GHI) forecasts are essential for effective dispatch, reserve scheduling, and grid operations. At these forecasting horizons, errors are predominantly driven by cloud induced ramps: relying solely on historical numerical data may struggle to anticipate an incoming cloud, making ground-based sky images a crucial complementary physical signal. Furthermore, forecast performance is highly sensitive to location and local observing conditions, creating a strong need for site-specific data that are often scarce. Recently, large language models (LLMs) have demonstrated competitive performance and high data efficiency in time-series forecasting. Despite their success, existing LLM-based forecasting methods remain predominantly unimodal, relying primarily on historical numerical time-series data. Effectively incorporating sky imagery into an LLM-based forecasting framework remains under-explored and an open challenge. In this paper, we propose SolCloudLLM, an LLM-based multimodal forecasting framework. SolCloudLLM aligns sky-image patches with time-series patches and fuses their corresponding representations through bidirectional multimodal fusion, yielding a unified representation that is subsequently mapped into the embedding space of an LLM. Extensive experiments on the SIRTA and SKIPP'D datasets demonstrate that SolCloudLLM consistently outperforms the best baseline methods in MSE across all forecasting horizons, achieving a maximum relative MSE reduction of 25.4%. Stratified analysis further indicates that the benefits of multimodal fusion are concentrated primarily under cloudy conditions. Notably, SolCloudLLM achieves the best performance in nearly all few-shot settings, whereas other deep learning baselines experience substantial performance degradation and are frequently outperformed by the non-learning physical method.
Authors 6
Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge
Specification
- Official page
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Arxiv announce type
- new
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- arXiv id
- 2609.11135
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Categories
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Primary category
- cs.LG
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
- Published
- 11 Sept 2026
Source:arXiv (Atom API + RSS)T1observed 46 min agohigh
Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →
Provenance
Attributed facts
9
Source tiers
T19
Freshest observation
46 min ago
Conflicts
None
No models linked to this paper yet.
- Authors
- Ken Chen, Maneesha Perera, Wei Wang
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history · Published
Publishedpublished_at1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| 11 Sept 2026 | → current | current | arXiv (Atom API + RSS)T1 | high | deterministic |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
- New paperPaperBidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models
New paper: Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models
arxiv
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| arXiv (Atom API + RSS) | rss.arxiv.org/rss/cs.LG | feed | T1· Official | 46 min ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.