Editorial illustration for NASA, IBM Release Lunar Science Model Built on 17 Years of Orbiter Data
NASA, IBM Release Lunar Science Model Built on 17 Years...
NASA's Lunar Reconnaissance Orbiter has circled the Moon since 2009, racking up 17 years of combined observation data across its instruments. That archive is massive, but most of it has sat underused because turning raw orbiter images into something a machine learning algorithm can act on takes labeled examples, and labels are expensive to produce at lunar scale. NASA and IBM Research, working with several academic partners, built a fix for that bottleneck: the NASA-IBM Lunar Foundation Model, now released as open source.
The model is trained on SomBench, a dataset the team calls the largest co-registered multimodal lunar corpus assembled so far, pulling together nearly 2 million tile bundles across 11 modalities and two spatial scales. Roughly 1 million of those come from the Narrow Angle Camera at about 1 meter per pixel resolution, with another 964,000 from the Wide Angle Camera at 100 meters per pixel. Instead of building a separate algorithm for every task, researchers can adapt this one pretrained model to specific jobs, like flagging ice deposits near the poles or picking out craters, using only a handful of labeled examples.
The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.
Why this matters
For anyone building AI tools on scientific data, this is a template worth watching. NASA and IBM didn't just publish a paper, they released an open source foundation model trained on 17 years of LRO observations plus GRAIL, Lunar Prospector, and Kaguya data, then pointed it at two concrete tasks: finding polar ice and detecting craters. That's a useful contrast to the usual "general-purpose" framing around foundation models.
Kevin Murphy's point, that collecting data is only half the job, is the real story here. Government agencies sit on mountains of unlabeled, multi-instrument data that's expensive to curate and harder to make ML-ready. If this model performs well on ice prediction, where false negatives matter for future lunar missions, it's a working example of how domain-specific foundation models can outperform bolting a generic model onto raw sensor data.
Researchers should watch benchmark numbers once they're public. Founders building vertical AI tools for geology, climate, or remote sensing should take note: NASA just handed the field a blueprint for turning fragmented archives into usable training infrastructure.
Further Reading
- NASA, IBM Launch AI Foundation Model for Lunar Science - NASA Science
- IBM and NASA Release Open-Source AI Model to Support Lunar Exploration - IBM Newsroom
- IBM, NASA launch AI model to help map ice, craters on Moon - Reuters
- NASA, IBM launch new AI model for studying the moon - Space.com
- NASA and IBM Launch Open-Source AI Model for Future Moon Exploration - CNET