NASA and IBM announced on Wednesday the open-source release of the NASA-IBM Lunar Foundation Model, an AI model designed to help scientists extract insights from decades of lunar observations — one of the first publicly available foundation models built for scientific exploration of the Moon. The model is available now, with Computer Weekly reporting that it can be downloaded from Hugging Face.

The release targets a problem familiar to planetary scientists: too much data, too few hours in the day. For decades, sensors and instruments have continuously observed the Moon, generating petabytes of information. To study the surface, scientists have had to sift through maps and images by hand or rely on low-resolution, task-specific machine learning models — approaches IBM describes as computationally intensive and often lacking the accuracy needed to identify and analyze geographic features. For more context on this story, see our ongoing AI news.

Finding Ice, Craters, and Ancient Volcanoes

The model is trained to surface hidden relationships across many different types and resolutions of lunar data, and IBM's announcement outlines three flagship applications.

Potential ice deposits. Permanently shadowed regions near the lunar poles are among the hardest places on the Moon to observe, yet they may hold ice below the surface — a resource considered essential for a future Moon base and for producing rocket fuel for missions to Mars. In a technical paper authored by the NASA-IBM team, the model reduced error (RMSE) in identifying areas with high potential for lunar ice by up to 22 percent compared with the widely used SwinV2-B (ImageNet) vision model. Crater detection. Craters reveal clues about the age of terrains, lunar geology, and the chemical composition of the early lunar interior, and mapping them helps NASA select safe landing sites and avoid hazards like steep slopes and boulders. The model can identify and classify craters at meter-scale resolution with accuracy comparable to state-of-the-art models while offering greater efficiency and lower fine-tuning costs, IBM said. At roughly 100-meter resolution, it outperforms SwinV2-B by nearly 19 percent using just half the training data. Volcanic history. The model also tracks Irregular Mare Patches — enigmatic volcanic features that help scientists understand the Moon's thermal evolution. Even when fed imperfect labels, it captures the extent of these features 3 percent better than the SwinV2-B baseline, according to the announcement.

Taken together, IBM says the model exceeds widely used methods by up to 23 percent in identifying key geographic features on the lunar surface.

The First Unified Open Lunar Dataset

Alongside the model, IBM and NASA scientists built what the companies describe as the first open-source lunar dataset of its kind: a unified, machine-learning-ready collection aggregating more than 30 spatially aligned layers from nine instruments across four missions. It combines tens of thousands of images and maps capturing the geophysical properties of the lunar surface, drawing on NASA's Lunar Reconnaissance Orbiter (LRO) and GRAIL missions, among others. Trade publication Unite.ai has described the accompanying benchmark dataset as SomBench.

The dataset matters as much as the model. Until now, IBM noted, no publicly available dataset brought the Moon's multimodal, multi-resolution observations into a common framework suitable for modern machine learning — meaning any research group wanting to apply deep learning to lunar science first faced months of data engineering.

What NASA and IBM Are Saying

"NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job," said Kevin Murphy, chief science data officer and acting chief data and AI officer at NASA Headquarters. "We also have to make data easier for scientists to explore and use. The NASA-IBM Lunar Foundation Model shows what's possible when we bring AI to NASA's petabytes of scientific data."

"Uncovering the mysteries of the Moon requires an ability to learn from an extraordinary volume of scientific data," said Juan Bernabe-Moreno, Director of IBM Research Europe, UK and Ireland. "The NASA-IBM Lunar Foundation Model gives scientists a foundation to explore the Moon at scale, connecting observations across instruments, revealing patterns that are difficult to see in isolation, and providing an open platform the global research community can build on."

Why Open Source Matters Here

The open-source framing is deliberate. By releasing weights publicly, NASA and IBM are inviting laboratories, universities, and space agencies worldwide to fine-tune the model for their own questions — from landing-site selection for crewed missions to fundamental studies of how the Moon's surface evolved. Lower fine-tuning costs, IBM said, make that practical even for small teams.

With multiple space programs working toward a sustained human presence on the Moon, the resources at stake are significant: water ice implies oxygen to breathe and hydrogen-oxygen propellant to refuel rockets for the journey to Mars. Mapping where that ice sits, and where the surface is safe to land, is foundational work — and it is exactly the kind of pattern-recognition problem foundation models have proven good at.

The release also signals a broader trend in AI for science: rather than chasing chatbots, public research institutions are increasingly pointing foundation-model techniques at the mountain of unanalyzed scientific data their instruments have accumulated. The Moon, it turns out, is a good place to start.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →