The core development
Google DeepMind has introduced AlphaEarth Foundations, an AI model designed to integrate massive Earth observation datasets into a unified digital representation, and has released annual outputs from the model as the Satellite Embedding dataset in Google Earth Engine.
The announcement, dated July 30, 2025, frames the model as a kind of “virtual satellite”: not a new spacecraft, but a system that can combine many streams of satellite and environmental data into a consistent computational layer for mapping Earth’s terrestrial land and coastal waters.
Why Earth observation needs a new layer
Satellites already provide frequent, information-rich views of the planet. The challenge is that these observations are fragmented. Optical imagery, radar, 3D laser mapping, climate simulations and other public data sources differ in format, timing, coverage and reliability. Cloud cover, irregular revisit schedules and uneven measurement types can make it difficult to compare one region or year with another.
AlphaEarth Foundations addresses this by producing an embedding for each location. In AI, an embedding is a compact numerical representation of complex information; instead of storing only a raw image, the model encodes many signals into a vector that computers can compare, search and use for downstream mapping tasks.
DeepMind says the model represents land and coastal waters in 10 by 10 meter squares and can track changes over time. Each embedding has 64 components, and the company’s visual examples show how selecting three dimensions and mapping them to red, green and blue can reveal agricultural fields in cloudy Ecuador, complex Antarctic surfaces and subtle Canadian agricultural land-use differences.
Key numbers and technical claims
DeepMind presents AlphaEarth Foundations as a response to two problems: too much data, and inconsistent data. The model combines information from dozens of public sources and produces compact summaries that are easier to use at planetary scale.
Notable figures include:
- Spatial unit: 10×10 meter squares across terrestrial land and coastal waters;
- Representation: 64-component embeddings;
- Dataset scale: more than 1.4 trillion embedding footprints per year in the Satellite Embedding dataset;
- Efficiency: summaries require 16 times less storage than those produced by other AI systems tested by DeepMind;
- Performance: an average 24% lower error rate than tested models;
- External testing: more than 50 organizations worked with the dataset over the past year.
DeepMind says the model performed well across tasks such as identifying land use and estimating surface properties, including situations where labeled data was scarce. Labeled data refers to examples that have already been categorized or verified, often by humans, and is a key resource for training and evaluating AI models.
Early users and practical mapping cases
The Satellite Embedding dataset is available through Google Earth Engine, a cloud platform widely used for geospatial analysis. By making annual embeddings available there, DeepMind is positioning the model as infrastructure for custom map generation rather than as a single finished map product.
Organizations mentioned in connection with the dataset include the United Nations Food and Agriculture Organization, Harvard Forest, Group on Earth Observations, MapBiomas, Oregon State University, the Spatial Informatics Group and Stanford University.
One highlighted use case is the Global Ecosystems Atlas, which aims to create a comprehensive resource for mapping and monitoring ecosystems. The project is using the dataset to help countries classify unmapped ecosystems, including categories such as coastal shrublands and hyper-arid deserts. Another example is Brazil’s MapBiomas, which is testing the dataset to better understand agricultural and environmental changes across the country, including in critical ecosystems such as the Amazon rainforest.
What this signals for geospatial AI
AlphaEarth Foundations reflects a broader shift in geospatial technology: moving from individual images and sensor products toward reusable AI foundation layers. If the approach proves reliable in more real-world settings, it could reduce the time and storage required to build maps for agriculture, deforestation monitoring, urban expansion, water resources and conservation planning.
The model does not remove the need for field surveys, local expertise or careful validation. But it could give scientists and public-interest organizations a more consistent starting point, especially in regions where data is fragmented or labels are limited. The next test is whether annual embeddings can move beyond technical demonstrations and become dependable inputs for operational decisions about food security, biodiversity and land management.




