The physical world is complex, dynamic, and constantly changing. Yet while AI has transformed how we work with language, code, and structured data, understanding the physical world remains a challenge. Geospatial information is vast, fragmented, and difficult for large language models to reason over, from satellite imagery and weather conditions to infrastructure and terrain.
Today at a16z’s Speedrun AI Faire, we’re launching a preview of Terra, an AI agent built on our patented spatial reasoning technology that enables language models to understand and analyze complex geospatial data. Terra executes complex geospatial workflows—from identifying optimal sites and monitoring real-time weather impacts to forecasting spatial trends and generating detailed maps—all through natural language.
How it works
Terra turns a natural-language request into an executable geospatial workflow. You describe the analysis you need; Terra investigates the available data, selects the computational tools, tests its approach, and assembles the steps required to produce the result.
Behind that experience are three core components:
Data Catalog: Terra explores a catalog of public datasets and connected private data to identify the inputs needed for a task. It inspects their structure and coverage to understand what they can support. The catalog uses open geospatial standards, including STAC and OGC standards, to publish and consume datasets, connecting Terra to existing geospatial infrastructure.
Processors: Terra’s processors perform the underlying analysis. Each processor has a defined interface specifying the data it accepts, the parameters it requires, and the outputs it produces. These interfaces give the agent a concrete way to determine which operations can answer a question and how to connect them.
Workflow System: The workflow system connects those processors, manages execution, and records the steps and parameters that produced each result. This record makes it possible to trace an output back through the analysis that created it.
Terra brings these capabilities together through a research, test, and validate loop, with distinct planning and execution phases.
During planning, Terra explores datasets and processors, tests candidate operations, inspects intermediate outputs, and revises its approach. The resulting plan is an executable workflow, displayed as a diagram so users can inspect the steps before execution.
During execution, the workflow system runs the analysis. Terra then inspects the resulting data and artifacts to assess whether they satisfy the request. It checks both whether the workflow ran successfully and whether the results meet the requested requirements, including the presence of expected map layers and outputs that match the specified parameters.
This architecture connects the flexibility of a language model with explicit computational tools and traceable workflows. Users can see both the result and the process that produced it.
Terra in Action: Mapping Snowfall
To see how Terra moves from a request to a validated result, consider this question from the benchmark:
Chart contour lines for accumulated snowfall in winter 2023–2024 in the USA.
Contour lines connect locations that received the same accumulated snowfall. For this evaluation, the source dataset and six contour levels were specified in advance: 20, 40, 80, 160, 320 and 640 inches.
Terra worked through the task in four steps. The task’s evidence page provides the recorded tool activity and verification results.
Understand the data. Terra inspected the supplied NOAA snowfall dataset, checking its units, geographic coverage, and missing data.
Recorded evidence
Terra called get_collection, get_item and describe_processors to inspect the supplied dataset and available computation. It then validated the processor requests with validate_processor.
Test the approach. After inspecting the data, Terra tested its approach on a smaller geographic area, reviewing the generated contour lines, summary statistics, and map preview before applying it to the full dataset.
Recorded evidence
The recorded findings included 2,657 contour features, six distinct requested levels, and 1,032,448 missing source cells that remained unknown. Terra recorded the following question for its test: “Does the supplied 2023–24 snowfall raster produce mask-aware contour lines at the requested inch levels on band 1, and what are the complete-source missing-cell count and observed range needed to report nonempty levels?”
Build and execute the workflow. After confirming that the test approach worked as intended, Terra built and executed a workflow to generate contour lines and compute summary statistics across the full dataset.
Recorded evidence
Terra called propose_workflow, followed by execute_workflow. The saved execution records show that both contour generation and source-statistics computation succeeded.
Validate the results. Terra inspected the workflow outputs, generated artifacts, and map layers. It checked that the map contained the expected data, reviewed the contour count, and confirmed that a supporting data table was present.
Recorded evidence
Terra called inspect_workflow_results, inspect_artifact and list_map_layers before recording its final assessment.
The benchmark harness then independently verified six expected contour levels and six observed levels. Both the computation and the requested answer passed the evaluation.
The full analysis took approximately 10 minutes, from initial data inspection through final validation. The table below shows the execution timeline.
Full Benchmark Results
We evaluated Terra using questions from GeoBenchX, adapting the execution and verification process to our processor-and-workflow system. This evaluates Terra as a complete system, including its model, tools, workflows, and outputs.
These results measure Terra’s task success under our evaluation setup. They are not directly comparable to the published GeoBenchX scores, which used different models, tools, and execution budgets.
Using GPT‑5.4, Terra passed 182 of 201 measured tasks, for an overall task success rate of 90.5%. GeoBenchX contains 202 questions. One question could not be evaluated because its required source data was unavailable and was excluded from the measured total.

What counts as a pass?
For computational tasks, we checked Terra’s outputs against reference results using task-specific criteria, including numerical accuracy, geometry, geographic coverage, and delivery of the requested map.
For qualitative tasks, we assessed the work against predefined criteria, including whether Terra identified limitations in the available source data or missing computational capabilities.
A task-by-task breakdown is available on our study page, with evaluation materials and execution records on GitHub.
Raising the Bar for Geospatial AI Evaluation
This evaluation gives us a starting point for measuring Terra’s capabilities. The failed tasks also point to concrete opportunities to improve our computational workflow layer and increase task success.
We see a need for a new benchmark that tests geospatial reasoning under the conditions users encounter in practice: fragmented sources, inconsistent reporting, nuanced datasets, and subtle errors that emerge when multiple computational steps are combined.
That means evaluating tasks such as:
Political sentiment toward data centers. Assess local support and opposition by county, tackling the challenge of measuring political sentiment from incomplete, qualitative, and often conflicting evidence.
Water access across counties. Analyze fragmented and inconsistent source data, reconciling differences in how counties define, record, and report water access.
Parcel-level site analysis. Evaluate individual parcels for water and road access, then aggregate the findings by county while preserving the detail behind each result.
To try Terra as we expand access, join the waitlist.






