Nature benchmarks to accelerate nature tech
Session Description
Vizzuality opened by defining benchmarks narrowly: not "benchmark" in the baseline-reference sense that ecology often uses, but shared, reproducible evaluation frameworks made up of a defined task, reference data, and agreed metrics, the kind of thing that let ImageNet drive a decade of progress in computer vision. They mapped the nature tech pipeline into four levels (field observations, modelled field outputs, derived metrics, and agentic decision-support systems) and showed how other fields have benchmarked at each level, from CASP's 25-year run ahead of AlphaFold to HealthBench's expert-scored patient conversations. Their read: nature tech has real benchmarks at the observation end (BirdCLEF, PlantNet, GeoLifeCLEF) but gets sparser and less coordinated the further up the pipeline you go.
Attendees then split into four groups by pipeline stage and worked through two rounds: first sharing what they're currently measuring and how, then trying to get specific about what ground truth would even look like for their stage and who'd need to own it.
Speakers
Mike Hartford, Vizzuality
Francis Gassert, Vizzuality
Watch the Session Recording
Key Takeaways
Nature tech doesn't have an ImageNet or CASP-scale benchmark yet to organise investment around, and the four pipeline stages (observation, modelled output, derived metric, agentic) need different benchmarking approaches, not one shared standard.
Derived metrics like BII and the Human Modification Index often trace back to the same handful of satellite and public datasets. When two independently-built indices agree, that may just mean they share inputs, not that either has been independently validated.
"Ground truth" turned into a genuine sticking point. Some things have a definitive answer if you're willing to pay for it (what species is in this image); others (is this ecosystem degraded, is this worth protecting) are judgment calls all the way down, and groups didn't agree on whether that's permanent or just a current limitation.
STAR (Species Threat Abatement and Restoration) came up as a live cautionary example: newer site-level ground-truthing work found the calibrated, on-the-ground version diverged significantly from the original modelled STAR score, prompting work toward a STAR 2.0.
The agentic/decision-support layer already has something close to a working benchmark loop for product performance: a roughly 1,000-item expert-validated "golden set" used to test whether model or system changes hold up (it caught, for instance, that a switch to a smaller, cheaper model held accuracy steady). Scientific accuracy benchmarking is progressing but earlier-stage. Benchmarking trust and communication quality (can a system convey that losing a quarter of a small forest isn't the same severity as losing a quarter of a large one) has no established methodology yet.
A recurring proposal across groups: some kind of shared registry, potentially government or fund-backed, where organisations could disclose what ecological data they already hold and where, anonymised if needed, since a lot of relevant data already exists privately but isn't discoverable.
Multiple groups converged on a bigger gap than the metrics themselves: nobody has mapped the actual range of real-world use cases these derived metrics get used for, which makes it hard to know what a benchmark should even be evaluating against.
Relevant Readings
Nature Benchmarks to accelerate Nature Tech - Workshop Summary