Breaking Biodiversity Data Silos: LUCA's Unified Ingestion Pipeline and Species Richness API
Session Description
SimplexDNA introduced LUCA, a new platform that pulls together two of the world's largest and most siloed biodiversity datasets, environmental DNA records and bioacoustic recordings, into a single searchable map of life on Earth. The session covered how biased and incomplete existing biodiversity data really is, how LUCA works, and the tricky open questions around who gets paid, who owns the data, and how to keep quality high as more organisations start feeding data into a shared commons.
Speakers
Kristy Deiner, SimplexDNA
Christopher Hempel, SimplexDNA
Watch the Session Recording
Key Takeaways
Existing global biodiversity data is heavily skewed. Most records come from a small number of countries, and roughly two-thirds of all entries in GBIF, the main open biodiversity database, are about birds. That leaves most of the planet essentially undocumented.
Environmental DNA (eDNA) works by sampling water or soil and sequencing the DNA in it, then matching those sequences against a reference database to identify which species were present. A single sample can reveal hundreds or thousands of species at once, far more than the one or two species most customers actually ask for.
SimplexDNA built LUCA by pulling eDNA records out of the Short Read Archive, a public scientific repository most people outside genetics have never heard of, and reprocessing years of previously siloed data into a usable, searchable layer.
LUCA also incorporates bioacoustic data from Rainforest Connection's sound recorders and the Arbimon platform, described in the session as possibly the largest sound database in the world, converting recorded animal sounds into species identifications the same way eDNA converts sequences into species identifications.
The platform currently holds around 25,000 distinct organisms and roughly 2.6 million observations, growing by an average of about 11,000 new observations a month.
Basic exploration of LUCA's data (an interactive global map broken into hexagons showing observation counts, species, and conservation status per area) is free with no account needed. Downloading detailed records, including coordinates and full taxonomic detail, requires a free account, currently offered on a three-month trial.
SimplexDNA is designing LUCA around four types of users: subscribers who pay to access and download data, suppliers who contribute their own data and get paid when it's used, data stewards who help set ethical and legal standards for what can be shared, and gap fillers who commission new data collection for locations where none currently exists.
The team named three specific data gaps LUCA aims to help close: a taxonomic gap (data overwhelmingly favours birds over other life), a spatial gap (most of the Earth's surface has no biodiversity data at all), and a temporal gap, described as the most important one for actual monitoring and restoration work, since real conservation impact requires repeated measurements over time, not one-off snapshots.
SimplexDNA plans to shift ownership of LUCA and most of its own company shares into a nonprofit foundation, similar to Patagonia's ownership structure, and to build a linked financial mechanism called the Proof of Life Protocol, where a tokenized asset (referred to in the session as a "Franklin") earns a return tied to how much a given dataset gets used, creating a financial incentive for organisations to contribute data rather than keep it locked away.
SimplexDNA has secured an 11 million euro EU Horizon grant with a consortium of 17 partners to apply its river-sampling model for eDNA across all of Europe by 2029, building on a Switzerland-specific pilot funded separately.
A question from UNEP-WCMC about quality assurance drew a direct answer: SimplexDNA is relying on domain experts to vet incoming data and plans to hand governance of data standards to the same nonprofit foundation, rather than keeping quality control as an internal, for-profit function.