Citizen science & machine learning for plant Identification: challenges from Pl@ntNet



Joseph Salmon

IMAG, Univ Montpellier, CNRS, Inria, Montpellier, France

Consortium Pl@ntnet

Plant importance


The “weight” of plants on Earth

Biomass on earth (Bar-On et al., 2018)

The services they provide to us

  • Provisioning: Food, Medicines …

Honey bee John Severns
Public domain

  • Supporting: Photosynthesis, soil formation, water cycle, …

Leaves on the ground, in a forest Malene Thyssen
Public domain

Leaves on the ground, in a forest Lynette Schimming − EOL
CC-BY-NC

  • Regulating: Pollination, Climate, …

Leaves on the ground, in a forest fuerst.ernst
CC-BY-SA

  • Cultural: Places for recreation, wonder, …

Pl@ntNet insight



Tree with roots

Pl@ntNet: ML for citizen science


A machine learning-driven citizen science platform for mobile plant identification





Pl@ntNet & Cooperative Learning

Chronology of Pl@ntNet


Note: I am mostly innocent, I started working with the Pl@ntNet team in 2020

Scientific challenges


A multidisciplinary collaboration:

  • Stat / ML researchers
  • Ecologists
  • Engineers
  • Citizen contributors

Addressing open challenges:

  • Theoretical foundations
  • Methodological advances
  • Large-scale computational needs
  • Cross-disciplinary integration

We need you: come and help us improve it, one way or another!

The current Pl@ntNet Team

Alexis Joly
Primary investigator, INRIA
ResearchGate

Pierre Bonnet
Primary investigator, CIRAD
ResearchGate

Jean-Christophe Lombardo
Software Manager, IA engineer, INRIA
LinkedIn

Hervé Goëau
Researcher, CIRAD
ResearchGate

Christophe Botella
Researcher, INRIA
ResearchGate

Joseph Salmon
Researcher, INRIA
Website

Benjamin Bourel
Researcher, INRIA
Website

Lee Sue Han
Senior Lecturer, Swinburne University of Technology Sarawak Campus
Scholar

Fabrice Vinatier
Researcher, INRAE
ResearchGate

Lydia Bousset-Vaslin
Researcher, INRAE
Website

Diego Marcos
Researcher, INRIA
ResearchGate

Vanessa Hequet
Botanist, IRD
LinkedIn

Murielle Simo-Droissart
Botanist, IRD
ResearchGate

Jean-Marc Sadaillan
Project manager, INRAE
LinkedIn

Antoine Affouard
Backend & DB engineer, INRIA
LinkedIn

Mathias Chouet
Backend engineer, CIRAD
GitHub

Hugo Gresse
Mobile engineer, INRIA
GitHub | LinkedIn

Thomas Paillot
Front engineer, INRIA
LinkedIn

Rémi Palard
Geo & Fullstack engineer, CIRAD
LinkedIn

Théo Simoes
Backend engineer, INRAE
LinkedIn

Théo Larcher
PhD candidate, INRIA
LinkedIn

Giulio Martellucci
PhD candidate, INRIA
LinkedIn

Raphaël Benerradi
PhD candidate, INRIA
LinkedIn

Ilyass Moummad
Post-doc, INRIA
Personal website | Google Scholar

Some personal contributions



  • Pl@ntNet-300K (Garcin et al., 2021): Creation and release of a large-scale dataset sharing the same property (Long Tail!) as Pl@ntNet; available for the community to improve learning systems


  • Prediction uncertainty quantification
    with long tail data (Ding et al., 2026): providing prediction sets with statistical guarantees with Conformal Prediction


  • Learning & crowd-sourced data (Lefort et al., 2024; Lefort et al., 2025): How to leverage multiple labels per image to improve the model? Need to assert quality: the workers, the images/labels, the model, etc.


Pl@ntNet,
challenging standard ML assumptions Tree with roots

Pl@ntNet classification challenges




Multi-class supervised learning pushed to the limit:



Massive scale

Millions of images across thousands of species (classes)

Structured labels

A hierarchical taxonomy, with visually similar species prone to confusion

Biased data

A citizen-science dataset, unevenly sampled across species and regions

Pl@ntNet : a digital commons project



9 540 044

Users with accounts

3X more users without accounts

86 814

(WCVP) Species with photos

out of 404 389

1 511 221 837

Queries

unlabeled images : without species

32 952 432

Observations

labeled images: with species

Diversity of plants Rkitko
(Wikimedia Commons)

436 TB

Storage

Source: Pl@ntNet stats

WCVP: World Checklist of Vascular Plants database.

Structured labels: the tree of life



Intra-class variability

Guizotia abyssinica (L.f.) Cass. Benoît Janichon
CC-BY-SA
Guizotia abyssinica (L.f.) Cass.
Diascia rigescens E.Mey. ex Benth. Patrice SIROT
CC-BY-SA
Diascia rigescens E.Mey. ex Benth.
Lapageria rosea Ruiz & Pav. Borquez Vicent
CC-BY-SA
Lapageria rosea Ruiz & Pav.
Casuarina cunninghamiana Miq. LC Наталья
CC-BY-SA
Casuarina cunninghamiana Miq. LC
Guizotia abyssinica (L.f.) Cass. Annette Bejany
CC-BY-SA
Guizotia abyssinica (L.f.) Cass.
Diascia rigescens E.Mey. ex Benth. A Lee
CC-BY-SA
Diascia rigescens E.Mey. ex Benth.
Lapageria rosea Ruiz & Pav.
Daniel Barthelemy
CC-BY-SA
Lapageria rosea Ruiz & Pav.
Casuarina cunninghamiana Miq. LC Campos Ignacio
CC-BY-SA
Casuarina cunninghamiana Miq. LC
Based on pictures only, plant species are challenging to discriminate!

Inter-class ambiguity

Cirsium rivulare (Jacq.) All. stefano mazzotti
CC-BY-SA
Cirsium rivulare (Jacq.) All.
Chaerophyllum aromaticum L. buqa Jarmil
CC-BY-SA
Chaerophyllum aromaticum L.
Adenostyles leucophylla Rchb. Walter Reider
CC-BY-SA
Adenostyles leucophylla Rchb.
Petrosedum montanum (Songeon & E.P.Perrier) Grulich furs
CC-BY-SA
Petrosedum montanum (Songeon & E.P.Perrier) Grulich
Cirsium tuberosum (L.) All. Rene Weck
CC-BY-SA
Cirsium tuberosum (L.) All.
Chaerophyllum temulum L. Jcm Arthur
CC-BY-SA
Chaerophyllum temulum L.
Adenostyles alpina (L.) Bluff & Fingerh. pierre Lamy
CC-BY-SA
Adenostyles alpina (L.) Bluff & Fingerh.
Petrosedum rupestre (L.) P.V.Heath Wolfi 41
CC-BY-SA
Petrosedum rupestre (L.) P.V.Heath
Based on pictures only, plant species are challenging to discriminate!

Bias in Citizen Science Data





Tree with roots

Geographic bias (Pl@ntNet observations)


Food bias



Top-5 most observed plant species in Pl@ntNet (13/04/2024):


25134 obs.
Echium vulgare L. Llandrich anna
CC-BY-SA
Echium vulgare L.
24720 obs.
Ranunculus ficaria L. JYCO
CC-BY-SA
Ranunculus ficaria L.
24103 obs.
Prunus spinosa L. fuerst.ernst
CC-BY-SA
Prunus spinosa L.
23288 obs.
Zea mays L. Uta Groger
CC-BY-SA
Zea mays L.
23075 obs.
Alliaria petiolata (M.Bieb.) Cavara & Grande Pascal Ollagnier
CC-BY-SA
Alliaria petiolata (M.Bieb.) Cavara & Grande

Beauty bias


10753 obs.
Centaurea jacea L. Dieter Wagner
CC-BY-SA
Centaurea jacea L.
6 obs.
Cenchrus agrimonioides Trin. David Eickhoff − EOL
CC-BY-SA
Cenchrus agrimonioides Trin.

Size bias


8376 obs.
Magnolia grandiflora L. Patrick Cartier
CC-BY-SA
Magnolia grandiflora L.
413 obs.
Moehringia trinervia (L.) Clairv. Maximilien Perrin
CC-BY-SA
Moehringia trinervia (L.) Clairv.

More biases … shaping the dataset


Selection bias

What gets photographed is not random.

Easy access

Visible species

Easy identification

Temporal bias

What we see depends on when we look.

Season

Flowering

Plant appearance

Subjective bias

Observations can reflect observers own choices.

Personal preferences

Rare species

Endangered species


more observations ≠ a more representative dataset

The Pl@ntNet-300K dataset



Tree with roots

Datafication & data sharing



Popular labeled datasets limitations:

  • structure of labels too simplistic (CIFAR-10, CIFAR-100)
  • tasks too easy to discriminate / low labels confusion (MNIST)
  • too well-balanced, same number of images per class (Imagenet)
  • duplicate, low-quality, irrelevant images (Recht et al., 2019)


Pl@ntNet-300K: a curated, shareable subset of Pl@ntNet for frictionless reproducibility (Donoho, 2024)

  • Advance plant identification research with real-world biodiversity data
  • Engage the ML community with meaningful challenges beyond synthetic benchmarks

Construction of Pl@ntNet-300K




Constraints:

  • Extract images from South-Western Europe (SWE): most covered regions
  • Provide approx. 300 000 images
  • Shareable on Zenodo (size < 50Go)
  • Need more than 4 samples per class for train/val/test split
  • Curated images / labels (well, really?)

Krizhevsky (2009): “Furthermore, we personally verified every label submitted by the labelers”, on CIFAR10 & CIFAR100 paper

Pl@ntNet South Western Europe (SWE)

Species distribution, long tail & Lorenz curve


80% species | 12% images \(\iff\) 20% species | 88% images


  • Pl@ntNet (SWE): has a very long tail
  • CIFAR-100: balanced
  • ImageNet: (approx.) balanced
  • Pl@ntNet (SWE) - Rand: subsampling images uniformly from Pl@ntNet (SWE) distorts the distribution

Hierarchical construction of Pl@ntNet-300K

  • Sample at genus level (10%) from all genera present in Pl@ntNet-300k
  • Collect all images of such genera from Pl@ntNet (SWE)

Species distribution, long tail & Lorenz curve



  • Pl@ntNet (SWE): has a very long tail
  • Pl@ntNet (SWE) - Rand: subsampling images uniformly from Pl@ntNet (SWE) distorts the distribution
  • PlantNet-300K: long tail preserved by genera subsampling

Pl@ntNet-300K: Long tail visualization



Y-Scale:
Species Preview

Details on Pl@ntNet-300K



Goal: Provide a real life scenario for supervised learning, with multi-class and long-tail label distribution


Pl@ntNet-300K v1.0 characteristics:


  • 306,146 color images
  • Size: 32 GB
  • Labels: 1 000+ species
  • Required 2 000 000 volunteers

Neurips (Datasets and Benchmarks track) paper

(Garcin et al., 2021)


Zenodo, 1 click download

https://zenodo.org/record/5645731


Code to train models

https://github.com/plantnet/PlantNet-300K


2026 Update: Pl@ntNet-300K v2.0 on Zenodo

Uncertainty quantification with long-tail

Long tail plot

Joint work with


Tiffany Ding

UC Berkeley

within

Jean-Baptiste Fermanian

Inria



Conformal Prediction for Long-Tailed Classification

T. Ding, J.-B. Fermanian and J. Salmon

ICLR 2026

Pl@ntNet: set prediction (recommendation)



Elements to help guide the users

  • provide a set of possible species/labels
  • display similar images from proposed species
  • give a score of confidence

(Split) Conformal Prediction (Vovk et al., 2005)



Goal:

For an input image \(X\), propose the most probable classes \(y\) with confidence level \(1-\alpha\) (with small \(\alpha\))



Main idea: Return classes with predicted score above a threshold:

\[ \mathcal{C}_{\alpha}(X) = \big\{ y : s(X,y) \geq t_\alpha \big\} \]

In classification: a common conformal score is \(s(x,y) = \hat{p}(y|x)\) (e.g., softmax scores).



Key assumption: The data calibration \((X_i, Y_i)\) are exchangeable with the tested points.



Conformal prediction: sets \(t_{\alpha}\) as the \((1-\alpha)\) quantile of the scores on a calibration set

Optimal sets (Sadinle et al., 2019)


Marginal (or standard) coverage targets:

\[\mathbb{P}\big[ Y \in \mathcal{C}_{\alpha}(X) \big ] \geq 1 - \alpha.\]

Class conditional coverage targets:

\[\forall y,\quad \mathbb{P}\big[ Y \in \mathcal{C}_\alpha(X) | Y=y \big ] \geq 1 - \alpha.\]

The optimal set of minimum size and marginal coverage of at least \(1-\alpha\) is: \[ \mathcal{C}_{\alpha}(x) = \left\{ y : p(y|x) \geq t_\alpha \right\} \]

The optimal set of minimum size and conditional coverage of at least \(1-\alpha\) is: \[ \mathcal{C}_{\alpha}(x) = \left\{ y : p(y|x) \geq t_\alpha^{y} \right\} \]


Marginal:

calibrate \(t_{\alpha}\) on whole calibration set \((X_i, Y_i)_{i=1}^n\)

Conditional:

calibrate \(t_{\alpha}^y\) only on \((X_i, Y_i)\) such that \(Y_i = y\)

Interactive optimal set visualization


Class conditional: useless for long-tail




FracBelow50%: proportion of classes with low coverage (below 50%)

Targeting Macro-Coverage


Goal: Average coverage across all classes (better for long-tail!)

\[ \text{MacroCoverage} = \frac{1}{|\mathcal{Y}|} \sum_{y \in \mathcal{Y}} \mathbb{P}\big( Y \in \mathcal{C}(X) \, | \, Y = y \big) \]

The optimal set of minimum size and Macro-Coverage of at least \(1-\alpha\) is: \[ \mathcal{C}_{\alpha}(x) = \left\{ y : \frac{p(y|x)}{p(y)} \geq t_\alpha \right\} \]


Introduce a new conformal score, Prevalence Adjusted Softmax (PAS) : \(s(x,y) = \frac{\hat{p}(y|x)}{\hat{p}(y)}\)

Interactive optimal set visualization (II)


Experiments on Pl@ntNet-300K



Generalization : weighted Macro-Coverage


For user-chosen class weights \(\omega\) with \(\omega(y)\geq 0\) and \(\sum_{y \in \mathcal{Y}} \omega(y)=1\), define the \(\omega\)-weighted macro-coverage :

\[ \begin{align} \mathrm{MacroCov}_{\omega}(\mathcal{C}) = \sum_{y \in \mathcal{Y}} \omega(y) \mathbb{P}(Y \in \mathcal{C}(X) \mid Y = y). \end{align} \]

The optimal set of minimum size and Macro-Coverage of at least \(1-\alpha\) is: \[ \begin{align} \mathcal{C}^*(x) = \left\{ y \in \mathcal{Y} : \omega(y) \dfrac{p(y|x)}{p(y)} \geq t\right\}, \end{align} \]


Introduce a new conformal score, Weighted Prevalence Adjusted Softmax (WPAS) : \(s(x,y) = \omega(y) \frac{\hat{p}(y|x)}{\hat{p}(y)}\)

Experiments on Pl@ntNet-300K: endangered species


Goal: handle at-risk species in Pl@ntNet-300K, weight by \(\gamma \geq 1\) times more at-risk species:

\[ \omega(y) = \begin{cases} \frac{\gamma}{W} & \text{if } y \in \mathcal{Y}_{\text{at-risk}} \quad (\text{with } W = \gamma|\mathcal{Y}_{\text{at-risk}}| + |\mathcal{Y} \setminus \mathcal{Y}_{\text{at-risk}}|)\\ \frac{1}{W} & \text{otherwise}, \end{cases} \]

Take home message



  • Challenges in citizen science: many, varied and needing more attention
  • Prediction: theory can guide which prediction set to display

Dataset release:

Code release:

Future work

  • Handling label hierarchy
  • Human–computer interaction / performative learning / model collapse
  • Improve robustness to adversarial users
  • Improve label agregation, leverage gamification (e.g., theplantgame.com), etc.

References

References


Bar-On, Y. M., Phillips, R., & Milo, R. (2018). The biomass distribution on earth. Proceedings of the National Academy of Sciences, 115(25), 6506–6511.
Ding, T., Fermanian, J.-B., & Salmon, J. (2026). Conformal Prediction for Long-Tailed Classification. ICLR.
Donoho, D. (2024). Data science at the singularity. Harvard Data Science Review, 6(1).
Garcin, C., Joly, A., Bonnet, P., Affouard, A., Lombardo, J.-C., Chouet, M., Servajean, M., Lorieul, T., & Salmon, J. (2021). Pl@ntNet-300K: A plant image dataset with high label ambiguity and a long-tailed distribution. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Krizhevsky, A. (2009). Learning multiple layers of features from tiny images. Univ. of Toronto.
Lefort, T., Affouard, A., Charlier, B., Lombardo, J.-C., Chouet, M., Goëau, H., Salmon, J., Bonnet, P., & Joly, A. (2025). Cooperative learning of pl@ntNet’s artificial intelligence algorithm: How does it work and how can we improve it? Methods in Ecology and Evolution.
Lefort, T., Charlier, B., Joly, A., & Salmon, J. (2024). Identify ambiguous tasks combining crowdsourced labels by weighting areas under the margin. TMLR.
Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do imagenet classifiers generalize to imagenet? ICML, 5389–5400.
Sadinle, M., Lei, J., & Wasserman, L. (2019). Least ambiguous set-valued classifiers with bounded error levels. Journal of the American Statistical Association, 114(525), 223–234.
Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic learning in a random world. Springer.