Keywords:
Zero-shot Segmentation Weak Validation Species Imaging Biodiversity Vision
Introduction:
Large species classification datasets have enabled the training of models that can identify the main species in an image from a large pool of species spanning different taxonomic groups, including rare species [1, 2, 6, 7].
Currently, no dataset allows this to be done for fine-grained computer vision tasks. Weakly-supervised approaches that rely on text labels, such as those described in references [8, 9 and 10], are a promising way to leverage existing datasets in order to tackle these fine-grained tasks. However, validating such methods still requires datasets with denser annotations than simple image-wise labels.

Objectives:
Your task in this project is to design a weak-validation framework for zero-shot semantic segmentation models that can be used to evaluate their performance in biodiversity/ecology imaging settings. Specifically, it should assess their ability to accurately segment a large number of living species.
-
The first step is to find relevant datasets that either cover:
-
fine-grained tasks over species images, such as segmentation or detection. Examples: [3, 5]
-
or coarse multi-label or single-label classification tasks over a large number of species. Examples: [4, 1, 2].
-
-
The second step is to provide the code necessary to use these datasets (i.e. download procedures, dataset classes, etc.).
-
The third step involves coding the transformations that convert the semantic maps produced by the evaluated models into bounding boxes, multi-labels, or single labels. This allows metrics to be computed by comparing them with the ground-truth labels of the validation datasets.
-
The fourth step involves inferring existing pre-trained state-of-the-art models using the crafted framework (e.g. [9, 10]), or training/fine-tuning one model and subsequently evaluating it. This step is flexible and can be adapted based on the number of credits of your project, your interests and level of expertise.
Deliverables:
-
A clean codebase of your evalutation framework
-
The performance evaluation of at least one model on the different datasets you selected
-
A concise report including information regarding the datasets you selected, your design choices, implementation, and evaluation results
What you take away:
-
The chance to contribute to a publication about a weakly-supervised semantic segmentation model based on [8] for images of living species, which will be evaluated and compared to other approaches using your framework.
-
Gaining knowledge about state-of-the-art weakly-supervised semantic segmentation models and computer vision of biodiversity.
-
Hands-on experience of working with deep learning datasets and pre-trained models on HPC infrastructure.
Prerequisites:
-
You are comfortable programming in Python, including using the PyTorch library.
-
You have a strong understanding of deep learning, as evidenced by your track record.
-
You are a highly committed individual with a genuine interest in the technical challenges and potential impact of the topic.
Contact:
Interested candidates are kindly asked to send their CV and a short motivation statement by email.
References:
[1]G. Van Horn et al., “The iNaturalist Species Classification and Detection Dataset,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT: IEEE, Jun. 2018, pp. 8769–8778. doi: 10.1109/CVPR.2018.00914.
[2]C. Garcin et al., “Pl@ntNet-300K image dataset.” Zenodo, Apr. 29, 2021. doi: 10.5281/ZENODO.5645731.
[3]J. Weyler et al., “PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain,” IEEE Trans. on Pattern Analysis and Machine Intelligence (T-PAMI), vol. 46, no. 12, pp. 9583–9594, 2024.
[4]L. Picek et al., “Overview of LifeCLEF 2025: Challenges on Species Presence Prediction and Identification, and Individual Animal Identification,” in Experimental IR Meets Multilinguality, Multimodality, and Interaction, vol. 16089, J. Carrillo-de-Albornoz, A. García Seco De Herrera, J. Gonzalo, L. Plaza, J. Mothe, F. Piroi, P. Rosso, D. Spina, G. Faggioli, and N. Ferro, Eds., Cham: Springer Nature Switzerland, 2026, pp. 338–362. doi: 10.1007/978-3-032-04354-2_19.
[5]J. Gallmann et al., “Flower Mapping in Grasslands With Drones and Deep Learning,” Front. Plant Sci., vol. 12, p. 774965, Feb. 2022, doi: 10.3389/fpls.2021.774965.
[6]S. Stevens et al., “BioCLIP: A Vision Foundation Model for the Tree of Life,” 2024, pp. 19412–19424. Accessed: Mar. 09, 2026. [Online]. Available: https://openaccess.thecvf.com/content/CVPR2024/html/Stevens_BioCLIP_A_Vision_Foundation_Model_for_the_Tree_of_Life_CVPR_2024_paper.html
[7]J. Gu et al., “BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning,” May 29, 2025, arXiv: arXiv:2505.23883. doi: 10.48550/arXiv.2505.23883.
[8]G. Braso, A. Osep, and L. Leal-Taixé, “Native Segmentation Vision Transformers,” Oct. 2025. Accessed: May 18, 2026. [Online]. Available: https://openreview.net/forum?id=V7RRnsAlbY
[9]M. Yi, Q. Cui, H. Wu, C. Yang, O. Yoshie, and H. Lu, “A Simple Framework for Text-Supervised Semantic Segmentation,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 7071–7080. doi: 10.1109/CVPR52729.2023.00683.
[10]J.-J. Wu et al., “Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation,” 2024, arXiv. doi: 10.48550/ARXIV.2404.04231.