<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \bartext{Research article}?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">BG</journal-id><journal-title-group>
    <journal-title>Biogeosciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">BG</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Biogeosciences</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1726-4189</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/bg-21-2909-2024</article-id><title-group><article-title>From simple labels to semantic image segmentation: leveraging citizen science plant photographs for tree species mapping<?xmltex \hack{\break}?> in drone imagery</article-title><alt-title>From simple labels to semantic image segmentation</alt-title>
      </title-group><?xmltex \runningtitle{From simple labels to semantic image segmentation}?><?xmltex \runningauthor{S. Soltani et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff3 aff4">
          <name><surname>Soltani</surname><given-names>Salim</given-names></name>
          <email>salim.soltani@uni-leipzig.de</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4 aff6">
          <name><surname>Ferlian</surname><given-names>Olga</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-2536-7592</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4 aff6">
          <name><surname>Eisenhauer</surname><given-names>Nico</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2 aff3 aff4 aff5">
          <name><surname>Feilhauer</surname><given-names>Hannes</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff4">
          <name><surname>Kattenborn</surname><given-names>Teja</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-7381-3828</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Sensor-based Geoinformatics (geosense), University of Freiburg, Freiburg, Germany</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Remote Sensing Centre for Earth System Research (RSC4Earth), Leipzig University, Leipzig, Germany</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI), Leipzig University, Leipzig, Germany</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>German Centre for Integrative Biodiversity Research (iDiv) Halle–Jena–Leipzig, Leipzig, Germany</institution>
        </aff>
        <aff id="aff5"><label>5</label><institution>Remote Sensing, Helmholtz Centre for Environmental Research, Leipzig, Germany</institution>
        </aff>
        <aff id="aff6"><label>6</label><institution>Institute of Biology, Leipzig University, Leipzig, Germany</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Salim Soltani (salim.soltani@uni-leipzig.de)</corresp></author-notes><pub-date><day>14</day><month>June</month><year>2024</year></pub-date>
      
      <volume>21</volume>
      <issue>11</issue>
      <fpage>2909</fpage><lpage>2935</lpage>
      <history>
        <date date-type="received"><day>2</day><month>November</month><year>2023</year></date>
           <date date-type="rev-request"><day>5</day><month>December</month><year>2023</year></date>
           <date date-type="rev-recd"><day>11</day><month>April</month><year>2024</year></date>
           <date date-type="accepted"><day>22</day><month>April</month><year>2024</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2024 Salim Soltani et al.</copyright-statement>
        <copyright-year>2024</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://bg.copernicus.org/articles/bg-21-2909-2024.html">This article is available from https://bg.copernicus.org/articles/bg-21-2909-2024.html</self-uri><self-uri xlink:href="https://bg.copernicus.org/articles/bg-21-2909-2024.pdf">The full text article is available as a PDF file from https://bg.copernicus.org/articles/bg-21-2909-2024.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e154">Knowledge of plant species distributions is essential for various application fields, such as nature conservation, agriculture, and forestry. Remote sensing data, especially high-resolution orthoimages from unoccupied aerial vehicles (UAVs), paired with novel pattern-recognition methods,  such as convolutional neural networks (CNNs), enable accurate mapping (segmentation) of plant species. Training transferable pattern-recognition models for species segmentation across diverse landscapes and data characteristics typically requires extensive training data. Training data are usually derived from labor-intensive field surveys or visual interpretation of remote sensing images. Alternatively, pattern-recognition models could be trained more efficiently with plant photos and labels from citizen science platforms, which include millions of crowd-sourced smartphone photos and the corresponding species labels. However, these pairs of citizen-science-based photographs and simple species labels (one label for the entire image) cannot be used directly for training state-of-the-art segmentation models used for UAV image analysis, which require per-pixel labels for training (also called masks). Here, we overcome the limitation of simple labels of citizen science plant observations with a two-step approach. In the first step, we train CNN-based image classification models using the simple labels and apply them in a moving-window approach over UAV orthoimagery to create segmentation masks. In the second phase, these segmentation masks are used to train state-of-the-art CNN-based image segmentation models with an encoder–decoder structure. We tested the approach on UAV orthoimages acquired in summer and autumn at a test site comprising 10 temperate deciduous tree species in varying mixtures. Several tree species could be mapped with surprising accuracy (mean F1 score <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.47</mml:mn></mml:mrow></mml:math></inline-formula>). In homogenous species assemblages, the accuracy increased considerably (mean F1 score <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.55</mml:mn></mml:mrow></mml:math></inline-formula>). The results indicate that several tree species can be mapped without generating new training data and by only using preexisting knowledge from citizen science. Moreover, our analysis revealed that the variability in citizen science photographs, with respect to acquisition data and context, facilitates the generation of models that are transferable through the vegetation season. Thus, citizen science data may greatly advance our capacity to monitor hundreds of plant species and, thus, Earth's biodiversity across space and time.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Deutsche Forschungsgemeinschaft</funding-source>
<award-id>44452490</award-id>
<award-id>04978936</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e186">Spatially explicit information on plant species is crucial for various domains and applications, including nature conservation, agriculture, and forestry. For instance, species information is required for the identification of threatened or invasive<?pagebreak page2910?> species, the location of weeds or crops in precision farming, or tree species classification for forest inventories.</p>
      <p id="d1e189">Remote sensing has emerged as a promising tool for mapping plant species <xref ref-type="bibr" rid="bib1.bibx38 bib1.bibx4 bib1.bibx14" id="paren.1"/>. Thereby, supervised machine learning algorithms are commonly used to identify species-specific features in spatial, temporal, or spectral patterns of remotely sensed signals <xref ref-type="bibr" rid="bib1.bibx46 bib1.bibx35 bib1.bibx32 bib1.bibx11 bib1.bibx51" id="paren.2"/>. In recent years, remote sensing imagery from drones, also known as unoccupied air vehicles (UAVs), has emerged as an effective source of information for mapping plant species <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx14 bib1.bibx41" id="paren.3"/>. By means of mosaicking a series of individual image frames, UAVs enable the creation of georeferenced orthoimagery of relatively large areas with extremely high spatial resolution, e.g., in the millimeter or centimeter range. The fine spatial grain of such imagery can reveal distinctive morphological plant features to identify specific plant species. Such plant features include the leaf shape, flowers, branching patterns, or crown structures <xref ref-type="bibr" rid="bib1.bibx46 bib1.bibx26" id="paren.4"/>. An effective way to harness this spatial detail is provided by deep-learning-based pattern-recognition techniques, in particular by convolutional neural networks (CNNs). A series of studies have demonstrated that CNNs allow one to precisely segment plant species' canopies in high-resolution UAV imagery <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx20 bib1.bibx7" id="paren.5"/>. Such CNN models learn the characteristic spatial features of the target (here, plant species) through a cascade of filter operations (convolutions). Given these high-dimensional computations, efficiently applying these models to UAV orthoimagery, which often have large spatial extents and high resolution, requires training and applying them sequentially using smaller subregions of an orthoimage (e.g., image tiles of 512 pixels <inline-formula><mml:math id="M3" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixels; Fig. <xref ref-type="fig" rid="Ch1.F1"/>c).</p>
      <p id="d1e217">However, generating models that are transferable across various landscapes and remote sensing data characteristics requires a large number of training data <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx18" id="paren.6"/>. In particular, when neighboring plant species bear a resemblance, a wealth of training data becomes essential, allowing the model to discern the subtle distinctions between these species <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx41" id="paren.7"/>. Commonly, the generation of training data is costly, as training data are usually derived from field surveys or visual interpretation of remote sensing images, also known as annotation or labeling. Both methods have limitations. For example, field surveys are often logistically challenged by site accessibility or travel costs. Moreover, they commonly only enable the acquisition of point observations or relative cover fractions of the target species <xref ref-type="bibr" rid="bib1.bibx30" id="paren.8"/>. Visual image interpretation is often much more effective <xref ref-type="bibr" rid="bib1.bibx27 bib1.bibx42" id="paren.9"/>; however, for some species, precise visual identification of species can be challenging due to subtle indicative morphological features, the variability in these features in the landscape, or the complexity of vegetation communities (e.g., smooth transitions of canopies of different species). Moreover, the representativeness of data derived from field surveys and visual interpretation is often limited to the location and time at which the data were acquired. This can reduce a model's generalizability to new regions or time periods <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx29" id="paren.10"/>. Therefore, the number and quality of training data obtained can be critical with respect to the performance and transferability of CNN models <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx40 bib1.bibx6" id="paren.11"/>.</p>
      <p id="d1e239">The challenge of limited training data for UAV-based plant species identification may be alleviated by the collective power of scientists and citizens openly sharing their plant observations on the web <xref ref-type="bibr" rid="bib1.bibx22 bib1.bibx16 bib1.bibx13" id="paren.12"/>. A particular data treasure in this regard is generated by citizen science projects for plant species identification. Examples are the iNaturalist and Pl@ntNet projects, which encourage tens of thousands of individuals to capture, share, and annotate photographs of global plant life <xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx13" id="paren.13"/>. The quantity of such citizen science observations is rapidly growing due to the increasing number of volunteers participating in such projects <xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx13" id="paren.14"/>.</p>
      <p id="d1e252">Currently, the iNaturalist project contains over 26 million globally distributed and annotated photographs of vascular plant species. The iNaturalist platform allows users to identify plant species manually or using a computer vision model integrated into the platform. The submitted observations are then evaluated by the community, and a research-grade classification is assigned if over two-thirds of the community agrees on the species identification. The Pl@ntNet project includes over 20 million observations of globally distributed vascular plants. Pl@ntNet requires users to photograph their observations and select an organ tag (e.g., leaf, flower, fruit, or stem). Pl@ntNet features an image-recognition algorithm to analyze the tagged photograph and suggest a plant species. Pl@ntNet's validation process uses a dynamic approach, combining automated algorithm confidence with community consensus <xref ref-type="bibr" rid="bib1.bibx24" id="paren.15"/>. The validated observations of iNaturalist and Pl@ntNet are shared via the Global Biodiversity Information Facility (GBIF), a global network that provides open access to biodiversity data <xref ref-type="bibr" rid="bib1.bibx19" id="paren.16"/>.</p>
      <p id="d1e261">Citizen-science-based plant photographs with species annotations provide a valuable, large, and continuously growing data source for training pattern-recognition models, such as CNNs <xref ref-type="bibr" rid="bib1.bibx49 bib1.bibx24" id="paren.17"/>. However, such citizen science data have a cardinal limitation: they only provide a simple species annotation for a plant photograph (the image <inline-formula><mml:math id="M4" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> shows species <inline-formula><mml:math id="M5" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>). Hence, these labels enable one to train image classification models that predict the likelihood of a species being present in an image, but they do not specify where the species is present in the<?pagebreak page2911?> image. Ideally, for species-mapping applications, the species labels would delineate the regions or pixels belonging to a species (e.g., the pixels in the right corner of image <inline-formula><mml:math id="M6" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> represent species <inline-formula><mml:math id="M7" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>). Such labels (known as masks) could be used to train CNN-based segmentation models, which can predict a species probability for each individual pixel of an image (or tile of an orthoimage) <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx41" id="paren.18"/>.</p>
      <p id="d1e299">In a pioneering study by <xref ref-type="bibr" rid="bib1.bibx45" id="text.19"/>, the limitation of the simple labels that come with citizen science photographs was overcome by a workaround. At first, image classification models were trained with citizen science data and simple labels to predict a species per image. The trained image classification models were then applied sequentially on tiles of UAV-based orthomosaics in a moving-window-like fashion with very high overlap (Fig. <xref ref-type="fig" rid="Ch1.F1"/>a). Lastly, the individual predictions derived from the moving-window steps were rasterized to a seamless segmentation map (Fig. <xref ref-type="fig" rid="Ch1.F1"/>b). However, this workaround is computationally intense and inefficient for large or multiple UAV orthomosaics, as segmentation maps can only be derived from many overlapping prediction steps. In contrast, the state-of-the-art CNN-based segmentation methods (typically an encoder–decoder structure) used in remote sensing applications are trained with reference data in the form of masks with dimensions (pixels) corresponding to the extent of the input imagery, where each pixel of the mask defines the absence or presence of a class (here, plant species) in the imagery <xref ref-type="bibr" rid="bib1.bibx28" id="paren.20"/>. Respective segmentation models are more efficient, as they segment multiple classes in a single prediction step. Moreover, they enable more detailed class representations in situations where multiple classes are arranged in complex patterns.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e314">Schematic representation of the proposed workflow, including <bold>(a, b)</bold> the moving-window approach by <xref ref-type="bibr" rid="bib1.bibx45" id="text.21"/> and <bold>(c)</bold> the use of state-of-the-art encoder–decoder segmentation algorithms. The photographs in panel <bold>(a)</bold> are sourced from <xref ref-type="bibr" rid="bib1.bibx21" id="text.22"/>.</p></caption>
        <?xmltex \igopts{width=227.622047pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f01.jpg"/>

      </fig>

      <p id="d1e338">Here, we propose a solution to overcome the limitation of simple annotations of citizen science plant observations with a two-step approach. In the first step, we apply the procedure of <xref ref-type="bibr" rid="bib1.bibx45" id="text.23"/>, involving CNN-based image classification models trained on citizen science photographs and simple species labels to predict plant species in UAV orthoimages using the moving-window approach described above (Fig. <xref ref-type="fig" rid="Ch1.F1"/>a, b). Although computationally demanding, this serves to create segmentation masks for UAV orthoimages. In the second step, these segmentation masks are used to train more efficient CNN-based image segmentation models with an encoder–decoder structure (Fig. <xref ref-type="fig" rid="Ch1.F1"/>c). These more efficient models could then be applied to larger spatial extents or to new UAV orthomosaics (e.g., of different sites or time steps).</p>
      <p id="d1e349">Hence, the present study addresses the following research questions: <list list-type="bullet"><list-item>
      <p id="d1e354">Can we harness weak labels from citizen science plant observations to train efficient state-of-the-art semantic segmentation models?</p></list-item><list-item>
      <p id="d1e358">Do those segmentation models also increase the accuracy compared with the simple moving-window approach?</p></list-item></list> These questions are evaluated on a tree species dataset acquired at an experimental site (MyDiv experiment, Bad Lauchstädt, Germany), where 10 temperate deciduous tree species were planted in stratified and complex mixtures. The selection of this location is attributed to its harmonious coexistence of various plant species within a compact area.</p>
</sec>
<?pagebreak page2912?><sec id="Ch1.S2">
  <label>2</label><title>Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Data acquisition and preprocessing </title>
<sec id="Ch1.S2.SS1.SSS1">
  <label>2.1.1</label><title>Study site and drone data acquisition</title>
      <p id="d1e384">The MyDiv experimental site is located in Bad Lauchstädt, Saxony-Anhalt, Germany (51°23<inline-formula><mml:math id="M8" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> N, 11°53<inline-formula><mml:math id="M9" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> E). It comprises 80 plots with different configurations of 10 deciduous tree species, including <italic>Acer pseudoplatanus</italic>, <italic>Aesculus hippocastanum</italic>, <italic>Betula pendula</italic>, <italic>Carpinus betulus</italic>, <italic>Fagus sylvatica</italic>, <italic>Fraxinus excelsior</italic>, <italic>Prunus avium</italic>, <italic>Quercus petraea</italic>, <italic>Sorbus aucuparia</italic>, and <italic>Tilia platyphyllos</italic> <xref ref-type="bibr" rid="bib1.bibx15" id="paren.24"/>. Each plot measures 12 m <inline-formula><mml:math id="M10" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 12 m and contains 140 trees planted at a distance of 1 m from one another (Fig <xref ref-type="fig" rid="Ch1.F2"/>). In total, all plots combined accommodate a total of 11 200 individual trees. Each plot contains varying tree species compositions, including one, two, and four tree species. This species variety, the balanced species composition, and plots of different canopy complexity (species mixtures) provide an ideal setting to test the proposed species segmentation approach.</p>
      <p id="d1e449">We collected UAV-based RGB aerial imagery over the MyDiv experimental site using a DJI Mavic 2 Pro and the DroneDeploy (version 5.0, USA) flight planning software. Two flights were conducted in 2022 in July and September; July corresponds to the peak of the growing season, whereas September corresponds to the senescence stage (Fig <xref ref-type="fig" rid="Ch1.F2"/>). The flight plan was set up with a forward overlap of 90 % and a side overlap of 70 % at an altitude of 16 m (ground sampling distance of approximately 0.22 cm per pixel). We used the generated images and Metashape (version 1.7.6, Agisoft LLC) to create orthoimages for both flight campaigns. Hereafter, the orthoimages for July and September are referred to as Ortho<inline-formula><mml:math id="M11" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  and Ortho<inline-formula><mml:math id="M12" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>, respectively.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e474">Overview of the MyDiv experimental site with close-ups of three plots with different species compositions. The MyDiv site is located at 51.3916° N, 11.8857° E.</p></caption>
            <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f02.jpg"/>

          </fig>

      <p id="d1e484">To evaluate the performance of the CNN models for tree species mapping, we created reference data by manually delineating the tree species in the UAV orthoimages in QGIS (version 3.32.3). To reduce the workload, we did not delineate the species for the entire plot; rather, they were specified for diagonal transects with a 20 m length and a 2 m width.</p>
</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <label>2.1.2</label><title>Citizen science training data</title>
      <p id="d1e495">We queried citizen science plant observations of the iNaturalist and Pl@ntNet datasets via the GBIF database for our target tree species using scientific names. For the iNaturalist data, we used the R package rinat (version 0.1.8), an application programming interface (API) for iNaturalist. The Pl@ntNet data for the selected tree species were acquired using the tabulated observation data from GBIF and the integrated uniform resource locators (URLs) for the images. The number of photographs available from iNaturalist and Pl@ntNet varied for the different tree species. Per species, we were able to acquire between 582 and 10 000 photographs (mean 7696) from the iNaturalist dataset and between 221 and 3304 images (mean 2238) from the Pl@ntNet dataset (see Table <xref ref-type="table" rid="App1.Ch1.S1.T1"/> in the Appendix for details ).</p>
      <p id="d1e500">In addition to the tree species, we added a background class to consider canopy gaps between trees. Training data for this background class were obtained using the Google Images API and queries of different keywords, e.g., “grass”, “forest floor”, and “forest ground”. After cleaning the obtained images for nonmeaningful results, the background class included 1100 photographs.</p>
      <p id="d1e503">We converted all photographs to a rectangular shape, by cropping them to the shorter side, and resampled them to a common size of 512 pixels <inline-formula><mml:math id="M13" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixels (the tile size used later for the CNN model generation). Figure <xref ref-type="fig" rid="Ch1.F3"/> shows examples of the downloaded photographs for the different tree species and a comparison with their appearance in Ortho<inline-formula><mml:math id="M14" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e527">Example citizen-science-based photographs derived from iNaturalist and tiles of UAV orthoimages (512 pixels <inline-formula><mml:math id="M15" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixels) for the 10 tree species in the MyDiv experiment. Photographs in the left column are sourced from <xref ref-type="bibr" rid="bib1.bibx21" id="text.25"/>.</p></caption>
            <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f03.jpg"/>

          </fig>

      <p id="d1e546">The acquisition settings of citizen science plant photographs are heterogeneous and differ considerably from the typical bird's-eye perspective of UAV orthoimages (Fig. <xref ref-type="fig" rid="Ch1.F3"/>). For instance, from the UAV perspective, canopies are mostly viewed from a relatively homogeneous distance, and the photographs represent mostly leaves and other crown components. In contrast, the citizen science data include a lot of close-ups, landscape imagery, or horizontal photographs of trunks. <xref ref-type="bibr" rid="bib1.bibx45" id="text.26"/> demonstrated that species recognition in UAV images can be improved by excluding crowd-sourced photographs that are exceptionally close (e.g., showing individual leaf veins) or too far away from the plant (e.g., landscape images). Therefore, we filtered the citizen-science-based training photos according to the camera–plant distance. Moreover, we filtered photos that exclusively contained tree stems. Because such information is unavailable in the citizen science datasets, we trained CNN-based regression and classification models to predict acquisition distance and tree trunk presence for each downloaded photograph. To train these CNN-based models, we visually estimated the acquisition distance (4500 photographs) and labeled tree trunk presence (1000 photographs). To ease the labeling process, we used previously labeled training data from <xref ref-type="bibr" rid="bib1.bibx45" id="text.27"/> and added 150 additional tree photographs from the tree species present at the MyDiv experimental site.</p>
      <p id="d1e557">To evaluate the models with respect to predicting the acquisition distance and trunk presence, we randomly split the citizen-science-based plant photographs into training and validation sets, with 80 % for training and 20 % for validation.</p>
      <p id="d1e560">For the distance regression and the trunk classification, we used the EfficientNetB7 backbone <xref ref-type="bibr" rid="bib1.bibx47" id="paren.28"/>. For the distance regression, we used the following top-layer settings: global average pooling, batch normalization, drop out (rate 0.1), and a final dense layer with one unit and linear activation function. We used the Adam optimizer (learning rate of 0.0001) and a mean-squared error (MSE) loss<?pagebreak page2913?> function. For the trunk classification, we used the following top-layer settings: global max pooling, a final dense layer with two units, and a softmax activation function. We used the Adam optimizer (learning rate of 0.0001) and the categorical cross-entropy loss function. Both models were trained using a batch size of 20 and 50 epochs.</p>
      <p id="d1e566">We used the model with the lowest loss from these epochs (details on the model performance are given in Appendix <xref ref-type="sec" rid="App1.Ch1.S1.SS3"/>) to predict the acquisition distance and tree trunk presence in all downloaded photographs for our target species. We filtered training photographs prior to training CNN-based species classification (see Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>) with acquisition distances less than 0.2 m and greater than 15 m as well as photographs classified as trunk (probability threshold of 0.5). Following this process, 82 628 of the 101 574 downloaded citizen science photographs remained.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>CNN-based creation of plant species segmentation masks using a moving-window approach </title>
      <p id="d1e582">The segmentation masks were obtained using a CNN image classification model trained on crowd-sourced plant photographs and simple species labels using a moving-window method (hereafter CNN<inline-formula><mml:math id="M16" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula>; Fig. <xref ref-type="fig" rid="Ch1.F1"/>b). Based on the results of previous studies, we chose a generic image size of 512 pixels <inline-formula><mml:math id="M17" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixels for the CNN classification model <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx45" id="paren.29"/>. Using the moving-window approach, the orthoimage is sequentially cropped into tiles of 512 pixels <inline-formula><mml:math id="M18" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixels; the image classification is then applied to these tiles to predict the species for each location. This procedure is applied with a dense overlap between tiles defined by a step size, resulting in a dense, regular grid of species predictions. We chose a vertical and horizontal distance of 51 pixels as the step size. The resulting predictions were then rasterized to a continuous species distribution grid with a spatial resolution of 8.31 cm per pixel <xref ref-type="bibr" rid="bib1.bibx45" id="paren.30"><named-content content-type="pre">see</named-content><named-content content-type="post">for details</named-content></xref>. The CNN<inline-formula><mml:math id="M19" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> model was implemented as a classification task with 11 classes, including the 10 tree species and the background class.</p>
      <p id="d1e630">The number of available photographs varied widely across tree species (see Sect. <xref ref-type="sec" rid="Ch1.S2.SS1.SSS2"/>), potentially biasing the model towards classes with more photographs. To address this imbalance, we equally sampled 4000 photographs for each class with replacement. Sampling with replacement randomly duplicates the existing photographs for underrepresented classes – in this case, classes with fewer than 4000 photographs. We applied data augmentation to increase the variance of the duplicated images. The augmentation consisted of random vertical and horizontal flips, random brightness with a maximum delta of 10 % (<inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>), and contrast alteration within a range of 90 % to 110 % (0.9 to 1.1) of training photographs. We randomly partitioned the training data into validation and training sets to ensure unbiased evaluation. From the training set, we allocated a holdout of 20 % for model selection, while the remaining 80 % was used for model training. Subsequently, we assessed the accuracy of the selected model using the validation set.</p>
      <p id="d1e645">After testing different architectures as model backbones, including ResNet50V2, EfficientNetB07, and EfficientNetV2L, we selected EfficientNetV2L because it resulted in the highest classification accuracies. The following layers were added on top of the EfficientNetV2L backbone: dropout with a ratio of 0.5, average pooling, dropout with a ratio of 0.5, a dense layer with 128 units, L2 kernel regularizer (0.001), a rectified linear unit (ReLU) activation function, and a final dense layer with a softmax activation function and 11 units (corresponding to the 10 tree species<?pagebreak page2915?> and the background class). We used root-mean-square propagation (RMSprop) as the optimizer with a learning rate of 0.0001 and categorical cross-entropy as a loss function. We trained the configured model with a batch size of 15 over 150 epochs. The model with the lowest loss (based on the 20 % holdout) was selected as the final model. This model was used to predict the tree species (probabilities) in the UAV orthoimages using the abovementioned CNN<inline-formula><mml:math id="M21" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> method (Fig. <xref ref-type="fig" rid="Ch1.F1"/>b). To filter uncertain predictions (predominantly in canopy gaps or at crown shadows), we only considered a tree species as predicted above a threshold higher than 0.6; otherwise, it was assigned to NA (not available), which accounted for approximately 7.8 % of the UAV orthoimages. To smooth the predictions and remove noise, we applied a sieve operation on the output of the CNN<inline-formula><mml:math id="M22" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> (threshold <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">50</mml:mn></mml:mrow></mml:math></inline-formula>, considering horizontal, vertical, and diagonal neighbors, using the R package terra, version 1.7).</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>CNN-based plant species segmentation using an encoder–decoder architecture</title>
      <p id="d1e686">As the encoder–decoder segmentation architecture (hereafter CNN<inline-formula><mml:math id="M24" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>), we chose U-Net <xref ref-type="bibr" rid="bib1.bibx39" id="paren.31"/>, which is the most widely applied segmentation method in remote sensing image segmentation <xref ref-type="bibr" rid="bib1.bibx28" id="paren.32"/>. The U-Net architecture is a CNN-based algorithm that performs semantic segmentation by predicting a class for each pixel of the input image. The architecture consists of an encoder–decoder structure with skip connections. The configured architecture has four levels of convolutional blocks. Each convolutional block consists of two convolutional layers and is followed by batch normalization and ReLU activation. The encoder gradually compresses feature maps and reduces their spatial dimensions via max pooling operations, while the decoder increases the feature map resolution by transposed convolution. The encoder and decoder blocks are connected through skip connections, which transfer the spatial context of the encoder feature maps to the decoder, enabling a segmentation at the resolution of the input imagery in the last layer. The final layer has 11 units (corresponding to the 10 tree species and a background class). A corresponding softmax activation function maps the features to class probabilities. Using a max function, the pixels of the segmentation output are assigned to the class with the highest probability (Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F18"/>).</p>
      <p id="d1e706">The segmentation masks for training CNN<inline-formula><mml:math id="M25" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> were obtained from the predictions of the  CNN<inline-formula><mml:math id="M26" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> method applied to both UAV orthoimages (Ortho<inline-formula><mml:math id="M27" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula> and Ortho<inline-formula><mml:math id="M28" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>; Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>). At first, we resampled the CNN<inline-formula><mml:math id="M29" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> prediction maps to the original spatial resolution of the orthoimages (0.22 cm pixel size). Afterward, we cropped the orthoimages and the prediction maps into nonoverlapping tiles, each with a size of 512 pixels <inline-formula><mml:math id="M30" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixels, resulting in a total of 44 980 and 37 113 tiles from Ortho<inline-formula><mml:math id="M31" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula> and Ortho<inline-formula><mml:math id="M32" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>, respectively.</p>
      <p id="d1e782">The training data obtained from the CNN<inline-formula><mml:math id="M33" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> approach were filtered to avoid training the CNN<inline-formula><mml:math id="M34" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> model with uncertain predictions. Thereby, we assumed that predictions for a tile are uncertain when the model predicts multiple classes with low relative cover. Thus, after initial tests, we included only those tiles in which the cover of at least one class exceeded 30 %. The number of training tiles per class after filtering varied between 1257 and 16 894 samples: <italic>Acer pseudoplatanus</italic> (6581), <italic>Aesculus hippocastanum</italic> (2054), <italic>Betula pendula</italic> (4955), <italic>Carpinus betulus</italic> (1535), <italic>Fagus sylvatica</italic> (16 894), <italic>Fraxinus excelsior</italic> (7901), <italic>Prunus avium</italic> (1257), <italic>Quercus petraea</italic> (1302), <italic>Sorbus aucuparia</italic> (5473), <italic>Tilia platyphyllos</italic> (1982), and background (5408).</p>
      <p id="d1e835">Similar to the previous CNN<inline-formula><mml:math id="M35" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> classification task, the availability of training tiles varied greatly across the tree species. This class imbalance may have partially stemmed from the more systematic misclassification of certain classes during the CNN<inline-formula><mml:math id="M36" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> prediction. To reduce the unfavorable effects of a class imbalance on model training, we sampled 4000 tiles per class with replacement (similar to the CNN<inline-formula><mml:math id="M37" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> procedure). We applied the same data augmentation strategy as that used for the CNN<inline-formula><mml:math id="M38" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> workflow to increase variance among duplicates. A total of 20 % of the training data were withheld for model selection.</p>
      <p id="d1e875">We trained the U-Net architecture (CNN<inline-formula><mml:math id="M39" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>) using RMSprop as the optimizer with a learning rate of 0.0001 and an adapted Dice loss function. We adapted the Dice loss to ignore the weights coming from pixels with NA mask values. The models were trained with a batch size of 20 over 150 epochs.</p>
      <p id="d1e887">The CNN<inline-formula><mml:math id="M40" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> was then applied to Ortho<inline-formula><mml:math id="M41" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  and Ortho<inline-formula><mml:math id="M42" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>. To reduce uncertain predictions of CNN<inline-formula><mml:math id="M43" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>, we assigned the pixels where predicted probabilities for any of the tree species did not exceed 30 % to the background class. Thereby, we assumed that uncertain predictions predominantly occur in canopy gaps. As image segmentation typically suffers from increased uncertainty at tile edges, we repeated the predictions with horizontal and vertical shifts of 256 pixels, which were subsequently aggregated using a majority vote.</p>
      <p id="d1e926">The final model performance of CNN<inline-formula><mml:math id="M44" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> was assessed and compared to CNN<inline-formula><mml:math id="M45" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> using the independent reference data (transects) obtained from the visual interpretation of the UAV orthoimages.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results</title>
      <p id="d1e956">For the CNN<inline-formula><mml:math id="M46" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> method, F1 scores differed considerably across the tree species, although these differences were relatively consistent across the two orthoimages, i.e., Ortho<inline-formula><mml:math id="M47" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  and Ortho<inline-formula><mml:math id="M48" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="Ch1.F4"/>a, b). At the plot level, comparably high model performance (mean F1 <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula>) was found for <italic>Acer pseudoplatanus</italic> and <italic>Fraxinus excelsior</italic>; followed by intermediate performance (mean F1 score 0.35–0.55) for <italic>Aesculus hippocastanum</italic>, <italic>Sorbus aucuparia</italic>, <italic>Tilia platyphyllos</italic>, <italic>Betula pendula</italic>, and <italic>Carpinus betulus</italic>; and low performance (mean F1 score <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.35</mml:mn></mml:mrow></mml:math></inline-formula>) for <italic>Quercus petraea</italic>, <italic>Fagus sylvatica</italic>, and <italic>Prunus avium</italic>. Averaged across species, there was a slight decrease in model performance from Ortho<inline-formula><mml:math id="M51" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>,  with a mean F1 score of 0.44, to Ortho<inline-formula><mml:math id="M52" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>,  with a mean F1 score of 0.4 (Fig. <xref ref-type="fig" rid="Ch1.F4"/>a, b). Note that Ortho<inline-formula><mml:math id="M53" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  corresponded<?pagebreak page2916?> to the peak of the season, where leaves and canopies were still fully developed.</p>
      <p id="d1e1070">The CNN<inline-formula><mml:math id="M54" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> model performance across species was similar but generally higher compared with the CNN<inline-formula><mml:math id="M55" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> method. For Ortho<inline-formula><mml:math id="M56" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>,  F1 scores increased from 0.44 to 0.48 (Fig. <xref ref-type="fig" rid="Ch1.F4"/>a vs. c), while F1 scores increased from 0.40 to 0.46 for Ortho<inline-formula><mml:math id="M57" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="Ch1.F4"/>b vs. d).</p>
      <p id="d1e1114">We observed notable differences in model performance (mean F1) across different species mixtures: plots with one, two, or four species per plot (Fig. <xref ref-type="fig" rid="Ch1.F5"/>). For both CNN<inline-formula><mml:math id="M58" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> and CNN<inline-formula><mml:math id="M59" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>, the model performance strongly increased for a lower number of species per plot (Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F19"/>; the results for CNN<inline-formula><mml:math id="M60" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> are given in the Appendix).</p>
      <p id="d1e1148">The model performance of CNN<inline-formula><mml:math id="M61" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> exceeded the model performance of CNN<inline-formula><mml:math id="M62" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula>, particularly in plots with an increased number of species: for monocultures, the relative increase in model performance (F1 score) amounted to 2.5 %; for two species plots, the relative increase in model performance amounted to 6.9 %; and in plots with four species, the relative increase in model performance amounted to 20.9 % (averaged for Ortho<inline-formula><mml:math id="M63" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  and Ortho<inline-formula><mml:math id="M64" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>). This increased performance can be attributed to the advantages of the encoder–decoder principle of the CNN<inline-formula><mml:math id="M65" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> method, enabling a pixel-wise and contextual prediction at the original resolution of the orthomosaics. These advantages are also visible in Fig. <xref ref-type="fig" rid="Ch1.F6"/>, where CNN<inline-formula><mml:math id="M66" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> resulted in more detailed and accurate tree species segmentation (particularly for plots 26 and 29).</p>
      <p id="d1e1209">The highest model performance for CNN<inline-formula><mml:math id="M67" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> was found in monoculture plots, where F1 scores <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula> were found for 8 out of 10 species for both Ortho<inline-formula><mml:math id="M69" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula> and Ortho<inline-formula><mml:math id="M70" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>. A considerably lower performance for the July and September acquisition was found for <italic>Prunus avium</italic>, which may correspond to similarities in leaf and canopy structure with <italic>Fagus sylvatica</italic> and <italic>Fraxinus excelsior</italic> (a confusion matrix is given Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F17"/> in the Appendix). The decreased performance for <italic>Carpinus betulus</italic> and <italic>Prunus avium</italic> in Ortho<inline-formula><mml:math id="M71" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula> can be attributed to the very advanced senescence and leaf loss.</p>
      <p id="d1e1276">In addition to the increase in model performance, our analysis revealed that the prediction on orthoimagery using CNN<inline-formula><mml:math id="M72" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> only required 10 % of the computation time compared with CNN<inline-formula><mml:math id="M73" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula>. The duration of applying the models to the whole MyDiv orthomosaics covering an area of (3.02 ha; 0.22 cm ground sampling distance) took approximately 27.05 h with CNN<inline-formula><mml:math id="M74" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> and 264.88 h with CNN<inline-formula><mml:math id="M75" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> (NVIDIA A6000 with 48 GB RAM).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e1317">F1 scores by tree species and background class for Ortho<inline-formula><mml:math id="M76" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  and Ortho<inline-formula><mml:math id="M77" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>  derived from CNN<inline-formula><mml:math id="M78" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> and CNN<inline-formula><mml:math id="M79" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>: <bold>(a)</bold> CNN<inline-formula><mml:math id="M80" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> on Ortho<inline-formula><mml:math id="M81" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>, with a mean F1 score of 0.44; <bold>(b)</bold> CNN<inline-formula><mml:math id="M82" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> on Ortho<inline-formula><mml:math id="M83" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>, with a mean F1 score of 0.42; <bold>(c)</bold> CNN<inline-formula><mml:math id="M84" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> on Ortho<inline-formula><mml:math id="M85" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>, with a mean F1 score of 0.48; <bold>(d)</bold> CNN<inline-formula><mml:math id="M86" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> on Ortho<inline-formula><mml:math id="M87" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>, with a mean F1 score of 0.46.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f04.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e1450">The model performance (F1 score) of the CNN<inline-formula><mml:math id="M88" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> model across a gradient of canopy complexity in Ortho<inline-formula><mml:math id="M89" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>  and Ortho<inline-formula><mml:math id="M90" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>. F1 scores decrease with increasing canopy complexity in plots. Panel <bold>(a)</bold> shows performance across species mixtures on Ortho<inline-formula><mml:math id="M91" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula> with mean F1 scores for one species (0.51), two species (0.44), and four species (0.41). Panel <bold>(b)</bold> displays performance across species mixtures on Ortho<inline-formula><mml:math id="M92" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>  with mean F1 scores for one species (0.58), two species (0.51), and four species (0.42).</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f05.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e1514">The 2 m <inline-formula><mml:math id="M93" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M94" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M95" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions. Visualizations for the remaining plots are given in the Appendix (Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS1"/>).</p></caption>
        <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f06.jpg"/>

      </fig>

</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Discussion</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Filtering of citizen science data for drone-related applications</title>
      <p id="d1e1565">To achieve better correspondence between plant features visible in the citizen science photographs and the UAV images, we filtered the crowd-sourced photographs based on their acquisition distance (less than 0.3 m or greater than 15 m) to exclude macro and landscape photographs. Moreover, we excluded photographs that predominantly displayed tree stems, facilitating a foliage-centric perspective, as intrinsic to high-resolution UAV images (Fig. <xref ref-type="fig" rid="Ch1.F3"/>). In the future, more criteria may be considered for filtering citizen science imagery, including metadata (labels) on the presence of specific plant organs within an image (e.g., fruits and flowers), which are provided as a by-product by some citizen science plant identification apps (e.g., Pl@ntNet).</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>The creation of segmentation masks from simple image labels</title>
      <p id="d1e1578">One of the challenges of generating segmentation masks for the encoder–decoder method (CNN<inline-formula><mml:math id="M96" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>) with the proposed workflow may be error propagation between the different steps. Firstly, the CNN image classification trained on the citizen science data has varying uncertainty for the different species, resulting from noisy citizen science observations or limitations with respect to the identification of some species solely by photographs <xref ref-type="bibr" rid="bib1.bibx49" id="paren.33"/>. Secondly, the moving-window approach (CNN<inline-formula><mml:math id="M97" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula>), which predicts one species for an entire tile, may be too coarse to resemble very complex canopies (e.g., in highly diverse plant communities). However, although the fact that the segmentation labels created with the CNN<inline-formula><mml:math id="M98" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> approach are partially inaccurate (Figs. <xref ref-type="fig" rid="Ch1.F4"/>a, <xref ref-type="fig" rid="Ch1.F6"/>), we found that the CNN<inline-formula><mml:math id="M99" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> procedure indeed resulted in higher performance than the CNN<inline-formula><mml:math id="M100" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> procedure. This is in line with other studies <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx10 bib1.bibx43" id="paren.34"/> which have reported that deep-learning-based pattern recognition can partially overcome noisy labels, whereas the intentional use of noisy reference data, also known as weakly supervised learning, is generally very promising in the absence of high-quality labels <xref ref-type="bibr" rid="bib1.bibx9 bib1.bibx52 bib1.bibx43" id="paren.35"/>. Here, we filtered the training data (masks) for regions where we expected extreme noise levels – that is, for tiles where none of the classes exceeded a relative cover of 30 %. These regions were, according to our observation, often canopy gaps and shadowed areas, where one naturally expects lower model performance, as species-specific textures are less visible <xref ref-type="bibr" rid="bib1.bibx32 bib1.bibx36 bib1.bibx12" id="paren.36"/>.</p>
      <p id="d1e1643">The enhanced segmentation performance of the CNN<inline-formula><mml:math id="M101" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> approach compared with CNN<inline-formula><mml:math id="M102" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> can be attributed to the spatially explicit and more finely<?pagebreak page2917?> resolved predictions of the U-Net segmentation algorithm (encoder–decoder principle), enabling a segmentation of the tree species at the native resolution of the orthoimagery. The CNN<inline-formula><mml:math id="M103" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> approach resulted in improved prediction results compared with the CNN<inline-formula><mml:math id="M104" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> method in plots with more species and, hence, more complex canopies. Thus, the presented two-step approach of creating segmentation masks from simple class labels (CNN<inline-formula><mml:math id="M105" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula>), as provided by iNaturalist and Pl@ntNet platforms, can indeed be used to create segmentation masks required for state-of-the-art image analysis methods (CNN<inline-formula><mml:math id="M106" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula>) and, thereby, result in high value for remote sensing applications. The increased value of these segmentation masks enables the training of algorithms with higher performance in species recognition. It greatly enhances the computational efficiency of applying the models on orthoimagery (approximately 10 times faster). Especially for recurrent applications, such as monitoring or large-scale undertakings, the two-step approach involving the creation of segmentation masks and encoder–decoder architectures is recommended.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>The role of canopy complexity</title>
      <?pagebreak page2918?><p id="d1e1709">Overall, the segmentation performance declined with increasing species richness per plot. We expect that this can mainly be attributed to the small size of individual trees at the MyDiv site, where there is a lower chance that a 512 pixel <inline-formula><mml:math id="M107" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 512 pixel tile includes clearly visible species-specific leaf and branching patterns in species-rich mixtures. This also explains why, in particular, trees with lower relative canopy height (e.g., <italic>Quercus petraea</italic> and <italic>Fagus sylvatica</italic>) were less likely to be accurately segmented in species mixtures. The observed effect of canopy complexity is in line with previous findings from <xref ref-type="bibr" rid="bib1.bibx45" id="text.37"/>, <xref ref-type="bibr" rid="bib1.bibx31" id="text.38"/>, <xref ref-type="bibr" rid="bib1.bibx14" id="text.39"/>, and <xref ref-type="bibr" rid="bib1.bibx17" id="text.40"/>, who reported that smaller patches of individual species were less likely to be accurately detected. Visual inspection also confirmed that false predictions were more likely at canopy edges between different tree species (Fig. <xref ref-type="fig" rid="Ch1.F6"/>). However, it should be noted that the small-scale canopy complexity of the plots used here is exceptionally high (Fig. <xref ref-type="fig" rid="Ch1.F3"/>). Most tree crowns in the MyDiv experiment do not exceed a diameter of 1.5 m, and the transition among tree crowns of multiple species is often very fuzzy. Thus, we expect reduced performance in canopy transitions to be less relevant in real-world settings, where tree species appear in more extensive, homogeneous patches and where individual crowns are commonly larger. Thus, the model performance in these species mixtures can be interpreted as a rather conservative estimate. The results obtained for the monocultures might be more representative in terms of real-world applications, as mature trees in temperate forests typically have crown diameters 5–20 times larger. Application tests of the presented approach in real forests are desirable. However, acquiring such a dataset is a logistical challenge because temperate forest stands commonly do not feature a comparably high and balanced occurrence of that many tree species.</p>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Spatial resolution of the UAV imagery is key</title>
      <p id="d1e1750">According to the results obtained in the monocultures, the CNN<inline-formula><mml:math id="M108" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> model successfully classified 7 out of 10 tree species (F1 <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula>). The lower F1 scores for <italic>Quercus petraea</italic> (mean F1 0.57), <italic>Prunus avium</italic> (mean F1 0.2), and <italic>Tilia platyphyllos</italic> (mean F1 0.53) may result from the spectral and morphological similarity at the current spatial resolution of the UAV imagery (0.22 cm) (Fig. <xref ref-type="fig" rid="Ch1.F3"/>). Hence, these species were often confused with each other (see confusion matrices in Appendix <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>). Such confusion among plants with a similar appearance has been confirmed by other studies <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx41" id="paren.41"><named-content content-type="pre">e.g.,</named-content></xref> and matches our experience based on the generation of reference data via visual interpretation, where a separation between these species was sometimes challenging. Initial CNN-based segmentation<?pagebreak page2919?> attempts (results not shown) in the preparation of this study were based on an orthoimage with a resolution of 0.3 cm, instead of a 0.22 cm resolution, resulting in clearly lower model performance. This aligns with the reported importance of the spatial resolution of UAV imagery for CNN segmentation found earlier studies <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx44 bib1.bibx33 bib1.bibx5" id="paren.42"/>. Thus, while the current orthoimages with a 0.22 cm resolution delivered promising results, further increasing the spatial resolution might be very promising for species in which characteristic leaf forms are only visible at fine spatial resolutions.</p>
</sec>
<sec id="Ch1.S4.SS5">
  <label>4.5</label><title>Model transferability across seasons and orthoimage acquisition properties</title>
      <p id="d1e1803">The variability in human behavior and electronic devices makes citizen-science-based plant photographs very heterogeneous. This can be a challenge for deep learning applications, such as species recognition or plant trait characterization <xref ref-type="bibr" rid="bib1.bibx43 bib1.bibx50 bib1.bibx48 bib1.bibx1" id="paren.43"/>, in which models have to identify features that hold across various viewing angles, distances, or illumination conditions. However, this<?pagebreak page2920?> heterogeneity might also be of great value, given that citizens depict the appearance of plants under various site, environmental, and phenological conditions. This, in turn, offers a unique setting for training models that are generic and transferable across these conditions. Here, we evaluated the transferability of our models across different datasets by applying them to two orthoimages acquired in different seasons (peak growing season and autumn). Both the CNN<inline-formula><mml:math id="M110" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> and CNN<inline-formula><mml:math id="M111" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> models could identify deciduous tree species in the orthoimages with surprising accuracy, suggesting that the models are transferable to different conditions.</p>
</sec>
<sec id="Ch1.S4.SS6">
  <label>4.6</label><title>Outlook</title>
      <p id="d1e1835">Overall, our results indeed highlight the value of citizen science photographs with simple class labels to create training data for state-of-the-art segmentation approaches. A great advantage of this citizen-science-based approach is that it often does not require costly training data obtained from visual interpretation or field surveys (here, reference data were only used to validate the models). This particularly highlights the potential of citizen science data for applications in which many species are of interest, such as biodiversity-related monitoring <xref ref-type="bibr" rid="bib1.bibx8 bib1.bibx23" id="paren.44"/>. In this regard, data or models of species-recognition platforms that incorporate excessive numbers of plant species and respective imagery are very promising, including iNaturalist <xref ref-type="bibr" rid="bib1.bibx3" id="paren.45"/>, Pl@ntNet <xref ref-type="bibr" rid="bib1.bibx1" id="paren.46"/>, ObsIdentify <xref ref-type="bibr" rid="bib1.bibx37" id="paren.47"/>, or Flora Incognita <xref ref-type="bibr" rid="bib1.bibx34" id="paren.48"/>. However, based on the current work and an aforementioned precursor study <xref ref-type="bibr" rid="bib1.bibx45" id="paren.49"/>, we expect that a preselection of citizen science photograph databases considering images more representative of the common UAV-based perspective is required to unleash the potential of these heterogeneous data.</p><?xmltex \hack{\newpage}?>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e1868">The transfer learning approach presented here demonstrates the value of freely available crowd-sourced plant photographs for remote sensing studies. This heterogeneous dataset can provide valuable training data for transferable CNN-based segmentation models. Here, this potential was highlighted via a very complex task, i.e., the differentiation of 10 temperate deciduous tree species in mixed-vegetation stands with a complex structure. The presented two-step approach demonstrated how we can transfer and harness generic knowledge gathered by citizens on how plants “look” to the bird's-eye perspective of high-resolution drone imagery. The presented moving-window approach overcomes the limitation of citizen-science-based photographs having only simple species labels. The segmentation maps derived from an image classification model applied in a moving-window setting can be harnessed to create segmentation masks for encoder–decoder-type segmentation models. The latter not only allows higher accuracy with respect to species segmentation but is also considerably more efficient. By building on the effort of thousands of citizens, this framework enables the mapping of plant species without any training data obtained from visual interpretation or ground-based field surveys. Due to the large number of plant photographs acquired under different conditions, such models can be assumed to have good transferability.</p><?xmltex \hack{\clearpage}?>
</sec>

      
      </body>
    <back><app-group>

<?pagebreak page2921?><app id="App1.Ch1.S1">
  <?xmltex \currentcnt{A}?><label>Appendix A</label><title/>
<sec id="App1.Ch1.S1.SS1">
  <label>A1</label><title>Prediction maps</title>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F7"><?xmltex \currentcnt{A1}?><?xmltex \def\figurename{Figure}?><label>Figure A1</label><caption><p id="d1e1891">The 2 m <inline-formula><mml:math id="M112" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M113" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M114" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f07.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F8"><?xmltex \currentcnt{A2}?><?xmltex \def\figurename{Figure}?><label>Figure A2</label><caption><p id="d1e1930">The 2 m <inline-formula><mml:math id="M115" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M116" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M117" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f08.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F9"><?xmltex \currentcnt{A3}?><?xmltex \def\figurename{Figure}?><label>Figure A3</label><caption><p id="d1e1970">The 2 m <inline-formula><mml:math id="M118" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M119" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M120" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f09.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F10"><?xmltex \currentcnt{A4}?><?xmltex \def\figurename{Figure}?><label>Figure A4</label><caption><p id="d1e2009">The 2 m <inline-formula><mml:math id="M121" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M122" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M123" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f10.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F11"><?xmltex \currentcnt{A5}?><?xmltex \def\figurename{Figure}?><label>Figure A5</label><caption><p id="d1e2048">The 2 m <inline-formula><mml:math id="M124" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M125" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M126" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f11.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F12"><?xmltex \currentcnt{A6}?><?xmltex \def\figurename{Figure}?><label>Figure A6</label><caption><p id="d1e2088">The 2 m <inline-formula><mml:math id="M127" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M128" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M129" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f12.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F13"><?xmltex \currentcnt{A7}?><?xmltex \def\figurename{Figure}?><label>Figure A7</label><caption><p id="d1e2127">The 2 m <inline-formula><mml:math id="M130" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M131" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M132" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f13.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F14"><?xmltex \currentcnt{A8}?><?xmltex \def\figurename{Figure}?><label>Figure A8</label><caption><p id="d1e2166">The 2 m <inline-formula><mml:math id="M133" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M134" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M135" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f14.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F15"><?xmltex \currentcnt{A9}?><?xmltex \def\figurename{Figure}?><label>Figure A9</label><caption><p id="d1e2206">The 2 m <inline-formula><mml:math id="M136" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 20 m transects of selected plots, including the orthoimage, the reference, CNN<inline-formula><mml:math id="M137" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> predictions, and CNN<inline-formula><mml:math id="M138" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">segment</mml:mi></mml:msub></mml:math></inline-formula> predictions.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f15.jpg"/>

        </fig>

<?xmltex \hack{\clearpage}?>
</sec>
<?pagebreak page2930?><sec id="App1.Ch1.S1.SS2">
  <label>A2</label><title>Confusion matrix</title>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F16"><?xmltex \currentcnt{A10}?><?xmltex \def\figurename{Figure}?><label>Figure A10</label><caption><p id="d1e2254">Normalized confusion matrix of the CNN segment model applied to Ortho<inline-formula><mml:math id="M139" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula>.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f16.png"/>

        </fig>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F17"><?xmltex \currentcnt{A11}?><?xmltex \def\figurename{Figure}?><label>Figure A11</label><caption><p id="d1e2276">Normalized confusion matrix of the CNN segment model applied to the Ortho<inline-formula><mml:math id="M140" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula>.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f17.png"/>

        </fig>

<?xmltex \hack{\clearpage}?>
</sec>
<?pagebreak page2931?><sec id="App1.Ch1.S1.SS3">
  <label>A3</label><title>Data preprocessing</title>
      <p id="d1e2306">To reduce the heterogeneity of crowd-sourced photographs and match them with the UAV perspective, we filtered the photographs based on their acquisition distance and plant leaf visibility. The model achieved an <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula> and F1 <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula> on independent test data for both variables. Using predicted acquisition distance and tree trunk presence information for each photograph, we tested different filtering thresholds and combinations prior to training the CNN<inline-formula><mml:math id="M143" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> model for plant species classification. The best result was achieved by filtering out photographs taken from distances outside the range of 0.3–15 m and excluding photographs that were identified by the trained CNN classifier as containing tree trunks with a probability <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="App1.Ch1.S1.SS4">
  <label>A4</label><title>Citizen science data availability</title>

<?xmltex \floatpos{h!}?><table-wrap id="App1.Ch1.S1.T1"><?xmltex \currentcnt{A1}?><label>Table A1</label><caption><p id="d1e2364">Number of downloaded photographs for selected tree species from the iNaturalist and Pl@ntNet datasets.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">No.</oasis:entry>
         <oasis:entry colname="col2">Species</oasis:entry>
         <oasis:entry colname="col3">iNaturalist</oasis:entry>
         <oasis:entry colname="col4">Pl@ntNet</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2"><italic>Acer pseudoplatanus</italic></oasis:entry>
         <oasis:entry colname="col3">9999</oasis:entry>
         <oasis:entry colname="col4">3205</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2"><italic>Aesculus hippocastanum</italic></oasis:entry>
         <oasis:entry colname="col3">9998</oasis:entry>
         <oasis:entry colname="col4">1444</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2"><italic>Betula pendula</italic></oasis:entry>
         <oasis:entry colname="col3">9998</oasis:entry>
         <oasis:entry colname="col4">1308</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2"><italic>Carpinus betulus</italic></oasis:entry>
         <oasis:entry colname="col3">7165</oasis:entry>
         <oasis:entry colname="col4">2633</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2"><italic>Fagus sylvatica</italic></oasis:entry>
         <oasis:entry colname="col3">9981</oasis:entry>
         <oasis:entry colname="col4">3304</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">6</oasis:entry>
         <oasis:entry colname="col2"><italic>Fraxinus excelsior</italic></oasis:entry>
         <oasis:entry colname="col3">7745</oasis:entry>
         <oasis:entry colname="col4">3130</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">7</oasis:entry>
         <oasis:entry colname="col2"><italic>Prunus avium</italic></oasis:entry>
         <oasis:entry colname="col3">9999</oasis:entry>
         <oasis:entry colname="col4">3022</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">8</oasis:entry>
         <oasis:entry colname="col2"><italic>Quercus petraea</italic></oasis:entry>
         <oasis:entry colname="col3">1491</oasis:entry>
         <oasis:entry colname="col4">221</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">9</oasis:entry>
         <oasis:entry colname="col2"><italic>Sorbus aucuparia</italic></oasis:entry>
         <oasis:entry colname="col3">10 000</oasis:entry>
         <oasis:entry colname="col4">2730</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">10</oasis:entry>
         <oasis:entry colname="col2"><italic>Tilia platyphyllos</italic></oasis:entry>
         <oasis:entry colname="col3">582</oasis:entry>
         <oasis:entry colname="col4">1449</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{A1}?></table-wrap>

</sec>
<sec id="App1.Ch1.S1.SS5">
  <label>A5</label><title>Segmentation model architecture</title>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F18"><?xmltex \currentcnt{A12}?><?xmltex \def\figurename{Figure}?><label>Figure A12</label><caption><p id="d1e2574">A modified version (adapted from <xref ref-type="bibr" rid="bib1.bibx41" id="altparen.50"/>) of the U-Net CNN architecture for segmenting plant species from UAV orthoimages  <xref ref-type="bibr" rid="bib1.bibx39" id="paren.51"/>.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f18.png"/>

        </fig>

<?xmltex \hack{\clearpage}?>
</sec>
<?pagebreak page2932?><sec id="App1.Ch1.S1.SS6">
  <label>A6</label><title>CNN window species mixture box plot</title>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F19"><?xmltex \currentcnt{A13}?><?xmltex \def\figurename{Figure}?><label>Figure A13</label><caption><p id="d1e2603">The model performance (F1 score) of the CNN<inline-formula><mml:math id="M145" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">window</mml:mi></mml:msub></mml:math></inline-formula> model across a gradient of canopy complexity in Ortho<inline-formula><mml:math id="M146" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">July</mml:mi></mml:msub></mml:math></inline-formula> <bold>(a)</bold> and Ortho<inline-formula><mml:math id="M147" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">September</mml:mi></mml:msub></mml:math></inline-formula> <bold>(b)</bold>.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://bg.copernicus.org/articles/21/2909/2024/bg-21-2909-2024-f19.png"/>

        </fig>

</sec>
</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d1e2652">The code used in this study is publicly accessible via the GitHub repository at <uri>https://github.com/salimsoltani28/CrowdVision2TreeSegment</uri> <xref ref-type="bibr" rid="bib1.bibx25" id="paren.52"/>. The data supporting the findings of this research are available on Zenodo at <ext-link xlink:href="https://doi.org/10.5281/zenodo.10019552" ext-link-type="DOI">10.5281/zenodo.10019552</ext-link> <xref ref-type="bibr" rid="bib1.bibx25" id="paren.53"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e2670">SS contributed to the conceptualization, methodology, formal analysis, data curation, visualization, and writing – original draft preparation. OF provided resources and contributed to writing – review and editing. NE also provided resources and contributed to writing – review and editing. HF contributed to funding acquisition, supervision, and writing – review and editing. TK contributed to conceptualization, data collection, funding acquisition, data curation, resource acquisition, supervision, and writing – original draft preparation.</p>
  </notes><?xmltex \hack{\newpage}?><?xmltex \hack{\vspace*{145mm}}?><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e2678">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e2684">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e2690">Salim Soltani and Teja Kattenborn acknowledge funding from the German Research Foundation (DFG) within the framework of BigPlantSens (Assessing the Synergies of Big Data and Deep Learning for the Remote Sensing of Plant Species; project no. 444524904) and PANOPS (Revealing Earth's plant functional diversity with citizen science; project no. 504978936). Salim Soltani and Hannes Feilhauer acknowledge financial<?pagebreak page2933?> support from the Federal Ministry of Education and Research of Germany (BMBF) and the Sächsische Staatsministerium für Wissenschaft, Kultur und Tourismus within the framework of the Center of Excellence for AI Research “Center for Scalable Data Analytics and Artificial Intelligence Dresden/Leipzig” program (project ID ScaDS.AI). Nico Eisenhauer and Olga Ferlian acknowledge funding from the DFG (German Centre for Integrative Biodiversity Research, FZT118, and Gottfried Wilhelm Leibniz Prize, Ei 862/29-1). Moreover, the authors acknowledge support from University of Freiburg with respect to open-access publishing.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e2695">This research has been supported by the German Research Foundation (DFG) within the framework of BigPlantSens (grant no. 44452490) and PANOPS (grant no. 04978936).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e2702">This paper was edited by Paul Stoy and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Affouard et~al.(2017)Affouard, Go{\"{e}}au, Bonnet, Lombardo, and
Joly}}?><label>Affouard et al.(2017)Affouard, Goëau, Bonnet, Lombardo, and Joly</label><?label affouard2017pl?><mixed-citation> Affouard, A., Goëau, H., Bonnet, P., Lombardo, J.-C., and Joly, A.: Pl@ntnet app in the era of deep learning, in: ICLR: International Conference on Learning Representations, April 2017, Toulon, France, ffhal-01629195f, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Bayraktar et~al.(2020)Bayraktar, Basarkan, and
Celebi}}?><label>Bayraktar et al.(2020)Bayraktar, Basarkan, and Celebi</label><?label bayraktar2020low?><mixed-citation>Bayraktar, E., Basarkan, M. E., and Celebi, N.: A low-cost UAV framework towards ornamental plant detection and counting in the wild, ISPRS J. Photogramm., 167, 1–11, <ext-link xlink:href="https://doi.org/10.1016/j.isprsjprs.2020.06.012" ext-link-type="DOI">10.1016/j.isprsjprs.2020.06.012</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Boone and Basille(2019)}}?><label>Boone and Basille(2019)</label><?label boone2019using?><mixed-citation> Boone, M. E. and Basille, M.: Using iNaturalist to contribute your nature observations to science, EDIS, 2019, 5–5, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Bouguettaya et~al.(2022)Bouguettaya, Zarzour, Kechida, and
Taberkit}}?><label>Bouguettaya et al.(2022)Bouguettaya, Zarzour, Kechida, and Taberkit</label><?label bouguettaya2022deep?><mixed-citation> Bouguettaya, A., Zarzour, H., Kechida, A., and Taberkit, A. M.: Deep learning techniques to classify agricultural crops through UAV imagery: A review, Neural Comput. Appl., 34, 9511–9536, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Braga et~al.(2020)Braga, Peripato, Dalagnol, P.~Ferreira, Tarabalka,
OC~Arag{\~{a}}o, F.~de Campos~Velho, Shiguemori, and Wagner}}?><label>Braga et al.(2020)Braga, Peripato, Dalagnol, P. Ferreira, Tarabalka, OC Aragão, F. de Campos Velho, Shiguemori, and Wagner</label><?label g2020tree?><mixed-citation>Braga, G., J. R., Peripato, V., Dalagnol, R., P. Ferreira, M., Tarabalka, Y., OC Aragão, L. E., F. de Campos Velho, H., Shiguemori, E. H., and Wagner, F. H.: Tree crown delineation algorithm based on a convolutional neural network, Remote Sens., 12, 1288, <ext-link xlink:href="https://doi.org/10.3390/rs12081288" ext-link-type="DOI">10.3390/rs12081288</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Brandt et~al.(2020)Brandt, Tucker, Kariryaa, Rasmussen, Abel, Small,
Chave, Rasmussen, Hiernaux, Diouf et~al.}}?><label>Brandt et al.(2020)Brandt, Tucker, Kariryaa, Rasmussen, Abel, Small, Chave, Rasmussen, Hiernaux, Diouf et al.</label><?label brandt2020unexpectedly?><mixed-citation>Brandt, M., Tucker, C. J., Kariryaa, A., Rasmussen, K., Abel, C., Small, J., Chave, J., Rasmussen, L. V., Hiernaux, P., Diouf, A. A., Kergoat, L., Mertz, O., Igel, C., Gieseke, F., Schöning, J., Li, S., Melocik, K., Meyer, J., Sinno, S., Romero, E., Glennie, E., Montagu, A., Dendoncker, M., and Fensholt, R.: An unexpectedly large count of trees in the West African Sahara and Sahel, Nature, 587, 78–82, <ext-link xlink:href="https://doi.org/10.1038/s41586-020-2824-5" ext-link-type="DOI">10.1038/s41586-020-2824-5</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Brodrick et~al.(2019)Brodrick, Davies, and
Asner}}?><label>Brodrick et al.(2019)Brodrick, Davies, and Asner</label><?label brodrick2019uncovering?><mixed-citation>Brodrick, P. G., Davies, A. B., and Asner, G. P.: Uncovering ecological patterns with convolutional neural networks, Trends Ecol. Evol., 34, 734–745, <ext-link xlink:href="https://doi.org/10.1016/j.tree.2019.03.006" ext-link-type="DOI">10.1016/j.tree.2019.03.006</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Chandler et~al.(2017)Chandler, See, Copas, Bonde, L{\'{o}}pez,
Danielsen, Legind, Masinde, Miller-Rushing, Newman
et~al.}}?><label>Chandler et al.(2017)Chandler, See, Copas, Bonde, López, Danielsen, Legind, Masinde, Miller-Rushing, Newman et al.</label><?label chandler2017contribution?><mixed-citation>Chandler, M., See, L., Copas, K., Bonde, A. M., López, B. C., Danielsen, F., Legind, J. K., Masinde, S., Miller-Rushing, A. J., Newman, G., Rosemartin, A., and Turak, E.: Contribution of citizen science towards international biodiversity monitoring, Biol. Conserv., 213, 280–294, <ext-link xlink:href="https://doi.org/10.1016/j.biocon.2016.09.004" ext-link-type="DOI">10.1016/j.biocon.2016.09.004</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{Cherif et~al.(2023)Cherif, Feilhauer, Berger, Dao, Ewald, Hank, He,
Kovach, Lu, Townsend et~al.}}?><label>Cherif et al.(2023)Cherif, Feilhauer, Berger, Dao, Ewald, Hank, He, Kovach, Lu, Townsend et al.</label><?label cherif2023spectra?><mixed-citation>Cherif, E., Feilhauer, H., Berger, K., Dao, P. D., Ewald, M., Hank, T. B., He, Y., Kovach, K. R., Lu, B., Townsend, P. A., and Kattenborn, T.: From spectra to plant functional traits: Transferable multi-trait models from heterogeneous and sparse data, Remote Sens. Environ., 292, 113580, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2023.113580" ext-link-type="DOI">10.1016/j.rse.2023.113580</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Cloutier et~al.(2023)Cloutier, Germain, and
Lalibert{\'{e}}}}?><label>Cloutier et al.(2023)Cloutier, Germain, and Laliberté</label><?label cloutier2023influence?><mixed-citation>Cloutier, M., Germain, M., and Laliberté, E.: Influence of Temperate Forest Autumn Leaf Phenology on Segmentation of Tree Species from UAV Imagery Using Deep Learning, bioRxiv,  2023–08, <ext-link xlink:href="https://doi.org/10.1101/2023.08.03.548604" ext-link-type="DOI">10.1101/2023.08.03.548604</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Curnick et~al.(2021)Curnick, Davies, Duncan, Freeman, Jacoby,
Shelley, Rossi, Wearn, Williamson, and Pettorelli}}?><label>Curnick et al.(2021)Curnick, Davies, Duncan, Freeman, Jacoby, Shelley, Rossi, Wearn, Williamson, and Pettorelli</label><?label curnick2021smallsats?><mixed-citation>Curnick, D. J., Davies, A. J., Duncan, C., Freeman, R., Jacoby, D. M., Shelley, H. T., Rossi, C., Wearn, O. R., Williamson, M. J., and Pettorelli, N.: SmallSats: a new technological frontier in ecology and conservation?, Remote Sensing in Ecology and Conservation, 8, 139–150, <ext-link xlink:href="https://doi.org/10.1002/rse2.239" ext-link-type="DOI">10.1002/rse2.239</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{De~Sa et~al.(2018)De~Sa, Castro, Carvalho, Marchante,
L{\'{o}}pez-N{\'{u}}{\~{n}}ez, and Marchante}}?><label>De Sa et al.(2018)De Sa, Castro, Carvalho, Marchante, López-Núñez, and Marchante</label><?label de2018mapping?><mixed-citation>De Sa, N. C., Castro, P., Carvalho, S., Marchante, E., López-Núñez, F. A., and Marchante, H.: Mapping the flowering of an invasive plant using unmanned aerial vehicles: is there potential for biocontrol monitoring?, Front. Plant Sci., 9, 293, <ext-link xlink:href="https://doi.org/10.3389/fpls.2018.00293" ext-link-type="DOI">10.3389/fpls.2018.00293</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Di~Cecco et~al.(2021)Di~Cecco, Barve, Belitz, Stucky, Guralnick, and
Hurlbert}}?><label>Di Cecco et al.(2021)Di Cecco, Barve, Belitz, Stucky, Guralnick, and Hurlbert</label><?label di2021observing?><mixed-citation>Di Cecco, G. J., Barve, V., Belitz, M. W., Stucky, B. J., Guralnick, R. P., and Hurlbert, A. H.: Observing the observers: How participants contribute data to iNaturalist and implications for biodiversity science, BioScience, 71, 1179–1188, <ext-link xlink:href="https://doi.org/10.1093/biosci/biab093" ext-link-type="DOI">10.1093/biosci/biab093</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Fassnacht et~al.(2016)Fassnacht, Latifi, Stere{\'{n}}czak, Modzelewska,
Lefsky, Waser, Straub, and Ghosh}}?><label>Fassnacht et al.(2016)Fassnacht, Latifi, Stereńczak, Modzelewska, Lefsky, Waser, Straub, and Ghosh</label><?label fassnacht2016review?><mixed-citation>Fassnacht, F. E., Latifi, H., Stereńczak, K., Modzelewska, A., Lefsky, M., Waser, L. T., Straub, C., and Ghosh, A.: Review of studies on tree species classification from remotely sensed data, Remote Sens. Environ., 186, 64–87, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2016.08.013" ext-link-type="DOI">10.1016/j.rse.2016.08.013</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Ferlian et~al.(2018)Ferlian, Cesarz, Craven, Hines, Barry,
Bruelheide, Buscot, Haider, Heklau, Herrmann et~al.}}?><label>Ferlian et al.(2018)Ferlian, Cesarz, Craven, Hines, Barry, Bruelheide, Buscot, Haider, Heklau, Herrmann et al.</label><?label ferlian2018mycorrhiza?><mixed-citation>Ferlian, O., Cesarz, S., Craven, D., Hines, J., Barry, K. E., Bruelheide, H., Buscot, F., Haider, S., Heklau, H., Herrmann, S., Kühn, P.,Pruschitzki, U., Schädler, M., Wagg, C., Weigelt, A., Wubet, T., and Eisenhauer, N.: Mycorrhiza in tree diversity–ecosystem function relationships: conceptual framework and experimental implementation, Ecosphere, 9, e02226, <ext-link xlink:href="https://doi.org/10.1002/ecs2.2226" ext-link-type="DOI">10.1002/ecs2.2226</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Fraisl et~al.(2022)Fraisl, Hager, Bedessem, Gold, Hsing, Danielsen,
Hitchcock, Hulbert, Piera, Spiers et~al.}}?><label>Fraisl et al.(2022)Fraisl, Hager, Bedessem, Gold, Hsing, Danielsen, Hitchcock, Hulbert, Piera, Spiers et al.</label><?label fraisl2022citizen?><mixed-citation>Fraisl, D., Hager, G., Bedessem, B., Gold, M., Hsing, P.-Y., Danielsen, F., Hitchcock, C. B., Hulbert, J. M., Piera, J., Spiers, H., Thiel, M., and Haklay, M.: Citizen science in environmental and ecological sciences, Nature Reviews Methods Primers, 2, 64, <ext-link xlink:href="https://doi.org/10.1038/s43586-022-00144-4" ext-link-type="DOI">10.1038/s43586-022-00144-4</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Fricker et~al.(2019)Fricker, Ventura, Wolf, North, Davis, and
Franklin}}?><label>Fricker et al.(2019)Fricker, Ventura, Wolf, North, Davis, and Franklin</label><?label fricker2019convolutional?><mixed-citation>Fricker, G. A., Ventura, J. D., Wolf, J. A., North, M. P., Davis, F. W., and Franklin, J.: A convolutional neural network classifier identifies tree species in mixed-conifer forest from hyperspectral imagery, Remote Sens., 11, 2326, <ext-link xlink:href="https://doi.org/10.3390/rs11192326" ext-link-type="DOI">10.3390/rs11192326</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Galuszynski et~al.(2022)Galuszynski, Duker, Potts, and
Kattenborn}}?><label>Galuszynski et al.(2022)Galuszynski, Duker, Potts, and Kattenborn</label><?label galuszynski2022automated?><mixed-citation>Galuszynski, N. C., Duker, R., Potts, A. J., and Kattenborn, T.: Automated mapping of Portulacaria afra canopies for restoration monitoring with convolutional neural networks and heterogeneous unmanned aerial vehicle imagery, PeerJ, 10, e14219, <ext-link xlink:href="https://doi.org/10.7717/peerj.14219" ext-link-type="DOI">10.7717/peerj.14219</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{GBIF(2019)}}?><label>GBIF(2019)</label><?label gbif2019gbif?><mixed-citation> GBIF: GBIF: the global biodiversity information facility, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Hoeser and Kuenzer(2020)}}?><label>Hoeser and Kuenzer(2020)</label><?label hoeser2020object?><mixed-citation>Hoeser, T. and Kuenzer, C.: Object detection and image segmentation with deep learning on earth observation data: A review-part i: Evolution and recent trends, Remote Sens., 12, 1667, <ext-link xlink:href="https://doi.org/10.3390/rs12101667" ext-link-type="DOI">10.3390/rs12101667</ext-link>, 2020.</mixed-citation></ref>
      <?pagebreak page2934?><ref id="bib1.bibx21"><?xmltex \def\ref@label{{iNaturalist(2023)}}?><label>iNaturalist(2023)</label><?label iNaturalist?><mixed-citation>iNaturalist: iNaturalist, <uri>https://www.inaturalist.org</uri>, last access: 4 September 2023.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Ivanova and Shashkov(2021)}}?><label>Ivanova and Shashkov(2021)</label><?label ivanova2021possibilities?><mixed-citation> Ivanova, N. and Shashkov, M.: The possibilities of GBIF data use in ecological research, Russ. J. Ecol., 52, 1–8, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Johnston et~al.(2023)Johnston, Matechou, and
Dennis}}?><label>Johnston et al.(2023)Johnston, Matechou, and Dennis</label><?label johnston2023outstanding?><mixed-citation>Johnston, A., Matechou, E., and Dennis, E. B.: Outstanding challenges and future directions for biodiversity monitoring using citizen science data, Methods in Ecol. Evol., 14, 103–116, <ext-link xlink:href="https://doi.org/10.1111/2041-210X.13834" ext-link-type="DOI">10.1111/2041-210X.13834</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Joly et~al.(2016)Joly, Bonnet, Go{\"{e}}au, Barbe, Selmi, Champ,
Dufour-Kowalski, Affouard, Carr{\'{e}}, Molino et~al.}}?><label>Joly et al.(2016)Joly, Bonnet, Goëau, Barbe, Selmi, Champ, Dufour-Kowalski, Affouard, Carré, Molino et al.</label><?label joly2016look?><mixed-citation> Joly, A., Bonnet, P., Goëau, H., Barbe, J., Selmi, S., Champ, J., Dufour-Kowalski, S., Affouard, A., Carré, J., Molino, J. F., Boujemaa, N., and Barthélémy, D.: A look inside the Pl@ ntNet experience: The good, the bias and the hope, Multimedia Syst., 22, 751–766, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Kattenborn and Soltani(2023)}}?><label>Kattenborn and Soltani(2023)</label><?label data?><mixed-citation>Kattenborn, T. and Soltani, S.: CrowdVision2TreeSegment, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.10019552" ext-link-type="DOI">10.5281/zenodo.10019552</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Kattenborn et~al.(2019{\natexlab{a}})Kattenborn, Eichel, and
Fassnacht}}?><label>Kattenborn et al.(2019a)Kattenborn, Eichel, and Fassnacht</label><?label kattenborn2019convolutional?><mixed-citation>Kattenborn, T., Eichel, J., and Fassnacht, F. E.: Convolutional Neural Networks enable efficient, accurate and fine-grained segmentation of plant species and communities from high-resolution UAV imagery, Sci. Rep., 9, 1–9, <ext-link xlink:href="https://doi.org/10.1038/s41598-019-53797-9" ext-link-type="DOI">10.1038/s41598-019-53797-9</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Kattenborn et~al.(2019{\natexlab{b}})Kattenborn, Lopatin,
F{\"{o}}rster, Braun, and Fassnacht}}?><label>Kattenborn et al.(2019b)Kattenborn, Lopatin, Förster, Braun, and Fassnacht</label><?label kattenborn2019uav?><mixed-citation>Kattenborn, T., Lopatin, J., Förster, M., Braun, A. C., and Fassnacht, F. E.: UAV data as alternative to field sampling to map woody invasive species based on combined Sentinel-1 and Sentinel-2 data, Remote Sens. Environ., 227, 61–73, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2019.03.025" ext-link-type="DOI">10.1016/j.rse.2019.03.025</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Kattenborn et~al.(2021)Kattenborn, Leitloff, Schiefer, and
Hinz}}?><label>Kattenborn et al.(2021)Kattenborn, Leitloff, Schiefer, and Hinz</label><?label kattenborn2021review?><mixed-citation>Kattenborn, T., Leitloff, J., Schiefer, F., and Hinz, S.: Review on Convolutional Neural Networks (CNN) in vegetation remote sensing, ISPRS J. Photogramm., 173, 24–49, <ext-link xlink:href="https://doi.org/10.1016/j.isprsjprs.2020.12.010" ext-link-type="DOI">10.1016/j.isprsjprs.2020.12.010</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Kattenborn et~al.(2022)Kattenborn, Schiefer, Frey, Feilhauer,
Mahecha, and Dormann}}?><label>Kattenborn et al.(2022)Kattenborn, Schiefer, Frey, Feilhauer, Mahecha, and Dormann</label><?label kattenborn2022spatially?><mixed-citation>Kattenborn, T., Schiefer, F., Frey, J., Feilhauer, H., Mahecha, M. D., and Dormann, C. F.: Spatially autocorrelated training and validation samples inflate performance assessment of convolutional neural networks, ISPRS Open Journal of Photogrammetry and Remote Sensing, 5, 100018, <ext-link xlink:href="https://doi.org/10.1016/j.ophoto.2022.100018" ext-link-type="DOI">10.1016/j.ophoto.2022.100018</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Leit{\~{a}}o et~al.(2018)Leit{\~{a}}o, Schwieder, P{\"{o}}tzschner, Pinto,
Teixeira, Pedroni, Sanchez, Rogass, van~der Linden, Bustamante
et~al.}}?><label>Leitão et al.(2018)Leitão, Schwieder, Pötzschner, Pinto, Teixeira, Pedroni, Sanchez, Rogass, van der Linden, Bustamante et al.</label><?label leitao2018sample?><mixed-citation>Leitão, P. J., Schwieder, M., Pötzschner, F., Pinto, J. R., Teixeira, A. M., Pedroni, F., Sanchez, M., Rogass, C., van der Linden, S., Bustamante, M. M. and Hostert, P.: From sample to pixel: multi-scale remote sensing data for upscaling aboveground carbon data in heterogeneous landscapes, Ecosphere, 9, e02298, <ext-link xlink:href="https://doi.org/10.1002/ecs2.2298" ext-link-type="DOI">10.1002/ecs2.2298</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Lopatin et~al.(2017)Lopatin, Fassnacht, Kattenborn, and
Schmidtlein}}?><label>Lopatin et al.(2017)Lopatin, Fassnacht, Kattenborn, and Schmidtlein</label><?label lopatin2017mapping?><mixed-citation>Lopatin, J., Fassnacht, F. E., Kattenborn, T., and Schmidtlein, S.: Mapping plant species in mixed grassland communities using close range imaging spectroscopy, Remote Sens. Environ., 201, 12–23, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2017.08.031" ext-link-type="DOI">10.1016/j.rse.2017.08.031</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Lopatin et~al.(2019)Lopatin, Dolos, Kattenborn, and
Fassnacht}}?><label>Lopatin et al.(2019)Lopatin, Dolos, Kattenborn, and Fassnacht</label><?label lopatin2019canopy?><mixed-citation>Lopatin, J., Dolos, K., Kattenborn, T., and Fassnacht, F. E.: How canopy shadow affects invasive plant species classification in high spatial resolution remote sensing, Remote Sens. Ecol. Conserv., 5, 302–317, <ext-link xlink:href="https://doi.org/10.1002/rse2.109" ext-link-type="DOI">10.1002/rse2.109</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{{Ma et~al.(2019)Ma, Liu, Zhang, Ye, Yin, and Johnson}}?><label>Ma et al.(2019)Ma, Liu, Zhang, Ye, Yin, and Johnson</label><?label ma2019deep?><mixed-citation>Ma, L., Liu, Y., Zhang, X., Ye, Y., Yin, G., and Johnson, B. A.: Deep learning in remote sensing applications: A meta-analysis and review, ISPRS J. Photogramm., 152, 166–177, <ext-link xlink:href="https://doi.org/10.1016/j.isprsjprs.2019.04.015" ext-link-type="DOI">10.1016/j.isprsjprs.2019.04.015</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{M{\"{a}}der et~al.(2021)M{\"{a}}der, Boho, Rzanny, Seeland, Wittich,
Deggelmann, and W{\"{a}}ldchen}}?><label>Mäder et al.(2021)Mäder, Boho, Rzanny, Seeland, Wittich, Deggelmann, and Wäldchen</label><?label mader2021flora?><mixed-citation>Mäder, P., Boho, D., Rzanny, M., Seeland, M., Wittich, H. C., Deggelmann, A., and Wäldchen, J.: The flora incognita app–interactive plant species identification, Methods in Ecol. Evol., 12,   1335–1342, <ext-link xlink:href="https://doi.org/10.1111/2041-210X.13611" ext-link-type="DOI">10.1111/2041-210X.13611</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Maes and Steppe(2019)}}?><label>Maes and Steppe(2019)</label><?label maes2019perspectives?><mixed-citation>Maes, W. H. and Steppe, K.: Perspectives for remote sensing with unmanned aerial vehicles in precision agriculture, Trends Plant Sci., 24, 152–164, <ext-link xlink:href="https://doi.org/10.1016/j.tplants.2018.11.007" ext-link-type="DOI">10.1016/j.tplants.2018.11.007</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{Milas et~al.(2017)Milas, Arend, Mayer, Simonson, and
Mackey}}?><label>Milas et al.(2017)Milas, Arend, Mayer, Simonson, and Mackey</label><?label milas2017different?><mixed-citation>Milas, A. S., Arend, K., Mayer, C., Simonson, M. A., and Mackey, S.: Different colours of shadows: Classification of UAV images, Int. J. Remote Sens., 38, 3084–3100, <ext-link xlink:href="https://doi.org/10.1080/01431161.2016.1274449" ext-link-type="DOI">10.1080/01431161.2016.1274449</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{{Molls(2021)}}?><label>Molls(2021)</label><?label molls2021obs?><mixed-citation> Molls, C.: The Obs-Services and their potentials for biodiversity data assessments with a test of the current reliability of photo-identification of Coleoptera in the field, Tijdschrift voor Entomologie, 164, 143–153, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{{M{\"{u}}llerov{\'{a}} et~al.(2023)M{\"{u}}llerov{\'{a}}, Brundu,
Gro{\ss}e-Stoltenberg, Kattenborn, and Richardson}}?><label>Müllerová et al.(2023)Müllerová, Brundu, Große-Stoltenberg, Kattenborn, and Richardson</label><?label mullerova2023pattern?><mixed-citation>Müllerová, J., Brundu, G., Große-Stoltenberg, A., Kattenborn, T., and Richardson, D. M.: Pattern to process, research to practice: remote sensing of plant invasions, Biol. Invasions,  25,  3651–3676, <ext-link xlink:href="https://doi.org/10.1007/s10530-023-03150-z" ext-link-type="DOI">10.1007/s10530-023-03150-z</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{{Ronneberger et~al.(2015)Ronneberger, Fischer, and
Brox}}?><label>Ronneberger et al.(2015)Ronneberger, Fischer, and Brox</label><?label ronneberger2015u?><mixed-citation>Ronneberger, O., Fischer, P., and Brox, T.: U-net: Convolutional networks for biomedical image segmentation, in: International Conference on Medical image computing and computer-assisted intervention, Springer, 234–241, <ext-link xlink:href="https://doi.org/10.1007/978-3-319-24574-4_28" ext-link-type="DOI">10.1007/978-3-319-24574-4_28</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{{Rzanny et~al.(2019)Rzanny, M{\"{a}}der, Deggelmann, Chen, and
W{\"{a}}ldchen}}?><label>Rzanny et al.(2019)Rzanny, Mäder, Deggelmann, Chen, and Wäldchen</label><?label rzanny2019flowers?><mixed-citation>Rzanny, M., Mäder, P., Deggelmann, A., Chen, M., and Wäldchen, J.: Flowers, leaves or both? How to obtain suitable images for automated plant identification, Plant Methods, 15, 1–11, <ext-link xlink:href="https://doi.org/10.1186/s13007-019-0462-4" ext-link-type="DOI">10.1186/s13007-019-0462-4</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{{Schiefer et~al.(2020)Schiefer, Kattenborn, Frick, Frey, Schall, Koch,
and Schmidtlein}}?><label>Schiefer et al.(2020)Schiefer, Kattenborn, Frick, Frey, Schall, Koch, and Schmidtlein</label><?label schiefer2020mapping?><mixed-citation>Schiefer, F., Kattenborn, T., Frick, A., Frey, J., Schall, P., Koch, B., and Schmidtlein, S.: Mapping forest tree species in high resolution UAV-based RGB-imagery by means of convolutional neural networks, ISPRS J. Photogramm., 170, 205–215, <ext-link xlink:href="https://doi.org/10.1016/j.isprsjprs.2020.10.015" ext-link-type="DOI">10.1016/j.isprsjprs.2020.10.015</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{{Schiefer et~al.(2023)Schiefer, Schmidtlein, Frick, Frey, Klinke,
Zielewska-B{\"{u}}ttner, Junttila, Uhl, and Kattenborn}}?><label>Schiefer et al.(2023)Schiefer, Schmidtlein, Frick, Frey, Klinke, Zielewska-Büttner, Junttila, Uhl, and Kattenborn</label><?label schiefer2023uav?><mixed-citation>Schiefer, F., Schmidtlein, S., Frick, A., Frey, J., Klinke, R., Zielewska-Büttner, K., Junttila, S., Uhl, A., and Kattenborn, T.: UAV-based reference data for the prediction of fractional cover of standing deadwood from Sentinel time series, ISPRS Open Journal of Photogrammetry and Remote Sensing, 8, 100034, <ext-link xlink:href="https://doi.org/10.1016/j.ophoto.2023.100034" ext-link-type="DOI">10.1016/j.ophoto.2023.100034</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{{Schiller et~al.(2021)Schiller, Schmidtlein, Boonman,
Moreno-Mart{\'{\i}}nez, and Kattenborn}}?><label>Schiller et al.(2021)Schiller, Schmidtlein, Boonman, Moreno-Martínez, and Kattenborn</label><?label schiller2021deep?><mixed-citation> Schiller, C., Schmidtlein, S., Boonman, C., Moreno-Martínez, A., and Kattenborn, T.: Deep learning and citizen science enable automated plant trait predictions from photographs, Sci. Rep., 11, 1–12, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx44"><?xmltex \def\ref@label{{Schmitt et~al.(2020)Schmitt, Prexl, Ebel, Liebel, and
Zhu}}?><label>Schmitt et al.(2020)Schmitt, Prexl, Ebel, Liebel, and Zhu</label><?label schmitt2020weakly?><mixed-citation>Schmitt, M., Prexl, J., Ebel, P., Liebel, L., and Zhu, X. X.: Weakly supervised semantic segmentation of satellite images for land cover mapping–challenges and opportunities, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.2002.08254" ext-link-type="DOI">10.48550/arXiv.2002.08254</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{Soltani et~al.(2022)Soltani, Feilhauer, Duker, and
Kattenborn}}?><label>Soltani et al.(2022)Soltani, Feilhauer, Duker, and Kattenborn</label><?label soltani2022transfer?><mixed-citation>Soltani, S., Feilhauer, H., Duker, R., and Kattenborn, T.: Transfer learning from citizen science photographs enables plant species identification in UAVs imagery, ISPRS Open Journal of Photogrammetry and Remote Sensing, 5, 100016, <ext-link xlink:href="https://doi.org/10.1016/j.ophoto.2022.100016" ext-link-type="DOI">10.1016/j.ophoto.2022.100016</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{{Sun et~al.(2021)Sun, Wang, Wang, Yang, Xie, and Huang}}?><label>Sun et al.(2021)Sun, Wang, Wang, Yang, Xie, and Huang</label><?label sun2021uavs?><mixed-citation>Sun, Z., Wang, X., Wang, Z., Yang, L., Xie, Y., and Huang, Y.: UAVs as remote sensing platforms in plant ecology: review of applications and challenges, J. Plant Ecol., 14, 1003–1023, <ext-link xlink:href="https://doi.org/10.1093/jpe/rtab089" ext-link-type="DOI">10.1093/jpe/rtab089</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{{Tan and Le(2019)}}?><label>Tan and Le(2019)</label><?label tan2019efficientnet?><mixed-citation>Tan, M. and Le, Q.: Efficientnet: Rethinking model scaling for convolutional neural networks, in: International conference on machine learning, 6105–6114, PMLR, Long Beach, California, 10–15 June 2019, <ext-link xlink:href="https://doi.org/10.48550/arXiv.1905.11946" ext-link-type="DOI">10.48550/arXiv.1905.11946</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{van Der~Velde et~al.(2023)van Der~Velde, Go{\"{e}}au, Bonnet,
d'Andrimont, Yordanov, Affouard, Claverie, Cz{\'{u}}cz, Elvekj{\ae}r,
Martinez-Sanchez et~al.}}?><label>van Der Velde et al.(2023)van Der Velde, Goëau, Bonnet, d'Andrimont, Yordanov, Affouard<?pagebreak page2935?>, Claverie, Czúcz, Elvekjær, Martinez-Sanchez et al.</label><?label van2023pl?><mixed-citation>van Der Velde, M., Goëau, H., Bonnet, P., d’Andrimont, R., Yordanov, M., Affouard, A., Claverie, M., Czúcz, B., Elvekjær, N., Martinez-Sanchez, L., and Rotllan-Puig, X.: Pl@ ntNet Crops: merging citizen science observations and structured survey data to improve crop recognition for agri-food-environment applications, Environ. Res. Lett., 18, 025005, <ext-link xlink:href="https://doi.org/10.1088/1748-9326/acadf3" ext-link-type="DOI">10.1088/1748-9326/acadf3</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{Van~Horn et~al.(2018)Van~Horn, Mac~Aodha, Song, Cui, Sun, Shepard,
Adam, Perona, and Belongie}}?><label>Van Horn et al.(2018)Van Horn, Mac Aodha, Song, Cui, Sun, Shepard, Adam, Perona, and Belongie</label><?label van2018inaturalist?><mixed-citation> Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S.: The inaturalist species classification and detection dataset, in: Proceedings of the IEEE conference on computer vision and pattern recognition,  Salt Lake City, Utah, USA 18–22 June 2018, 8769–8778, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx50"><?xmltex \def\ref@label{{Van~Horn et~al.(2021)Van~Horn, Cole, Beery, Wilber, Belongie, and
Mac~Aodha}}?><label>Van Horn et al.(2021)Van Horn, Cole, Beery, Wilber, Belongie, and Mac Aodha</label><?label van2021benchmarking?><mixed-citation>Van Horn, G., Cole, E., Beery, S., Wilber, K., Belongie, S., and Mac Aodha, O.: Benchmarking Representation Learning for Natural World Image Collections, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, virtual, 19–25 June 2021, <ext-link xlink:href="https://doi.org/10.48550/arXiv.2103.16483" ext-link-type="DOI">10.48550/arXiv.2103.16483</ext-link>, 12884–12893, 2021. </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx51"><?xmltex \def\ref@label{{Wagner(2021)}}?><label>Wagner(2021)</label><?label wagner2021flowering?><mixed-citation>Wagner, F. H.: The flowering of Atlantic Forest Pleroma trees, Sci. Rep., 11, 1–20, <ext-link xlink:href="https://doi.org/10.1038/s41598-021-99304-x" ext-link-type="DOI">10.1038/s41598-021-99304-x</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx52"><?xmltex \def\ref@label{{Zhou(2018)}}?><label>Zhou(2018)</label><?label zhou2018brief?><mixed-citation>Zhou, Z.-H.: A brief introduction to weakly supervised learning, Natl. Sci. Rev., 5, 44–53, <ext-link xlink:href="https://doi.org/10.1093/nsr/nwx106" ext-link-type="DOI">10.1093/nsr/nwx106</ext-link>, 2018.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>From simple labels to semantic image segmentation: leveraging citizen science plant photographs for tree species mapping in drone imagery</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Affouard et al.(2017)Affouard, Goëau, Bonnet, Lombardo, and
Joly</label><mixed-citation>
      
Affouard, A., Goëau, H., Bonnet, P., Lombardo, J.-C., and Joly, A.:
Pl@ntnet app in the era of deep learning, in: ICLR: International Conference on
Learning Representations, April 2017, Toulon, France, ffhal-01629195f, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Bayraktar et al.(2020)Bayraktar, Basarkan, and
Celebi</label><mixed-citation>
      
Bayraktar, E., Basarkan, M. E., and Celebi, N.: A low-cost UAV framework
towards ornamental plant detection and counting in the wild, ISPRS J.
Photogramm., 167, 1–11,
<a href="https://doi.org/10.1016/j.isprsjprs.2020.06.012" target="_blank">https://doi.org/10.1016/j.isprsjprs.2020.06.012</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Boone and Basille(2019)</label><mixed-citation>
      
Boone, M. E. and Basille, M.: Using iNaturalist to contribute your nature
observations to science, EDIS, 2019, 5–5, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Bouguettaya et al.(2022)Bouguettaya, Zarzour, Kechida, and
Taberkit</label><mixed-citation>
      
Bouguettaya, A., Zarzour, H., Kechida, A., and Taberkit, A. M.: Deep learning
techniques to classify agricultural crops through UAV imagery: A review,
Neural Comput. Appl., 34, 9511–9536, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Braga et al.(2020)Braga, Peripato, Dalagnol, P. Ferreira, Tarabalka,
OC Aragão, F. de Campos Velho, Shiguemori, and Wagner</label><mixed-citation>
      
Braga, G., J. R., Peripato, V., Dalagnol, R., P. Ferreira, M., Tarabalka, Y.,
OC Aragão, L. E., F. de Campos Velho, H., Shiguemori, E. H., and Wagner,
F. H.: Tree crown delineation algorithm based on a convolutional neural
network, Remote Sens., 12, 1288, <a href="https://doi.org/10.3390/rs12081288" target="_blank">https://doi.org/10.3390/rs12081288</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Brandt et al.(2020)Brandt, Tucker, Kariryaa, Rasmussen, Abel, Small,
Chave, Rasmussen, Hiernaux, Diouf et al.</label><mixed-citation>
      
Brandt, M., Tucker, C. J., Kariryaa, A., Rasmussen, K., Abel, C., Small, J., Chave, J., Rasmussen, L. V., Hiernaux, P., Diouf, A. A., Kergoat, L., Mertz, O., Igel, C., Gieseke, F., Schöning, J., Li, S., Melocik, K., Meyer, J., Sinno, S., Romero, E., Glennie, E., Montagu, A., Dendoncker, M., and Fensholt, R.: An
unexpectedly large count of trees in the West African Sahara and Sahel,
Nature, 587, 78–82, <a href="https://doi.org/10.1038/s41586-020-2824-5" target="_blank">https://doi.org/10.1038/s41586-020-2824-5</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Brodrick et al.(2019)Brodrick, Davies, and
Asner</label><mixed-citation>
      
Brodrick, P. G., Davies, A. B., and Asner, G. P.: Uncovering ecological
patterns with convolutional neural networks, Trends Ecol. Evol.,
34, 734–745, <a href="https://doi.org/10.1016/j.tree.2019.03.006" target="_blank">https://doi.org/10.1016/j.tree.2019.03.006</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Chandler et al.(2017)Chandler, See, Copas, Bonde, López,
Danielsen, Legind, Masinde, Miller-Rushing, Newman
et al.</label><mixed-citation>
      
Chandler, M., See, L., Copas, K., Bonde, A. M., López, B. C., Danielsen, F., Legind, J. K., Masinde, S., Miller-Rushing, A. J., Newman, G., Rosemartin, A., and Turak, E.:
Contribution of citizen science towards international biodiversity
monitoring, Biol. Conserv., 213, 280–294,
<a href="https://doi.org/10.1016/j.biocon.2016.09.004" target="_blank">https://doi.org/10.1016/j.biocon.2016.09.004</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Cherif et al.(2023)Cherif, Feilhauer, Berger, Dao, Ewald, Hank, He,
Kovach, Lu, Townsend et al.</label><mixed-citation>
      
Cherif, E., Feilhauer, H., Berger, K., Dao, P. D., Ewald, M., Hank,
T. B., He, Y., Kovach, K. R., Lu, B., Townsend, P. A., and Kattenborn, T.: From spectra to plant
functional traits: Transferable multi-trait models from heterogeneous and
sparse data, Remote Sens. Environ., 292, 113580,
<a href="https://doi.org/10.1016/j.rse.2023.113580" target="_blank">https://doi.org/10.1016/j.rse.2023.113580</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Cloutier et al.(2023)Cloutier, Germain, and
Laliberté</label><mixed-citation>
      
Cloutier, M., Germain, M., and Laliberté, E.: Influence of Temperate Forest
Autumn Leaf Phenology on Segmentation of Tree Species from UAV Imagery Using
Deep Learning, bioRxiv,  2023–08, <a href="https://doi.org/10.1101/2023.08.03.548604" target="_blank">https://doi.org/10.1101/2023.08.03.548604</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Curnick et al.(2021)Curnick, Davies, Duncan, Freeman, Jacoby,
Shelley, Rossi, Wearn, Williamson, and Pettorelli</label><mixed-citation>
      
Curnick, D. J., Davies, A. J., Duncan, C., Freeman, R., Jacoby, D. M., Shelley,
H. T., Rossi, C., Wearn, O. R., Williamson, M. J., and Pettorelli, N.:
SmallSats: a new technological frontier in ecology and conservation?, Remote Sensing in Ecology and Conservation,
8,
139–150, <a href="https://doi.org/10.1002/rse2.239" target="_blank">https://doi.org/10.1002/rse2.239</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>De Sa et al.(2018)De Sa, Castro, Carvalho, Marchante,
López-Núñez, and Marchante</label><mixed-citation>
      
De Sa, N. C., Castro, P., Carvalho, S., Marchante, E., López-Núñez,
F. A., and Marchante, H.: Mapping the flowering of an invasive plant using
unmanned aerial vehicles: is there potential for biocontrol monitoring?,
Front. Plant Sci., 9, 293, <a href="https://doi.org/10.3389/fpls.2018.00293" target="_blank">https://doi.org/10.3389/fpls.2018.00293</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Di Cecco et al.(2021)Di Cecco, Barve, Belitz, Stucky, Guralnick, and
Hurlbert</label><mixed-citation>
      
Di Cecco, G. J., Barve, V., Belitz, M. W., Stucky, B. J., Guralnick, R. P., and
Hurlbert, A. H.: Observing the observers: How participants contribute data to
iNaturalist and implications for biodiversity science, BioScience, 71,
1179–1188, <a href="https://doi.org/10.1093/biosci/biab093" target="_blank">https://doi.org/10.1093/biosci/biab093</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Fassnacht et al.(2016)Fassnacht, Latifi, Stereńczak, Modzelewska,
Lefsky, Waser, Straub, and Ghosh</label><mixed-citation>
      
Fassnacht, F. E., Latifi, H., Stereńczak, K., Modzelewska, A., Lefsky, M.,
Waser, L. T., Straub, C., and Ghosh, A.: Review of studies on tree species
classification from remotely sensed data, Remote Sens. Environ., 186,
64–87, <a href="https://doi.org/10.1016/j.rse.2016.08.013" target="_blank">https://doi.org/10.1016/j.rse.2016.08.013</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Ferlian et al.(2018)Ferlian, Cesarz, Craven, Hines, Barry,
Bruelheide, Buscot, Haider, Heklau, Herrmann et al.</label><mixed-citation>
      
Ferlian, O., Cesarz, S., Craven, D., Hines, J., Barry, K. E., Bruelheide,
H., Buscot, F., Haider, S., Heklau, H., Herrmann, S., Kühn, P.,Pruschitzki, U., Schädler, M., Wagg, C., Weigelt, A., Wubet, T., and Eisenhauer, N.: Mycorrhiza in tree
diversity–ecosystem function relationships: conceptual framework and
experimental implementation, Ecosphere, 9, e02226, <a href="https://doi.org/10.1002/ecs2.2226" target="_blank">https://doi.org/10.1002/ecs2.2226</a>,
2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Fraisl et al.(2022)Fraisl, Hager, Bedessem, Gold, Hsing, Danielsen,
Hitchcock, Hulbert, Piera, Spiers et al.</label><mixed-citation>
      
Fraisl, D., Hager, G., Bedessem, B., Gold, M., Hsing, P.-Y.,
Danielsen, F., Hitchcock, C. B., Hulbert, J. M., Piera, J.,
Spiers, H., Thiel, M., and Haklay, M.: Citizen
science in environmental and ecological sciences,
Nature Reviews Methods Primers, 2, 64, <a href="https://doi.org/10.1038/s43586-022-00144-4" target="_blank">https://doi.org/10.1038/s43586-022-00144-4</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Fricker et al.(2019)Fricker, Ventura, Wolf, North, Davis, and
Franklin</label><mixed-citation>
      
Fricker, G. A., Ventura, J. D., Wolf, J. A., North, M. P., Davis, F. W., and
Franklin, J.: A convolutional neural network classifier identifies tree
species in mixed-conifer forest from hyperspectral imagery, Remote Sens.,
11, 2326, <a href="https://doi.org/10.3390/rs11192326" target="_blank">https://doi.org/10.3390/rs11192326</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Galuszynski et al.(2022)Galuszynski, Duker, Potts, and
Kattenborn</label><mixed-citation>
      
Galuszynski, N. C., Duker, R., Potts, A. J., and Kattenborn, T.: Automated
mapping of Portulacaria afra canopies for restoration monitoring with
convolutional neural networks and heterogeneous unmanned aerial vehicle
imagery, PeerJ, 10, e14219, <a href="https://doi.org/10.7717/peerj.14219" target="_blank">https://doi.org/10.7717/peerj.14219</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>GBIF(2019)</label><mixed-citation>
      
GBIF: GBIF: the global biodiversity information facility, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Hoeser and Kuenzer(2020)</label><mixed-citation>
      
Hoeser, T. and Kuenzer, C.: Object detection and image segmentation with deep
learning on earth observation data: A review-part i: Evolution and recent
trends, Remote Sens., 12, 1667, <a href="https://doi.org/10.3390/rs12101667" target="_blank">https://doi.org/10.3390/rs12101667</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>iNaturalist(2023)</label><mixed-citation>
      
iNaturalist: iNaturalist, <a href="https://www.inaturalist.org" target="_blank"/>, last access:
4 September 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Ivanova and Shashkov(2021)</label><mixed-citation>
      
Ivanova, N. and Shashkov, M.: The possibilities of GBIF data use in ecological
research, Russ. J. Ecol., 52, 1–8, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Johnston et al.(2023)Johnston, Matechou, and
Dennis</label><mixed-citation>
      
Johnston, A., Matechou, E., and Dennis, E. B.: Outstanding challenges and
future directions for biodiversity monitoring using citizen science data,
Methods in Ecol. Evol., 14, 103–116,
<a href="https://doi.org/10.1111/2041-210X.13834" target="_blank">https://doi.org/10.1111/2041-210X.13834</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Joly et al.(2016)Joly, Bonnet, Goëau, Barbe, Selmi, Champ,
Dufour-Kowalski, Affouard, Carré, Molino et al.</label><mixed-citation>
      
Joly, A., Bonnet, P., Goëau, H., Barbe, J., Selmi, S., Champ, J., Dufour-Kowalski, S., Affouard, A., Carré, J., Molino, J. F., Boujemaa, N., and Barthélémy, D.: A
look inside the Pl@ ntNet experience: The good, the bias and the hope,
Multimedia Syst., 22, 751–766, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Kattenborn and Soltani(2023)</label><mixed-citation>
      
Kattenborn, T. and Soltani, S.: CrowdVision2TreeSegment, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.10019552" target="_blank">https://doi.org/10.5281/zenodo.10019552</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Kattenborn et al.(2019a)Kattenborn, Eichel, and
Fassnacht</label><mixed-citation>
      
Kattenborn, T., Eichel, J., and Fassnacht, F. E.: Convolutional Neural Networks
enable efficient, accurate and fine-grained segmentation of plant species and
communities from high-resolution UAV imagery, Sci. Rep., 9, 1–9,
<a href="https://doi.org/10.1038/s41598-019-53797-9" target="_blank">https://doi.org/10.1038/s41598-019-53797-9</a>, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Kattenborn et al.(2019b)Kattenborn, Lopatin,
Förster, Braun, and Fassnacht</label><mixed-citation>
      
Kattenborn, T., Lopatin, J., Förster, M., Braun, A. C., and Fassnacht,
F. E.: UAV data as alternative to field sampling to map woody invasive
species based on combined Sentinel-1 and Sentinel-2 data, Remote Sens.
Environ., 227, 61–73, <a href="https://doi.org/10.1016/j.rse.2019.03.025" target="_blank">https://doi.org/10.1016/j.rse.2019.03.025</a>,
2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Kattenborn et al.(2021)Kattenborn, Leitloff, Schiefer, and
Hinz</label><mixed-citation>
      
Kattenborn, T., Leitloff, J., Schiefer, F., and Hinz, S.: Review on
Convolutional Neural Networks (CNN) in vegetation remote sensing, ISPRS
J. Photogramm., 173, 24–49,
<a href="https://doi.org/10.1016/j.isprsjprs.2020.12.010" target="_blank">https://doi.org/10.1016/j.isprsjprs.2020.12.010</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Kattenborn et al.(2022)Kattenborn, Schiefer, Frey, Feilhauer,
Mahecha, and Dormann</label><mixed-citation>
      
Kattenborn, T., Schiefer, F., Frey, J., Feilhauer, H., Mahecha, M. D., and
Dormann, C. F.: Spatially autocorrelated training and validation samples
inflate performance assessment of convolutional neural networks,
ISPRS Open Journal of Photogrammetry and Remote Sensing, 5, 100018,
<a href="https://doi.org/10.1016/j.ophoto.2022.100018" target="_blank">https://doi.org/10.1016/j.ophoto.2022.100018</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Leitão et al.(2018)Leitão, Schwieder, Pötzschner, Pinto,
Teixeira, Pedroni, Sanchez, Rogass, van der Linden, Bustamante
et al.</label><mixed-citation>
      
Leitão, P. J., Schwieder, M., Pötzschner, F., Pinto, J. R., Teixeira, A. M., Pedroni, F., Sanchez, M., Rogass, C., van der Linden, S., Bustamante, M. M. and Hostert, P.: From sample to pixel: multi-scale remote sensing data for
upscaling aboveground carbon data in heterogeneous landscapes, Ecosphere, 9,
e02298, <a href="https://doi.org/10.1002/ecs2.2298" target="_blank">https://doi.org/10.1002/ecs2.2298</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Lopatin et al.(2017)Lopatin, Fassnacht, Kattenborn, and
Schmidtlein</label><mixed-citation>
      
Lopatin, J., Fassnacht, F. E., Kattenborn, T., and Schmidtlein, S.: Mapping
plant species in mixed grassland communities using close range imaging
spectroscopy, Remote Sens. Environ., 201, 12–23,
<a href="https://doi.org/10.1016/j.rse.2017.08.031" target="_blank">https://doi.org/10.1016/j.rse.2017.08.031</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Lopatin et al.(2019)Lopatin, Dolos, Kattenborn, and
Fassnacht</label><mixed-citation>
      
Lopatin, J., Dolos, K., Kattenborn, T., and Fassnacht, F. E.: How canopy shadow
affects invasive plant species classification in high spatial resolution
remote sensing, Remote Sens. Ecol. Conserv., 5, 302–317,
<a href="https://doi.org/10.1002/rse2.109" target="_blank">https://doi.org/10.1002/rse2.109</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Ma et al.(2019)Ma, Liu, Zhang, Ye, Yin, and Johnson</label><mixed-citation>
      
Ma, L., Liu, Y., Zhang, X., Ye, Y., Yin, G., and Johnson, B. A.: Deep learning
in remote sensing applications: A meta-analysis and review, ISPRS J.
Photogramm., 152, 166–177,
<a href="https://doi.org/10.1016/j.isprsjprs.2019.04.015" target="_blank">https://doi.org/10.1016/j.isprsjprs.2019.04.015</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Mäder et al.(2021)Mäder, Boho, Rzanny, Seeland, Wittich,
Deggelmann, and Wäldchen</label><mixed-citation>
      
Mäder, P., Boho, D., Rzanny, M., Seeland, M., Wittich, H. C., Deggelmann,
A., and Wäldchen, J.: The flora incognita app–interactive plant species
identification, Methods in Ecol. Evol., 12,   1335–1342,
<a href="https://doi.org/10.1111/2041-210X.13611" target="_blank">https://doi.org/10.1111/2041-210X.13611</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Maes and Steppe(2019)</label><mixed-citation>
      
Maes, W. H. and Steppe, K.: Perspectives for remote sensing with unmanned
aerial vehicles in precision agriculture, Trends Plant Sci., 24,
152–164, <a href="https://doi.org/10.1016/j.tplants.2018.11.007" target="_blank">https://doi.org/10.1016/j.tplants.2018.11.007</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Milas et al.(2017)Milas, Arend, Mayer, Simonson, and
Mackey</label><mixed-citation>
      
Milas, A. S., Arend, K., Mayer, C., Simonson, M. A., and Mackey, S.: Different
colours of shadows: Classification of UAV images, Int. J.
Remote Sens., 38, 3084–3100, <a href="https://doi.org/10.1080/01431161.2016.1274449" target="_blank">https://doi.org/10.1080/01431161.2016.1274449</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Molls(2021)</label><mixed-citation>
      
Molls, C.: The Obs-Services and their potentials for biodiversity data
assessments with a test of the current reliability of photo-identification of
Coleoptera in the field, Tijdschrift voor Entomologie, 164, 143–153, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Müllerová et al.(2023)Müllerová, Brundu,
Große-Stoltenberg, Kattenborn, and Richardson</label><mixed-citation>
      
Müllerová, J., Brundu, G., Große-Stoltenberg, A., Kattenborn, T.,
and Richardson, D. M.: Pattern to process, research to practice: remote
sensing of plant invasions, Biol. Invasions,  25,  3651–3676, <a href="https://doi.org/10.1007/s10530-023-03150-z" target="_blank">https://doi.org/10.1007/s10530-023-03150-z</a>,
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Ronneberger et al.(2015)Ronneberger, Fischer, and
Brox</label><mixed-citation>
      
Ronneberger, O., Fischer, P., and Brox, T.: U-net: Convolutional networks for
biomedical image segmentation, in: International Conference on Medical image
computing and computer-assisted intervention, Springer, 234–241,
<a href="https://doi.org/10.1007/978-3-319-24574-4_28" target="_blank">https://doi.org/10.1007/978-3-319-24574-4_28</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Rzanny et al.(2019)Rzanny, Mäder, Deggelmann, Chen, and
Wäldchen</label><mixed-citation>
      
Rzanny, M., Mäder, P., Deggelmann, A., Chen, M., and Wäldchen, J.:
Flowers, leaves or both? How to obtain suitable images for automated plant
identification, Plant Methods, 15, 1–11, <a href="https://doi.org/10.1186/s13007-019-0462-4" target="_blank">https://doi.org/10.1186/s13007-019-0462-4</a>,
2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Schiefer et al.(2020)Schiefer, Kattenborn, Frick, Frey, Schall, Koch,
and Schmidtlein</label><mixed-citation>
      
Schiefer, F., Kattenborn, T., Frick, A., Frey, J., Schall, P., Koch, B., and
Schmidtlein, S.: Mapping forest tree species in high resolution UAV-based
RGB-imagery by means of convolutional neural networks, ISPRS J.
Photogramm., 170, 205–215,
<a href="https://doi.org/10.1016/j.isprsjprs.2020.10.015" target="_blank">https://doi.org/10.1016/j.isprsjprs.2020.10.015</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Schiefer et al.(2023)Schiefer, Schmidtlein, Frick, Frey, Klinke,
Zielewska-Büttner, Junttila, Uhl, and Kattenborn</label><mixed-citation>
      
Schiefer, F., Schmidtlein, S., Frick, A., Frey, J., Klinke, R.,
Zielewska-Büttner, K., Junttila, S., Uhl, A., and Kattenborn, T.:
UAV-based reference data for the prediction of fractional cover of standing
deadwood from Sentinel time series, ISPRS Open Journal of Photogrammetry and
Remote Sensing, 8, 100034, <a href="https://doi.org/10.1016/j.ophoto.2023.100034" target="_blank">https://doi.org/10.1016/j.ophoto.2023.100034</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Schiller et al.(2021)Schiller, Schmidtlein, Boonman,
Moreno-Martínez, and Kattenborn</label><mixed-citation>
      
Schiller, C., Schmidtlein, S., Boonman, C., Moreno-Martínez, A., and
Kattenborn, T.: Deep learning and citizen science enable automated plant
trait predictions from photographs, Sci. Rep., 11, 1–12, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Schmitt et al.(2020)Schmitt, Prexl, Ebel, Liebel, and
Zhu</label><mixed-citation>
      
Schmitt, M., Prexl, J., Ebel, P., Liebel, L., and Zhu, X. X.: Weakly supervised
semantic segmentation of satellite images for land cover mapping–challenges
and opportunities, arXiv [preprint],
<a href="https://doi.org/10.48550/arXiv.2002.08254" target="_blank">https://doi.org/10.48550/arXiv.2002.08254</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Soltani et al.(2022)Soltani, Feilhauer, Duker, and
Kattenborn</label><mixed-citation>
      
Soltani, S., Feilhauer, H., Duker, R., and Kattenborn, T.: Transfer learning
from citizen science photographs enables plant species identification in UAVs
imagery, ISPRS Open Journal of Photogrammetry and Remote Sensing, 5, 100016,
<a href="https://doi.org/10.1016/j.ophoto.2022.100016" target="_blank">https://doi.org/10.1016/j.ophoto.2022.100016</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Sun et al.(2021)Sun, Wang, Wang, Yang, Xie, and Huang</label><mixed-citation>
      
Sun, Z., Wang, X., Wang, Z., Yang, L., Xie, Y., and Huang, Y.: UAVs as remote
sensing platforms in plant ecology: review of applications and challenges,
J. Plant Ecol., 14, 1003–1023, <a href="https://doi.org/10.1093/jpe/rtab089" target="_blank">https://doi.org/10.1093/jpe/rtab089</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Tan and Le(2019)</label><mixed-citation>
      
Tan, M. and Le, Q.: Efficientnet: Rethinking model scaling for convolutional
neural networks, in: International conference on machine learning,
6105–6114, PMLR, Long Beach, California,
10–15 June 2019,
<a href="https://doi.org/10.48550/arXiv.1905.11946" target="_blank">https://doi.org/10.48550/arXiv.1905.11946</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>van Der Velde et al.(2023)van Der Velde, Goëau, Bonnet,
d'Andrimont, Yordanov, Affouard, Claverie, Czúcz, Elvekjær,
Martinez-Sanchez et al.</label><mixed-citation>
      
van Der Velde, M., Goëau, H., Bonnet, P., d’Andrimont, R., Yordanov, M., Affouard, A., Claverie, M., Czúcz, B., Elvekjær, N., Martinez-Sanchez, L., and Rotllan-Puig, X.: Pl@ ntNet Crops: merging citizen science
observations and structured survey data to improve crop recognition for
agri-food-environment applications, Environ. Res. Lett., 18,
025005, <a href="https://doi.org/10.1088/1748-9326/acadf3" target="_blank">https://doi.org/10.1088/1748-9326/acadf3</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Van Horn et al.(2018)Van Horn, Mac Aodha, Song, Cui, Sun, Shepard,
Adam, Perona, and Belongie</label><mixed-citation>
      
Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H.,
Perona, P., and Belongie, S.: The inaturalist species classification and
detection dataset, in: Proceedings of the IEEE conference on computer vision
and pattern recognition,  Salt Lake City, Utah, USA
18–22 June 2018, 8769–8778, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Van Horn et al.(2021)Van Horn, Cole, Beery, Wilber, Belongie, and
Mac Aodha</label><mixed-citation>
      
Van Horn, G., Cole, E., Beery, S., Wilber, K., Belongie, S., and Mac Aodha, O.:
Benchmarking Representation Learning for Natural World Image Collections, in:
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, virtual, 19–25 June 2021,
<a href="https://doi.org/10.48550/arXiv.2103.16483" target="_blank">https://doi.org/10.48550/arXiv.2103.16483</a>, 12884–12893, 2021.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Wagner(2021)</label><mixed-citation>
      
Wagner, F. H.: The flowering of Atlantic Forest Pleroma trees, Sci.
Rep., 11, 1–20, <a href="https://doi.org/10.1038/s41598-021-99304-x" target="_blank">https://doi.org/10.1038/s41598-021-99304-x</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Zhou(2018)</label><mixed-citation>
      
Zhou, Z.-H.: A brief introduction to weakly supervised learning,
Natl. Sci. Rev., 5, 44–53, <a href="https://doi.org/10.1093/nsr/nwx106" target="_blank">https://doi.org/10.1093/nsr/nwx106</a>, 2018.

    </mixed-citation></ref-html>--></article>
