Articles | Volume 23, issue 18
https://doi.org/10.5194/bg-23-6725-2026
© Author(s) 2026. This work is distributed under the Creative Commons Attribution 4.0 License.
Exploratory characterization of bacterial communities and predicted functional profiles in six water samples from four Colombian Andean lakes using 16S rRNA gene amplicon sequencing
Download
- Final revised paper (published on 24 Sep 2026)
- Supplement to the final revised paper
- Preprint (discussion started on 21 May 2026)
- Supplement to the preprint
Interactive discussion
Status: closed
Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor
| : Report abuse
-
RC1: 'Comment on egusphere-2026-2137', Anonymous Referee #1, 29 Jun 2026
- AC1: 'Reply on RC1', Andres Gomez-Palacio, 28 Jul 2026
-
RC2: 'Comment on egusphere-2026-2137', Anonymous Referee #2, 02 Jul 2026
- AC2: 'Reply on RC2', Andres Gomez-Palacio, 28 Jul 2026
-
RC3: 'Comment on egusphere-2026-2137', Anonymous Referee #3, 14 Jul 2026
- AC3: 'Reply on RC3', Andres Gomez-Palacio, 28 Jul 2026
Peer review completion
AR – Author's response | RR – Referee report | ED – Editor decision | EF – Editorial file upload
ED: Submit a revised manuscript (31 Jul 2026) by Pierre Amato
AR by Andres Gomez-Palacio on behalf of the Authors (31 Jul 2026)
Author's response
Manuscript
EF by Katja Gänger (05 Aug 2026)
Author's tracked changes
EF by Katja Gänger (05 Aug 2026)
Author's tracked changes
ED: Referee Nomination & Report Request started (17 Aug 2026) by Pierre Amato
RR by Anonymous Referee #1 (17 Aug 2026)
RR by Beatriz Sánchez-Parra (08 Sep 2026)
ED: Publish subject to technical corrections (10 Sep 2026) by Pierre Amato
AR by Andres Gomez-Palacio on behalf of the Authors (10 Sep 2026)
Author's response
Manuscript
The authors used 16S rRNA gene amplicon sequencing combined with PICRUSt2 functional prediction to systematically analyze the bacterial community structure, co-occurrence patterns, and potential metabolic functions of four high-altitude lakes (F ú quene, Tota, Calderona, Colorado) with different conservation states in the eastern Andes Mountains of Colombia. Research has found that lakes with greater human influence, such as F ú quene, exhibit higher microbial richness and diversity, while lakes with better protection, such as Calderona and Colorado, have more stable communities but lower diversity; Spatial heterogeneity (such as inflow rivers and coastal zones) and vertical stratification further shape community composition. Functional prediction reveals potential differences in bacterial communities in pathways such as carbohydrate metabolism and exogenous substance degradation under different protection states. This study provides baseline data for understanding the response of microorganisms in Andean mountain lakes to anthropogenic stress, and has certain ecological and conservation biological value.
However, despite the practical significance of the research topic, there are several significant deficiencies in the rigor of experimental design, data analysis, and generalizability of the conclusions. Taking all factors into consideration, I believe that the manuscript is currently not sufficient for publication in this journal. I suggest making significant revisions and reconsidering it.
The main issue with this manuscript is the severe lack of sample size and sampling design, which greatly undermines the reliability of all statistical inferences and ecological conclusions. In addition, the explanation of functional prediction results is too bold and lacks necessary cautious wording.
1.Rows 91-130, Table 1: The sample size is extremely small and the statistical power is seriously insufficient. This study only involved 4 lakes, and the number of sampling points for each lake was extremely small (2 samples from F ú quene, 2 samples from Tota, and 1 sample each from Calderona and Colorado). Such a small sample size cannot support any meaningful inter group statistical tests (such as comparisons between different conservation states), nor can it distinguish individual differences in lakes from conservation state effects. The so-called 'significant differences' (such as lines 237-241) are likely driven by a single outlier.
Suggest the author to significantly increase the sampling intensity. Each lake should have multiple sampling stations (at least 3-5) and repeat sampling in different seasons. If it cannot be achieved, this study should be honestly positioned as a "preliminary exploratory case study" and avoid making generalizations in the title, abstract, and conclusion.
2.Rows 116-120, Table 1: Pseudo repetition and time confusion. Two samples from F ú quene come from two different periods in 2019, one sample from Tota comes from 2018 and the other from 2019, while Calderona and Colorado only have one sample from 2019. This uneven and mixed time factor makes it difficult to distinguish the impact of "conservation status", "lake characteristics", and "sampling time" on community variation. For example, the high diversity of F ú quene may be due to its two samplings capturing seasonal fluctuations, rather than its low conservation state.
It is recommended that all lakes be sampled at the same time. If historical data (such as Tota's 2018 data) must be used, it should be analyzed separately as a control, rather than mixed with other 2019 data for comparison.
3.Lines 28-29, 85-88, 197-208, 354-370, 467-50: Overinterpretation of functional prediction (PICRUSt2) results. Although the author pointed out the limitations of PICRUSt2 in the methods section (lines 85-88), deterministic language such as "predicted metabolic potential variables" and "pathways related to... were identified" is frequently used in the results and discussion section. Especially when directly associating KEGG pathway annotations such as "infectious diseases" with potential pathogenic risks in environmental samples (lines 368-370), it can easily lead to misunderstandings. These are only computer predictions based on genome homology, not evidence.
Suggest revising this type of expression throughout the entire text. The word 'identified' should be changed to 'predicted to be present based on 16S rRNA gene induced metagenomics'; Change 'variation in... pathways' to' variation in the predicted abundance of genes associated with... pathways'. For the "infectious diseases" pathway, it must be clearly stated that this is only an artificial product annotated in the database and does not represent any pathogenic activity.
4.Lines 189-196, 319-338: Species level identification is unreliable. The author constructed a phylogenetic tree based on the V3-V4 16S rRNA gene fragment (approximately 460bp) and conducted species level identification of Mycobacteria. As is well known, the resolution of 16S rRNA genes at the species level is limited, especially for highly similar NTM groups. The author himself acknowledges this (lines 326-328), but still provides a detailed species level discussion below (lines 444-456).
Suggest either deleting all species level identification and discussions based on this fragment, and keeping the analysis at the genus level; Alternatively, it should be clarified that these are only "tentative species level assignments" and it is recommended to confirm them through whole genome sequencing or specific marker genes in the future.
5.The article discusses multiple driving factors such as nutrients, oxygen, and temperature, but lacks environmental physicochemical data.
During the discussion, it was repeatedly mentioned that factors such as "nutrient inputs," "temperature," and "oxygen gradients" shape microbial communities, but the entire text did not provide any synchronously measured environmental variable data (such as total phosphorus, total nitrogen, dissolved oxygen, chlorophyll-a, pH, etc.). This makes all inferences about environmental drivers unfounded. This is a fatal flaw.
It is recommended to supplement at least basic on-site hydrochemical parameters. If the data has been lost, it should be explicitly stated in the discussion that environmental microbiological correlation analysis cannot be conducted and the relevant inferences should be significantly weakened.
6.Lines 175-188, 286-296: The limitations of co-occurrence network analysis have not been fully discussed. The network is constructed based on only 6 samples (N=6), with 20 nodes and 226 edges. Calculating correlation on such extremely sparse datasets results in extremely unstable results, almost entirely driven by individual samples. Although the author mentioned in lines 186-188 and 526-527 that statistical associations do not equal ecological interactions, they did not emphasize the fundamental issues brought about by sample size.
It is not recommended to conduct network analysis under the current sample size. Alternatively, the author should explicitly downgrade this section to 'exploratory visualization' and warn readers that the results cannot be extrapolated.
The statement "Temporary stability" in lines 297-306 is not valid.
Only two time points of data (F ú quene and Tota twice each) claim 'temporary stability in dominant taxa'. Two time points can only describe changes and cannot prove stability.
Suggest deleting the word 'stability' and replacing it with 'temporary variation between two sampling events'.
7.It is recommended to conduct detailed language polishing and formatting proofreading throughout the entire text.