For 16S rRNA amplicon sequencing data, you should submit your raw data to the Sequence Read Archive (SRA), organized under a BioProject and linked to individual BioSamples.
You should not submit raw 16S sequencing data primarily to GEO (Gene Expression Omnibus).
Here is a breakdown of why, how the NCBI submission hierarchy works, and how to correctly link your data if processed files are also required by the journal.
1. Why SRA and not GEO for Raw Data?
- SRA (Sequence Read Archive) is the official NCBI repository for raw high-throughput sequencing reads, including 16S rRNA amplicon data (FASTQ files). NCBI guidelines and most scientific journals explicitly require raw microbiome sequencing data to be deposited here.
- GEO is designed primarily for functional genomics data (e.g., RNA-Seq, microarrays, ChIP-Seq, ATAC-Seq) where the focus is on gene expression or epigenetic profiles. While researchers occasionally upload processed microbiome data (like an OTU/ASV count table) to GEO as a supplementary dataset, the raw sequencing files belong in SRA.
2. The Correct NCBI Submission Hierarchy
When you submit, you will build the submission in this exact order:
-
BioProject (The Umbrella)
- What it is: The overarching description of your entire study.
- What you provide: Project title (e.g., “Sex-specific gut microbiota and IL-17A response in aged mice after experimental stroke”), study type (Metagenomics or Amplicon), and a brief abstract. This generates a PRJNAxxxxxx accession number.
-
BioSample (The Biological Source)
- What it is: A record for each individual biological sample you are submitting (e.g., Sample A1, Sample C3, Sample J10).
- What you provide: Metadata describing the mouse (e.g., organism: Mus musculus, age: 14-16 months, sex: male/female, tissue: feces, treatment group: post-stroke, pre-FMT, etc.). This generates a SAMNxxxxxx accession number for each sample.
-
SRA (The Sequencing Data)
- What it is: The actual raw data files and sequencing metadata.
- What you provide: You will upload your demultiplexed FASTQ files (or a tarball of them) and link each file to its corresponding BioSample. You will also specify the sequencing platform (e.g., Illumina MiSeq), library strategy (AMPLICON), and target gene (16S rRNA). This generates an SRRxxxxxx (or ERR/DRR) accession number for each run.
3. The Two-Step Strategy: Linking SRA Raw Data to GEO Processed Data
If the target journal require you to also submit processed data (e.g., the final OTU/ASV abundance table used to generate Figures 4 and 5) to GEO, NCBI provides a streamlined workflow to avoid duplicate uploads.
GEO submissions require a metadata spreadsheet. Metadata refers to descriptive information about the overall study, individual samples, all protocols, and references to processed and raw data file names. Information is supplied by completing all fields of a metadata template spreadsheet (guidelines are provided within the file).
š” Important: Provide enough details so that users can get a general understanding of the study and samples from the GEO records. Please spell out all acronyms and abbreviations. Submit a separate metadata spreadsheet for each data type.
Have you already submitted raw data to SRA and now want to submit to GEO?
If you already have your raw data in SRA, you do not need to submit it again to GEO. NCBI only needs the processed data and a specialized metadata file in order to create GEO records and link them to your raw data records previously submitted to SRA.
To do this, you must choose the second option in the GEO submission portal: š “Download metadata spreadsheet with SRA accessions”
- What you need to do: You will need to enter the PRJNA, SAMN, and SRX or SRR accession numbers for all samples with raw data already submitted to SRA.
- Where to find this: You can get this information for your SUB ID on the NCBI Submission Portal page after your SRA submission is initiated.
4. Actionable Next Steps & Summary Workflow
- Finalize the sample list with your co-authors (confirming Groups 1ā11 and excluding Groups 12ā14 and the 2022 legacy data).
- Prepare a Metadata Spreadsheet: Use the NCBI template. You will need one row per sample (e.g., A1, A2… B1, B2…) with columns for:
sample_name,bioSample_model(mouse),sex,age,tissue(feces),collection_date, andtreatment(e.g., “post-stroke day 3”, “pre-FMT baseline”). - Step 1: Submit to SRA First: Go to the NCBI SRA Submission Wizard (https://submit.ncbi.nlm.nih.gov/). Create the BioProject, batch-upload the BioSample metadata, and upload the raw FASTQ files. Save your generated PRJNA, SAMN, and SRR numbers.
- Step 2: Submit Processed Data to GEO: Go to the GEO submission portal. Choose the option to “Download metadata spreadsheet with SRA accessions”. Fill it out with your processed data file names and the SRA accession numbers you just generated. Upload this to link everything together.
Pro Tip: If the journal is flexible, processed data can often just be included as a Supplementary File (e.g., a .csv or .xlsx file) with the manuscript, or deposited in a repository like Figshare or Zenodo. However, if GEO is explicitly requested, the two-step SRA-first workflow above is the correct and most efficient method.
TODOs: drafting the BioProject abstract and formatting the BioSample metadata spreadsheet!
I am organizing the 16S rRNA sequencing data for deposition in the NCBI Sequence Read Archive (SRA).
Because only a subset of the sequenced samples was ultimately used in the final manuscript figures, I have compiled an exact list of the specific samples to be submitted. I want to ensure our public dataset perfectly aligns with the content of the manuscript.
Here is the exact list of samples I will upload to NCBI, mapped to the manuscript figures. Could you please confirm that they are correct?
1. Stroke Model Groups (Used in Fig. 4DāF for blood/brain SCFA)
- Group 1 (Aged ā, Post-stroke): Submitting A1āA11
- Group 2 (Aged ā, Post-stroke): Submitting B1āB16
2. Baseline Donor Groups (Used in Fig. 4AāC, Supp. Fig. 4, Fig. 5C)
- Group 3 (Aged ā FMT Donor): Submitting C1āC6 (n=6).
- Excluded: C7āC10 (C8āC9 excluded due to age; C10 excluded as an outlier/low sequencing depth).
- Group 4 (Aged ā FMT Donor): Submitting E1āE8 (n=8).
- Excluded: E9āE10 (low sequencing depth/outliers).
- Group 5 (Young ā Control Donor): Submitting F1āF5 (Control, not shown in main figures).
3. Pre-FMT Baseline Groups (Used in Fig. 5B as purple dots, n=18 total)
- Group 6 (Aged ā, Pre-FMT): Submitting G1āG6
- Group 7 (Aged ā, Pre-FMT): Submitting H1āH6
- Group 8 (Young ā, Pre-FMT): Submitting I1āI6
4. FMT Recipient Groups (Used in Fig. 5B, 5C, 5D, 5E)
- Group 9 (Young ā recipient, Aged ā donor): Submitting J1āJ4, J10, J11 (n=6).
- Excluded: J5āJ9 (insufficient sequencing depth/QC exclusion).
- Group 10 (Young ā recipient, Aged ā donor): Submitting K1āK6 (n=6).
- Excluded: K7āK15 (insufficient sequencing depth/QC exclusion).
- Group 11 (Young ā recipient, Young ā donor): Submitting L2āL6 (n=5).
- Excluded: L1, L7āL15 (insufficient sequencing depth/QC exclusion).
Points for Your Final Confirmation:
- Exclusion of Groups 12, 13, and 14 (Samples M, N, O), as well as Group 5, since they are not shown in the final manuscript.
- Exclusion of the 2022 Legacy Dataset: I also identified an older sequencing run from 2022 (containing samples labeled Group 1 to Group 8, covering f.aged, f.young, m.aged, and m.young in pre/post-stroke conditions). Since this dataset is from an earlier pilot phase and is not referenced or utilized in the current manuscript, I assume we should NOT submit this 2022 data as part of this paper’s NCBI submission. Could you please confirm if this is correct?