Tier: D (cMD paper) · Type: composite · Category: Statistical Analysis
What
Where a disease is represented by several datasets, pool within the disease first; then pool the resulting per-disease summaries across diseases in a second stage — keeping the two levels of between-study variance distinct rather than throwing every dataset into one flat pool.
Why it matters
Flattening would let the four multi-dataset diseases (colorectal cancer, type 2 diabetes, Crohn's, ulcerative colitis) dominate a 15-disease synthesis. The hierarchy is the correct handling and it is a genuinely non-obvious design that deserves to be written down separately from the pooling arithmetic.
Source material
waldronlab/curatedMetagenomicDataAnalyses — python_tools/hierarchical_metaanalysis.py, python_tools/hierarchical_binary_test.py, cMD3_paper_analyses/all_command_lines_py.sh (SECTION 2: hierarchical meta-analysis on 12 diseases)
- Paper: 10.1038/s41467-025-66888-1
Scope
In: the rule for which diseases enter stage one; how a single-dataset disease is carried into stage two; propagation of the stage-one standard errors; whether heterogeneity is re-estimated at stage two; multiple testing applied at which stage.
Out: the pooling arithmetic itself (component protocol).
Frontmatter starting point
type: "composite"
category: "Statistical Analysis"
citation: "10.1038/s41467-025-66888-1"
protocols_used:
- name: "per-dataset-smd-covariate-adjusted"
- name: "random-effects-meta-analysis-pm"
Acceptance criteria
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue in waldronlab/agent-protocol-standard — the standard may need a way to express
"classical method, no primary source".
- If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Tier: D (cMD paper) · Type:
composite· Category: Statistical AnalysisWhat
Where a disease is represented by several datasets, pool within the disease first; then pool the resulting per-disease summaries across diseases in a second stage — keeping the two levels of between-study variance distinct rather than throwing every dataset into one flat pool.
Why it matters
Flattening would let the four multi-dataset diseases (colorectal cancer, type 2 diabetes, Crohn's, ulcerative colitis) dominate a 15-disease synthesis. The hierarchy is the correct handling and it is a genuinely non-obvious design that deserves to be written down separately from the pooling arithmetic.
Source material
waldronlab/curatedMetagenomicDataAnalyses—python_tools/hierarchical_metaanalysis.py,python_tools/hierarchical_binary_test.py,cMD3_paper_analyses/all_command_lines_py.sh(SECTION 2: hierarchical meta-analysis on 12 diseases)Scope
In: the rule for which diseases enter stage one; how a single-dataset disease is carried into stage two; propagation of the stage-one standard errors; whether heterogeneity is re-estimated at stage two; multiple testing applied at which stage.
Out: the pooling arithmetic itself (component protocol).
Frontmatter starting point
Acceptance criteria
Cite the method's origin, not its users
PROTOCOL_STANDARD.mdis explicit: an atomic protocol carries "strictly 1 citation... correspondingto the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue in
waldronlab/agent-protocol-standard— the standard may need a way to express"classical method, no primary source".
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read
CONTRIBUTING.mdandPROTOCOL_STANDARD.md.The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-varianceprotocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR: