This pipeline runs multiple proteomics search engines on ProteoBench benchmark datasets and collects output files for downstream submission to ProteoBench. It supports DIA-NN, AlphaDIA, Sage, FragPipe, MaxQuant, and MetaMorpheus across DDA and DIA acquisition modes.
Everything runs through Nextflow (proteobench.nf), on a local machine or on a cluster (SLURM, …): it checks its own docker setup and runs the setup wizard itself when needed. Pull a tagged release directly from GitHub — no git clone needed — and nextflow run ProteoBench/ProteoRunners -r v1.0.5 is the only command most users ever have to type. Add --setup to that command to add or re-enable a tool later.
Docker is required. Every search engine (DIA-NN, AlphaDIA, Sage, FragPipe, MaxQuant, MetaMorpheus) now runs from a docker image, there is nothing left to compile or install natively for any tool.
| Dependency | Required for | Install |
|---|---|---|
| Docker | every tool — mandatory | docs.docker.com/get-docker |
Python 3.11+ with pyyaml, available as python3 |
Config checks and job enumeration used internally by the pipeline (rich is only needed for the standalone run_proteobench.py) |
conda env create -f environment.yml then conda activate proteobench-pipeline, or pip install -r requirements.txt in a virtual environment |
| Nextflow 23.10+ | Running the pipeline and setup wizard | curl -s https://get.nextflow.io | bash then move to a directory on $PATH |
curl, tar, unzip |
Downloading benchmark datasets | Usually already installed (sudo apt install curl tar unzip) |
git |
Building the DIA-NN 2.x image only | Usually already installed; see git-scm.com |
Check that each dependency is available:
docker info # should print server info, not a connection error
python3 -c "import yaml; print('pyyaml ok')" # run this in the environment you start Nextflow from
nextflow -version # needed for setup.nf and the Nextflow runnerYour user must be able to run docker without sudo (on Linux: sudo usermod -aG docker $USER, then log out/in). The pipeline checks Docker and Python before it does anything else and says what to fix if one of them is missing.
Make a folder for your benchmark, go into it, and run the latest release directly from GitHub:
mkdir my_benchmark && cd my_benchmark
nextflow run ProteoBench/ProteoRunners -r v1.0.5Everything you own is kept in the folder you run the command from:
| What | Default location | Change with |
|---|---|---|
config.yaml |
./config.yaml |
--config |
| Downloaded datasets | data/ next to config.yaml |
--data_dir |
| FragPipe licensed JARs | tools/fragpipe_jars/ next to config.yaml |
(asked by the wizard) |
Results and run_summary_nf.tsv |
results/ next to config.yaml (asked by the wizard, stored as global.output_dir) |
global.output_dir, --publish_dir |
Nextflow caches the pipeline code itself under ~/.nextflow/assets/ProteoBench/ProteoRunners; nothing of yours is written there. (Older versions defaulted config.yaml to that cache. If there is no config.yaml in your folder but there is one next to the pipeline code, that one is still used.) Relative paths in config.yaml are relative to the folder that holds config.yaml.
Developing the pipeline itself? Clone the repo and run it from inside your checkout instead — then these defaults resolve to the repo root, matching every other example in this README:
git clone https://github.com/ProteoBench/ProteoRunners.git
cd ProteoRunners
nextflow run proteobench.nfThat's it — the pipeline sets itself up on the way in:
- No
config.yamlyet? It runs the interactive setup wizard first: for each tool, a yes/no prompt to pull its image, then the dataset download, then where the results should go. It writesconfig.yamland asks whether to start the runs now. Press Enter to accept the default answer (the capital letter in[Y/n]). config.yamlexists and looks complete (every enabled tool's docker image — and, for FragPipe, its licensed JARs — is present, and every dataset an enabled tool uses is on disk)? Setup is skipped entirely.config.yamlexists but something's missing (an image was removed, a FragPipe JAR went missing, or a dataset that an enabled tool uses is not on disk)? Setup runs again, but only for those problems:- A missing image is pulled again with the exact tag in
config.yaml, so a pinned version stays pinned. Nothing else in that tool's entry changes. - If you decline, or the pull fails, setup offers to set that version to
enabled: false, so the pipeline stops asking. Only versions withenabled: trueare checked. - A missing dataset is offered for download. If you do not download it, setup offers to comment it out of the tools'
datasets:lists, so the pipeline stops asking. - Tools that are complete, disabled, or never set up are left alone.
- A missing image is pulled again with the exact tag in
- Want to add a tool you skipped, re-enable a disabled one, or add FragPipe's JARs later? Add
--setupto the same command. Setup then also offers every tool that is not set up (default answer: no). FragPipe without its JARs is offered with default yes. - Before any job starts, the whole
config.yamlis checked once (paths, FASTA files,output_dir, images). If something is wrong, you get one numbered list of problems instead of failing jobs. config.yamlis edited as text, so your comments and layout are kept. It is only written when something changed, and the previous version is kept asconfig.yaml.bak.
Setup-wizard details per tool:
- MaxQuant, Sage, MetaMorpheus, AlphaDIA — a yes/no prompt each; a plain
docker pullif you say yes. - FragPipe — the
fcyucn/fragpipeimage does not include MSFragger, IonQuant, or diaTracer (Nesvilab Academic License, separate from FragPipe's own license). For each of the three, the wizard asks whether you already have it downloaded as a.zipor extracted folder; if not, it prints the download URL and lets you skip — FragPipe is written to the config but staysenabled: falseuntil all three are present. Run the pipeline with--setupto add them later. When you point the wizard at an MSFragger folder, it also copies theext/folder shipped next to the jar (the Thermo.rawand Bruker.dnative readers, run under the mono runtime already in the image) and mounts it at run time, so FragPipe reads.raw/.ddirectly; ifext/is missing, FragPipe will need mzML input instead. FragPipe also needs decoys already appended to the FASTA (unlike the other tools); iffasta_decoy:isn't set for a dataset, one is generated automatically the first time that dataset is searched, via the Philosopher CLI already bundled in the image (the same command the FragPipe GUI's "Add decoys" button runs), and cached next to the source FASTA for reuse. - DIA-NN — always pulls the free
biocontainers/diann:v1.8.1_cv1image. It also asks whether to build DIA-NN 2.x images (needed for DDA support and native Thermo.rawreading on Linux); if you say yes, itgit clones bigbio/quantms-containers, lists the recipes it ships (currently 1.8.1, 1.9.2, 2.0.2, 2.1.0, 2.2.0, 2.3.2, 2.5.0, 2.5.1) and lets you pick one or several as a comma-separated list. Each one is built withdocker buildlocally and written to the config as its own version entry (diann:<version>), so several DIA-NN versions can be benchmarked side by side on the same datasets. DIA-NN itself is downloaded from the public vdemichev/DiaNN releases during the build, so no registry account or token is needed (requiresgit; a few minutes per version).supports_ddais set automatically:truefrom 2.1.0 onward. If you decline, only 1.8.1 is configured. Already-built images are detected and reused instead of rebuilt, so re-running the wizard to add another version is cheap. On a later run, the wizard lists which configured versions are enabled and present, which are disabled, and which areenabled: truebut have no local image. For each of the last group it offers to pull or rebuild the image, or to set that version toenabled: false. Disabled versions are never pulled or rebuilt. - Datasets — after the tools above are set up, the wizard offers to download benchmark datasets from
nextflow/datasets_catalog.yaml, scoped to the datasets relevant to your enabled tools (by DDA/DIA acquisition). They are stored indata/next toconfig.yaml(change with--data_dir). It lists the relevant datasets with their size and the free disk space, and lets you pickall,none, or specific ones by number. Each dataset is downloaded as an archive (.tar.gzor.zip) and extracted. An interrupted download resumes on the next setup run where the server allows it, and a half-extracted dataset is detected and extracted again. After extraction the wizard checks that the MS files of the dataset's format are really there. Datasets downloaded during a--setuprun are also added to thedatasets:lists of the enabled tools that can use them. The ProteoBench archives unpack asraws/<acquisition>/plusfasta/<name>.zip; the wizard moves the MS files up into the dataset folder and extracts the FASTA next to them, which is the flat layout the runners expect (datasets downloaded earlier are fixed the same way on the next setup run). The FASTA (and decoy, if present) are detected automatically, and anmzml/subfolder is moved to a sibling<name>_mzmldirectory — not a separate dataset entry, but a fallback location tools automatically use whenever they can't read the dataset's native format and need mzML instead (e.g. Sage always; DIA-NN < 2.0 for Thermo.raw). Real, resolved paths are written intodatasets:, and each configured tool'sdatasets:list is filled in automatically — no moreCHANGE_MEfor anything that was downloaded. Datasets already present on disk from an earlier run are reused, not re-downloaded. See "Adding a downloadable dataset" below.
Non-interactive / CI use (skips all prompts, uses these flags instead):
nextflow run proteobench.nf --non_interactive --skip_fragpipe \
--build_diann_v2 --diann_version 2.1.0,2.5.0 \
--download_datasets all --data_dir /path/to/data--setup_tools sage,diann sets up exactly those tools without asking (none for no tools); without it, non-interactive mode sets up every tool. --skip_datasets skips the dataset step entirely; --download_datasets also accepts a comma-separated list of dataset names instead of all. With no --download_datasets given, non-interactive mode downloads nothing (safe default for CI).
Available dataset names (source of truth: nextflow/datasets_catalog.yaml):
| Name | Acquisition | Format | Instrument |
|---|---|---|---|
HYE_DDA_Orbitrap |
DDA | raw | Orbitrap |
HYE_DDA_Astral |
DDA | raw | Astral |
HYE_Astral |
DIA | raw | Astral |
HYE_Astral_Single_Cell |
DIA | raw | Astral |
HYE_AIF |
DIA | raw | Orbitrap |
HYE_diaPASEF |
DIA | d | timstof |
HYE_ZenoSWATH |
DIA | wiff | ZenoTOF |
PYE_diaPASEF |
DIA | d | timstof |
Entrapment_DIA |
DIA | raw | Orbitrap |
To run the wizard again regardless of completeness (e.g. to add a tool you skipped), add --setup:
nextflow run ProteoBench/ProteoRunners -r v1.0.4 --setup # pulled release
nextflow run proteobench.nf --setup # inside a cloneIn a clone, nextflow run setup.nf does the same without starting any runs afterwards.
The pipeline reads config.yaml from the folder you run it from by default. Each job writes its actual result files under global.output_dir from that file. Relative paths in config.yaml (output_dir, and path:, fasta: and fasta_decoy: under datasets:) are resolved relative to the directory that holds config.yaml. The repo's nextflow.config also reads config.yaml: Nextflow's concurrency (maxForks) defaults to global.max_parallel_jobs (or 6 if unset), and run_summary_nf.tsv is published to global.output_dir (or results/ next to config.yaml while it is unset or still CHANGE_ME). Override either with --max_parallel_jobs / --publish_dir. nextflow.config also lists the defaults of every --flag of proteobench.nf and setup.nf.
The scripts use Nextflow's strict syntax, which Nextflow 26.04+ requires by default, and they still run with the legacy parser (NXF_SYNTAX_PARSER=v1). See Tests for how to check changes.
| Flag | Description |
|---|---|
--config /path/to/config.yaml |
Path to config file (default: config.yaml in the folder you run from) |
--setup |
Run the setup wizard first and offer every tool that is not set up yet |
--data_dir /path |
Where setup downloads datasets (default: data/ next to config.yaml) |
--tool diann |
Restrict run to one tool |
--dataset Entrapment_DIA |
Restrict run to one dataset |
--no_preflight |
Skip preflight checks before each job |
--max_parallel_jobs 4 |
Override Nextflow concurrency (default: global.max_parallel_jobs, or 6) |
--publish_dir /path |
Where run_summary_nf.tsv is published (default: global.output_dir, or ./results) |
Example — run only DIA-NN jobs, skip preflight:
nextflow run proteobench.nf --tool diann --no_preflightExample — run against a custom config and limit concurrency:
nextflow run proteobench.nf --config /data/my_config.yaml --max_parallel_jobs 2Nextflow caches each completed job in the work/ directory. If a run is interrupted, resume it without re-running successful jobs:
nextflow run proteobench.nf -resumeJobs that already have a .done marker in the output directory are also skipped by the runner logic itself, so both layers protect against redundant work.
Each job's actual result files are written to global.output_dir (from config.yaml). The summary file, run_summary_nf.tsv, is published to global.output_dir as well, or to ./results while output_dir is unset or still CHANGE_ME (override with --publish_dir). Columns: tool, version, dataset, success, skipped, runtime_s, output_dir, error_msg.
Nextflow task working directories are placed under Nextflow's default work/ directory in the project root (override with -w /path/to/dir). To delete them after a successful run:
nextflow clean -fTo also redirect the Nextflow log into the results directory, pass -log on the command line:
nextflow run proteobench.nf -log /path/to/results/.nextflow.logTo run on SLURM (or another executor), add executor settings to a nextflow.config in the directory you run nextflow run from (Nextflow merges it with the repo's own nextflow.config automatically — this works the same whether you're running from a clone or a pulled release):
// nextflow.config in your working directory
process.executor = 'slurm'
process.queue = 'gpu'
process.clusterOptions = '--mem=64G --time=04:00:00'See the Nextflow executor documentation for other executors (PBS, LSF, Kubernetes, etc.). No changes to proteobench.nf itself are needed.
| Tool | Acquisition | Input format | Docker image |
|---|---|---|---|
| DIA-NN | DDA (v2.1+), DIA | raw, mzML | biocontainers/diann:v1.8.1_cv1 (public) or diann:2.x (built locally from bigbio/quantms-containers) |
| AlphaDIA | DIA | raw, mzML, .d | mannlabs/alphadia:latest |
| Sage | DDA | mzML, MGF | ghcr.io/lazear/sage:latest |
| FragPipe | DDA, DIA | raw, mzML, .d | fcyucn/fragpipe:latest + licensed MSFragger/IonQuant/diaTracer JARs |
| MaxQuant | DDA, DIA | raw | quay.io/medbioinf/maxquant:2.6.3.0, quay.io/medbioinf/maxquant:2.8.1.0 |
| MetaMorpheus | DDA | raw, mzML | smithchemwisc/metamorpheus:latest |
All six images are pulled by the setup wizard (run automatically by nextflow run proteobench.nf, or with --setup) — see Quick start above. There is no native/manual install path any more; every tool runs from its image.
setup.nf always pulls :latest (except DIA-NN, which is described above, and MaxQuant, which pulls two fixed versions: 2.6.3.0 and 2.8.1.0). To pin or add a different tag, edit config.yaml directly — add a new entry under that tool's versions: list with the new image: tag and enabled: true. Any *_bin/*_dir/fragpipe_root path may need updating too if the new image version changes its internal layout; docker run --rm --entrypoint find <image> / -maxdepth 4 -iname <binary-name> (as setup.nf does internally) will locate it.
Each tool version has an enabled flag in config.yaml. Setting it to false skips that version without removing its configuration:
tools:
diann:
versions:
- id: "2.5.0"
image: ghcr.io/bigbio/diann:2.5.0
diann_bin: /usr/diann/2.5.0/diann
enabled: true # ← run this version
- id: "1.8.1"
image: biocontainers/diann:v1.8.1_cv1
diann_bin: /usr/diann/1.8.1/diann
enabled: false # ← skip this versionAdd a block under datasets: in config.yaml, then add the dataset name to the datasets: list of each tool that should run on it:
datasets:
My_New_Dataset:
path: /data/my_experiment # directory containing the MS files
acquisition: DIA # DDA or DIA
format: raw # raw | mzml | d | wiff | mgf
instrument: Orbitrap # Orbitrap | Astral | timstof | ZenoTOF
fasta: /data/fastas/human.fasta
fasta_decoy: /data/fastas/human_decoy.fasta # optional; FragPipe generates one automatically if omitted
tools:
diann:
datasets:
- Entrapment_DIA
- My_New_Dataset # ← add hereTo let setup download a dataset automatically instead of a user pointing path:/fasta: at existing files by hand, add an entry to nextflow/datasets_catalog.yaml with a real download URL:
My_New_Dataset:
url: https://example.org/path/to/My_New_Dataset.zip
acquisition: DIA # DDA or DIA
format: raw # raw | mzml | d | wiff | mgf
instrument: Orbitrap # Orbitrap | Astral | timstof | ZenoTOFThe archive may use the ProteoBench layout (raws/<acquisition>/ + fasta/<name>.zip) or contain, at its top level: the MS files, one *.fasta (+ optionally one decoy fasta with decoy in the name), and optionally an mzml/ subfolder. Run the pipeline with --setup — it offers the new dataset (scoped to tools it's relevant to), downloads and unzips it, and writes the resolved path:/fasta: into config.yaml automatically.
Copy an existing version block, update id, image, and the tool's in-container binary path, then set enabled: true. Each tool uses its own key for that path — diann_bin (DIA-NN), sage_bin (Sage), maxquant_dll (MaxQuant), fragpipe_root (FragPipe); AlphaDIA and MetaMorpheus need no path key:
tools:
diann:
versions:
- id: "2.6.0" # new version
image: diann:2.6.0 # docker image tag
diann_bin: /usr/diann-2.6.0/diann # in-container binary path
supports_dda: true # DIA-NN only: whether this build supports DDA
enabled: trueEach tool run creates a subdirectory under output_dir:
results/
└── Entrapment_DIA/
├── diann_v2.5.0/
│ ├── report.tsv ← DIA-NN result file
│ ├── stdout.log
│ ├── stderr.log
│ └── .done ← marker: job succeeded; delete to re-run
├── alphadia_v2.1.1/
│ └── ...
└── run_summary_20260603_120000.tsv ← summary of all runs
The .done marker causes the pipeline to skip that job on the next run (useful for resuming after interruption). Set overwrite: true in config.yaml to force all jobs to re-run.
There are two test suites. Run both before a release or after changing the setup logic.
Nextflow (nf-test) tests the helpers in setup.nf and the setup gate of proteobench.nf end to end:
curl -fsSL https://get.nf-test.com | bash # once; puts ./nf-test here, move it onto your PATH
nf-test test tests/nextflow # about 2 minutestests/nextflow/setup_functions.nf.test: how setup decides what to do per tool, and how it editsconfig.yamlas text.tests/nextflow/pipeline.nf.test: non-interactive runs from a fresh folder (first run, missing image, missing dataset, broken downloads, nothing to do). Datasets come from tiny archives intests/fixtures/datasetsviafile://URLs, so no network is needed for them. These tests need Docker; they use fake image tags, the tinyhello-worldimage, and the Sage image for the first-run test.
Python (pytest) tests config_validator.py and the path handling of run_single_job.py, with a fake Docker:
python3 -m venv --system-site-packages .venv && .venv/bin/pip install -r requirements-dev.txt # once
.venv/bin/pytestThe interactive prompts are not covered by the automated tests: nf-test cannot type answers. Check those by hand after changing a prompt.
| Error message | Likely cause | Fix |
|---|---|---|
still contains 'CHANGE_ME' |
Path not updated in config | Edit config.yaml and replace CHANGE_ME |
This computer is not ready yet |
Docker is missing, not running, or not usable without sudo; or python3 has no pyyaml |
Follow the fix printed under the message |
docker image not pulled locally |
Image not pulled yet | Run the pipeline again; setup offers to pull it |
No 'raw' files found |
Wrong format: or wrong path: |
Verify files exist and format: matches |
MSFragger/IonQuant/diaTracer JAR not found in jars_dir |
Licensed FragPipe JAR missing | Run the pipeline with --setup and provide the downloaded zip/folder when asked |
'ext/thermo' folder was not found next to MSFragger jar |
MSFragger's native readers weren't copied (pointed the wizard at a bare .jar, not its folder) |
Delete the MSFragger jar in jars_dir, run the pipeline with --setup and point at the MSFragger folder or zip (which has ext/), or switch that dataset to mzML input |
build_command failed: Philosopher decoy generation failed |
FASTA's directory isn't writable, or the FASTA is malformed | Check permissions on the FASTA's directory; or set fasta_decoy: explicitly to a pre-built one |
no download URL set in the dataset catalog |
Catalog entry still has url: CHANGE_ME |
Add the real URL to nextflow/datasets_catalog.yaml, then run the pipeline with --setup |
not ready to run yet (N problem(s) above) |
config.yaml has problems that would make jobs fail |
Fix the numbered problems, or run the pipeline with --setup |
is not valid YAML |
A typo in a hand edit of config.yaml |
Fix the line shown in the message, or restore config.yaml.bak |
| Setup keeps asking about a dataset | An enabled tool lists a dataset that is not on disk | Download it, or let setup comment it out of the tools' datasets: lists |
| Dataset download skipped in CI | Non-interactive mode with no --download_datasets (safe default) |
Pass --download_datasets all or a comma-separated list |
exit code 1 — check log: ... |
Tool crashed during search | Open the stderr.log file shown in the error |
| Job is skipped unexpectedly | .done marker exists |
Delete the .done file in the output directory, or set overwrite: true |
No jobs to run |
All versions have enabled: false, or the enabled tools list no dataset |
Set enabled: true for at least one version, and list dataset names from datasets: under that tool's datasets: |
| Setup keeps asking about a DIA-NN version | That version is enabled: true but its image is not present locally |
Let setup pull or build it, or set enabled: false for it (setup offers this) |
Groovy import declarations are not supported or for loops are no longer supported |
An old checkout with Nextflow 26.04+ (strict parser) | Update to the current version, or run with NXF_SYNTAX_PARSER=v1 |
enumerate_jobs.py failed (Nextflow) |
Python or YAML not on PATH | Run Nextflow from the activated conda/pip env; check python3 -c "import yaml" |
No module named 'yaml' (Nextflow) |
pyyaml not installed | pip install pyyaml |
| Nextflow process hangs | max_parallel_jobs too high for available cores/RAM |
Lower global.max_parallel_jobs in config.yaml or pass --max_parallel_jobs N |