Repository navigation
Conversation
…wipes the folder even if overwrite is set to false, probably because it's a separate run
Bypassing review because it's Saturday and the last day of the contract so no one is available to review.
Externalize hardcoded mpox-specific v-annotate.pl flags into a new vadr_opts param so each organism profile controls its own VADR options. Add measles profile using Greninger Lab MeV models (vadr-models-mev) and set correct vadr_opts for RSV per upstream model documentation.
Add test FASTA (AF266288 Edmonston strain) and test metadata Excel for the measles profile. Set meta_path in measles.config to match the pattern used by all other organism profiles.
The measles CM file (mev.cm, ~150MB) exceeds GitHub's file size limit and ships as a Git LFS pointer. Rather than requiring git-lfs or manual setup steps, add a VADR_MODEL_SETUP process that runs inside the VADR container before annotation. It downloads the CM via curl when vadr_cm_url is set and the local copy is missing or invalid, then rebuilds indices with cmpress. Cached on -resume so it only runs once. Also adds vadr_cm_url param (default empty) to nextflow.config, sets it in measles.config, and registers vadr_opts and vadr_cm_url in the nextflow schema to suppress validation warnings.
The GFF file written in line_cleanup() was never closed, so its write buffer had not been flushed to disk when gff2tbl() opened the same file for reading. This caused gff2tbl to see an empty file and produce a .tbl with only the header line. Close the file handle after writing so the data is on disk before the conversion step reads it.
The previous gff2tbl loop (anns[0:-2]) skipped the second-to-last attribute, which was typically the product qualifier. Also, the GFF3 ID attribute was written directly to the .tbl file, but ID is not a valid NCBI qualifier and caused table2asn to reject the file. Fix the loop to iterate all attributes, convert ID back to protein_id for CDS features (stripping the GFF prefix), drop ID for gene features, and skip invalid numeric keys from multi-CDS gene parsing.
For genes with multiple CDS entries (e.g. measles P gene which encodes P, V, and C proteins), later CDS qualifiers were overwriting earlier ones. This caused the primary gene product (phosphoprotein) to be replaced by the last CDS product (C protein). Preserve the first encountered value for each qualifier key instead of overwriting.
Once the first CDS block is fully parsed (indicated by protein_id being set), stop accepting new qualifier keys. This prevents qualifiers from subsequent CDS entries (e.g. V protein exception and note) from leaking into the primary gene product annotation.
Genes encoding multiple proteins (e.g. measles P gene with P, V, C proteins) now have all CDS entries preserved through the cleanup pipeline and inserted into the final .tbl file for complete GenBank annotation.
codon_start is not a valid qualifier on gene features in NCBI feature tables. table2asn rejects it with "Unrecognized qualifier name". Filter out CDS-only qualifiers (codon_start, product, protein_id, transl_except, transl_table) when writing gene lines in gff2tbl.
…tion_date Sequence_ID falls back to sample_id when ncbi-spuid-sra is empty (non-SRA submissions). Country falls back to country + state when geo_loc_name is absent (CREATE_BATCH_TSVS bypasses validation). Collection_date strips time component from pandas datetime strings.
…llbacks extract_biosample_metadata() did not include country or state, so the geo_loc_name fallback could never read them. Also handle pandas NaN (truthy) for ncbi-spuid-sra and geo_loc_name so fallbacks trigger correctly.
RNA viruses (measles, RSV) need viral cRNA instead of the default genomic molecule type. Added mol_type parameter to profile configs and passed through to table2asn via the -j flag. Removed the Seqdesc pub block from authorset.sbt generation per NCBI preference to manage publication references separately.
Older table2asn versions don't support mol_type via -j. The standard approach is to add [mol_type=viral cRNA] to the FASTA definition line, which works across all table2asn versions.
The container's table2asn is too old to support mol_type via -j flag or FASTA defline modifiers. Instead, run table2asn normally and then replace 'biomol genomic' with 'biomol cRNA' in the output .sqn file when mol_type is set to viral cRNA in the profile config.
NCBI prefers pub citation and DBLink/BioProject blocks managed through their portal rather than embedded in .sqn files. Added strip_pub_block parameter (default false) that post-processes the .sqn to remove these blocks. Enabled by default in measles profile.
When metadata row has biosample_accession populated with a SAMN id, emit <PrimaryId db="BioSample">SAMN…</PrimaryId> instead of the SPUID-based BioSample reference. Mirrors the GenBank submission pattern. Falls back to existing SPUID behavior when biosample_accession is absent, preserving backwards compatibility.
…t-in submission Title/Comment
Read metadata with dtype=str (submission_prep, submission_update) so a numeric sample_name stays a string and matches sample_id; it was inferred as int, giving an empty match that crashed BioSample/SRA and made GenBank use the wrong row. Raise on an empty match instead of crashing or falling back.
Re-running a batch into a submit/<mode>/<id>_<batch>_<db> folder that a prior submission already populated corrupts the submission at NCBI. The folder name isn't run-unique, so re-running with a different batch_size reshuffles samples into a directory that already has files. Check the remote directory after connecting and before creating it, and stop with a clear message if it exists and is non-empty. Add --allow_dir_collision to override when the reuse is intentional, plus dir_exists and list_dir helpers on the FTP and SFTP clients.
metadata_template.xlsx used ncbi_sequence_name_sra, which nothing in the code reads. The submission modules and every other package template use ncbi-spuid-sra for the SRA SPUID. Rename the header so the template lines up with the validator and submission code.
Split authors on ";" (optional whitespace) so ";"-separated lists don't collapse to one name, and add fix_name's missing fallback return.
…mple-primaryid-bugfix
…ugfix Fix genbank authors source object
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 113 out of 266 changed files in this pull request and generated 7 comments.
Suppressed comments (2)
workflows/biosample_update.nf:86
- The dry_run branch logs that the workflow "ends here", but execution continues into UPDATE_SUBMISSION. If dry_run is meant to stop before any submission/update steps, return early (or gate the remaining steps under an else).
workflows/biosample_update.nf:59 - System.exit(1) on missing batch_summary.json will terminate the JVM abruptly. Use workflow.abort(...) so Nextflow fails gracefully and still reports a clear error upstream (Tower/CLI/logs).
Comment on lines
+57
to
+61
| CHECK_VALIDATION_ERRORS.out.status.subscribe { status -> | ||
| if (status == "ERROR") { | ||
| log.error "Validation failed. Please check ${params.outdir}/${params.validation_outdir}/error.txt" | ||
| System.exit(1) | ||
| } |
Comment on lines
39
to
43
| CHECK_VALIDATION_ERRORS.out.status.subscribe { status -> | ||
| if (status == "ERROR") { | ||
| log.info "Validation failed. Please check ${params.outdir}/${params.metadata_basename}/${params.validation_outdir}/error.txt" | ||
| workflow.abort() | ||
| log.info "Validation failed. Please check ${params.outdir}/${params.validation_outdir}/error.txt" | ||
| System.exit(1) | ||
| } |
Comment on lines
+77
to
+80
| def parsed = new groovy.json.JsonSlurper().parseText(json_file.text) | ||
| def batch_id = parsed.meta.batch_id | ||
| def matching_tsv = tsv_list.find { it.name == "${batch_id}.tsv" } | ||
| tuple(parsed.meta, parsed.samples, parsed.enabled, matching_tsv) |
Comment on lines
41
to
45
| CHECK_VALIDATION_ERRORS.out.status.subscribe { status -> | ||
| if (status == "ERROR") { | ||
| log.info "Validation failed. Please check ${params.outdir}/${params.metadata_basename}/${params.validation_outdir}/error.txt" | ||
| workflow.abort() | ||
| log.info "Validation failed. Please check ${params.outdir}/${params.validation_outdir}/error.txt" | ||
| System.exit(1) | ||
| } |
Comment on lines
+20
to
+22
| # Build fail list from VADR pass/fail TSV | ||
| awk -F'\\t' 'NR>1 && \$8>0 {gsub(/_mev\\.vadr\\.mdl/,"",\$1); print \$1}' ${pass_fail_tsv} > fail_list.txt | ||
|
|
Comment on lines
+13
to
+15
| REPO = os.path.join(HERE, "..", "..") | ||
| BUNDLE = os.path.join(REPO, "vadr_runs_bundle") | ||
| sys.path.insert(0, BIN) |
Comment on lines
+35
to
+39
| if [ -n "${cm_url}" ]; then | ||
| if [ ! -f "\$cm" ] || [ "\$(wc -c < "\$cm" 2>/dev/null || echo 0)" -lt 1000 ]; then | ||
| curl -L -o "\$cm" "${cm_url}" | ||
| needs_cmpress=true | ||
| fi |
NCBI's submission schema requires Hold after Organization. The BioSample/SRA
builder emitted it before Comment and Organization, so any submission with a
Specified_Release_Date was rejected ("Element 'Hold': This element is not
expected"). Move Hold to the end of the Description block, matching the
GenBank builder.
Fix the Submission Config link (main to master, the repo default branch), add the Specified_Release_Date format (YYYY-MM-DD) to the config fields table, and document the Org_ID field.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This is the version of tostadas being used for MeV submissions.
@jforstedt may provide more details
Checklist
Go Through Checklist Below and Place A ✔️ (X Inside the Box) if Completed
General Checks
[] Have you run appropriate tests (unit/integration/end-to-end) to check logic across run environments (Conda/Docker/Singularity on Scicomp/AWS/NF Tower/Local)?
For each relevant configuration:
[] Have you conducted proper linting procedures?
[] Have you updated existing documentation (README.md, etc.) or created new ones within docs?
CDC Checks
Are additional approvals needed for this change? If so, please mention them below:
Are there potential vulnerabilities or licensing issues with any new dependencies introduced? If so, please mention them below: