Skip to content

Measles prod code - Feature/sra biosample primaryid - #363

Open
donutbrew wants to merge 151 commits into
masterfrom
feature/sra-biosample-primaryid
Open

donutbrew wants to merge 151 commits into
masterfrom
feature/sra-biosample-primaryid

Conversation

@donutbrew

Copy link
Copy Markdown
Collaborator

Description

This is the version of tostadas being used for MeV submissions.

@jforstedt may provide more details

Checklist

Go Through Checklist Below and Place A ✔️ (X Inside the Box) if Completed

General Checks

  • [] Have you run appropriate tests (unit/integration/end-to-end) to check logic across run environments (Conda/Docker/Singularity on Scicomp/AWS/NF Tower/Local)?

    For each relevant configuration:

    • Can the program run completely through without erroring out?
    • Does it produce the expected outputs, given the inputs provided?
  • [] Have you conducted proper linting procedures?

    • Numpy formatted docstrings for functions
    • Comments explaining lines of code
    • Consistent and intuitive naming conventions for variables, functions, classes, methods, attributes, and scripts
    • Single empty line between class functions, two lines between non-class functions, and two lines between imports and code body
    • Camel case formatting for class names
  • [] Have you updated existing documentation (README.md, etc.) or created new ones within docs?

CDC Checks

  • [] Did you check for sensitive data, and remove any?
  • [] If you added or modified HTML, did you check that it was 508 compliant?

Are additional approvals needed for this change? If so, please mention them below:

Are there potential vulnerabilities or licensing issues with any new dependencies introduced? If so, please mention them below:

Jessica Rowell and others added 30 commits September 20, 2025 12:07
…wipes the folder even if overwrite is set to false, probably because it's a separate run
Bypassing review because it's Saturday and the last day of the contract so no one is available to review.
Externalize hardcoded mpox-specific v-annotate.pl flags into a new
vadr_opts param so each organism profile controls its own VADR options.
Add measles profile using Greninger Lab MeV models (vadr-models-mev)
and set correct vadr_opts for RSV per upstream model documentation.
Add test FASTA (AF266288 Edmonston strain) and test metadata Excel
for the measles profile. Set meta_path in measles.config to match
the pattern used by all other organism profiles.
The measles CM file (mev.cm, ~150MB) exceeds GitHub's file size limit
and ships as a Git LFS pointer. Rather than requiring git-lfs or manual
setup steps, add a VADR_MODEL_SETUP process that runs inside the VADR
container before annotation. It downloads the CM via curl when
vadr_cm_url is set and the local copy is missing or invalid, then
rebuilds indices with cmpress. Cached on -resume so it only runs once.

Also adds vadr_cm_url param (default empty) to nextflow.config, sets it
in measles.config, and registers vadr_opts and vadr_cm_url in the
nextflow schema to suppress validation warnings.
The GFF file written in line_cleanup() was never closed, so its write
buffer had not been flushed to disk when gff2tbl() opened the same file
for reading. This caused gff2tbl to see an empty file and produce a .tbl
with only the header line. Close the file handle after writing so the
data is on disk before the conversion step reads it.
The previous gff2tbl loop (anns[0:-2]) skipped the second-to-last
attribute, which was typically the product qualifier. Also, the GFF3
ID attribute was written directly to the .tbl file, but ID is not a
valid NCBI qualifier and caused table2asn to reject the file.

Fix the loop to iterate all attributes, convert ID back to protein_id
for CDS features (stripping the GFF prefix), drop ID for gene features,
and skip invalid numeric keys from multi-CDS gene parsing.
For genes with multiple CDS entries (e.g. measles P gene which encodes
P, V, and C proteins), later CDS qualifiers were overwriting earlier
ones. This caused the primary gene product (phosphoprotein) to be
replaced by the last CDS product (C protein). Preserve the first
encountered value for each qualifier key instead of overwriting.
Once the first CDS block is fully parsed (indicated by protein_id being
set), stop accepting new qualifier keys. This prevents qualifiers from
subsequent CDS entries (e.g. V protein exception and note) from leaking
into the primary gene product annotation.
Genes encoding multiple proteins (e.g. measles P gene with P, V, C
proteins) now have all CDS entries preserved through the cleanup
pipeline and inserted into the final .tbl file for complete GenBank
annotation.
codon_start is not a valid qualifier on gene features in NCBI feature
tables. table2asn rejects it with "Unrecognized qualifier name". Filter
out CDS-only qualifiers (codon_start, product, protein_id, transl_except,
transl_table) when writing gene lines in gff2tbl.
…tion_date

Sequence_ID falls back to sample_id when ncbi-spuid-sra is empty (non-SRA
submissions). Country falls back to country + state when geo_loc_name is
absent (CREATE_BATCH_TSVS bypasses validation). Collection_date strips
time component from pandas datetime strings.
…llbacks

extract_biosample_metadata() did not include country or state, so the
geo_loc_name fallback could never read them. Also handle pandas NaN
(truthy) for ncbi-spuid-sra and geo_loc_name so fallbacks trigger correctly.
RNA viruses (measles, RSV) need viral cRNA instead of the default
genomic molecule type. Added mol_type parameter to profile configs
and passed through to table2asn via the -j flag.

Removed the Seqdesc pub block from authorset.sbt generation per
NCBI preference to manage publication references separately.
Older table2asn versions don't support mol_type via -j. The standard
approach is to add [mol_type=viral cRNA] to the FASTA definition line,
which works across all table2asn versions.
The container's table2asn is too old to support mol_type via -j flag
or FASTA defline modifiers. Instead, run table2asn normally and then
replace 'biomol genomic' with 'biomol cRNA' in the output .sqn file
when mol_type is set to viral cRNA in the profile config.
NCBI prefers pub citation and DBLink/BioProject blocks managed through
their portal rather than embedded in .sqn files. Added strip_pub_block
parameter (default false) that post-processes the .sqn to remove these
blocks. Enabled by default in measles profile.
jforstedt and others added 16 commits March 26, 2026 10:43
When metadata row has biosample_accession populated with a SAMN id,
emit <PrimaryId db="BioSample">SAMN…</PrimaryId> instead of the
SPUID-based BioSample reference. Mirrors the GenBank submission
pattern. Falls back to existing SPUID behavior when biosample_accession
is absent, preserving backwards compatibility.
Read metadata with dtype=str (submission_prep, submission_update) so a numeric sample_name stays a string and matches sample_id; it was inferred as int, giving an empty match that crashed BioSample/SRA and made GenBank use the wrong row. Raise on an empty match instead of crashing or falling back.
Re-running a batch into a submit/<mode>/<id>_<batch>_<db> folder that a
prior submission already populated corrupts the submission at NCBI. The
folder name isn't run-unique, so re-running with a different batch_size
reshuffles samples into a directory that already has files.

Check the remote directory after connecting and before creating it, and
stop with a clear message if it exists and is non-empty. Add
--allow_dir_collision to override when the reuse is intentional, plus
dir_exists and list_dir helpers on the FTP and SFTP clients.
metadata_template.xlsx used ncbi_sequence_name_sra, which nothing in the
code reads. The submission modules and every other package template use
ncbi-spuid-sra for the SRA SPUID. Rename the header so the template lines
up with the validator and submission code.
Split authors on ";" (optional whitespace) so ";"-separated lists don't
collapse to one name, and add fix_name's missing fallback return.
@donutbrew
donutbrew requested review from jforstedt and a lite review from Copilot August 28, 2026 00:57
@donutbrew donutbrew added the enhancement New feature or request label Aug 28, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@donutbrew
donutbrew requested a lite review from Copilot August 28, 2026 01:22
@donutbrew donutbrew changed the title Feature/sra biosample primaryid Measles prod code - Feature/sra biosample primaryid Aug 28, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 113 out of 266 changed files in this pull request and generated 7 comments.

Suppressed comments (2)

workflows/biosample_update.nf:86

  • The dry_run branch logs that the workflow "ends here", but execution continues into UPDATE_SUBMISSION. If dry_run is meant to stop before any submission/update steps, return early (or gate the remaining steps under an else).
    workflows/biosample_update.nf:59
  • System.exit(1) on missing batch_summary.json will terminate the JVM abruptly. Use workflow.abort(...) so Nextflow fails gracefully and still reports a clear error upstream (Tower/CLI/logs).

Comment thread workflows/genbank.nf
Comment on lines +57 to +61
CHECK_VALIDATION_ERRORS.out.status.subscribe { status ->
if (status == "ERROR") {
log.error "Validation failed. Please check ${params.outdir}/${params.validation_outdir}/error.txt"
System.exit(1)
}
Comment on lines 39 to 43
CHECK_VALIDATION_ERRORS.out.status.subscribe { status ->
if (status == "ERROR") {
log.info "Validation failed. Please check ${params.outdir}/${params.metadata_basename}/${params.validation_outdir}/error.txt"
workflow.abort()
log.info "Validation failed. Please check ${params.outdir}/${params.validation_outdir}/error.txt"
System.exit(1)
}
Comment on lines +77 to +80
def parsed = new groovy.json.JsonSlurper().parseText(json_file.text)
def batch_id = parsed.meta.batch_id
def matching_tsv = tsv_list.find { it.name == "${batch_id}.tsv" }
tuple(parsed.meta, parsed.samples, parsed.enabled, matching_tsv)
Comment on lines 41 to 45
CHECK_VALIDATION_ERRORS.out.status.subscribe { status ->
if (status == "ERROR") {
log.info "Validation failed. Please check ${params.outdir}/${params.metadata_basename}/${params.validation_outdir}/error.txt"
workflow.abort()
log.info "Validation failed. Please check ${params.outdir}/${params.validation_outdir}/error.txt"
System.exit(1)
}
Comment on lines +20 to +22
# Build fail list from VADR pass/fail TSV
awk -F'\\t' 'NR>1 && \$8>0 {gsub(/_mev\\.vadr\\.mdl/,"",\$1); print \$1}' ${pass_fail_tsv} > fail_list.txt

Comment on lines +13 to +15
REPO = os.path.join(HERE, "..", "..")
BUNDLE = os.path.join(REPO, "vadr_runs_bundle")
sys.path.insert(0, BIN)
Comment on lines +35 to +39
if [ -n "${cm_url}" ]; then
if [ ! -f "\$cm" ] || [ "\$(wc -c < "\$cm" 2>/dev/null || echo 0)" -lt 1000 ]; then
curl -L -o "\$cm" "${cm_url}"
needs_cmpress=true
fi
NCBI's submission schema requires Hold after Organization. The BioSample/SRA
builder emitted it before Comment and Organization, so any submission with a
Specified_Release_Date was rejected ("Element 'Hold': This element is not
expected"). Move Hold to the end of the Description block, matching the
GenBank builder.
Fix the Submission Config link (main to master, the repo default branch),
add the Specified_Release_Date format (YYYY-MM-DD) to the config fields
table, and document the Org_ID field.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants