Skip to content
CaizhaohuiPublic

About

Single-molecule time series Hidden Markov Model (HMM) state classification tool. Inspired by HaMMy, rewritten from scratch in Python, supporting cross-platform operation, batch processing, and GUI interaction.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

FretHMM

A single-molecule time-series Hidden Markov Model (HMM) state classification tool. Inspired by HaMMy, rewritten from scratch in Python with cross-platform support, batch processing, and a full GUI.

中文文档

Features

Feature Description
HMM Engine Baum-Welch training + Viterbi decoding (via hmmlearn), with customizable initial guesses
Data Modes Auto-detect / Single-channel signal / Dual-channel Donor-Acceptor (auto-computes FRET efficiency)
Batch Processing Multi-file parallel processing (ProcessPoolExecutor), directory scanning with multi-worker support
Review Grid Batch classification + paginated multi-panel PNG visual review for quick quality screening
Lowest-State Window Filtering Two-pass HMM fitting that scans forward for the first persistent lowest-state window and discards later signal
CLI Six subcommands: run, tdp, review-grid, events, dwell-stats, gui
GUI CustomTkinter interface with dark/light themes, English/Chinese switching, threaded background analysis, and batch review grid export
Output Formats *_classified.csv, *_summary.json, *report.dat, *path.dat, *dwell.dat (selectable in GUI)
TDP Transition Density Plot visualization + Gaussian rate fitting
Packaging PyInstaller one-click build for Windows executables (directory mode / --onefile mode)

Installation

git clone https://github.com/Caizhaohui/FretHMM.git
cd FretHMM
pip install -e .

Requirements:

  • Python >= 3.10
  • NumPy >= 1.24
  • SciPy >= 1.10
  • hmmlearn >= 0.3.0
  • matplotlib >= 3.7 (required for TDP and Review Grid visualization)
  • customtkinter >= 5.2.0 (required for GUI)

Optional dependencies:

pip install -e ".[dev]"    # Install pytest testing framework
pip install -e ".[gui]"    # Install PyInstaller packaging tool

Usage

CLI

FretHMM provides six subcommands: run (HMM analysis), review-grid (visual review), tdp (transition density plot), events (ON/OFF event analysis), dwell-stats (dwell-time statistics + rate-constant fit), and gui (graphical interface).

run — HMM State Classification

# Single file analysis (2 states, auto-detect data format)
frethmm run --files trace.csv --states 2 --output-dir ./results/

# Batch process all trace files in a directory (4 parallel workers)
frethmm run --input-dir ./traces/ --states 5 --workers 4 --output-dir ./results/

# Process multiple files at once
frethmm run --files trace1.csv trace2.csv trace3.csv --states 3 --output-dir ./results/

# Provide initial guesses (useful when state spacing is small)
frethmm run --files data.csv --states 2 --guesses "0.3,0.7"

# Specify single-channel mode and signal column
frethmm run --files data.csv --states 2 --mode single_channel --signal-column 1

# Use lowest-state window filtering (keep through the first 5-second lowest-state window, then re-classify)
frethmm run --files trace.csv --states 2 --low-state-tail-trim-seconds 5.0

# Output only classified.csv
frethmm run --files data.csv --states 2 --classified-only

# Verbose mode (show all warnings)
frethmm run --files data.csv --states 3 -v

run subcommand parameters:

Parameter Default Description
--files — One or more trace file paths (mutually exclusive with --input-dir, required)
--input-dir — Input directory to scan for trace files (mutually exclusive with --files, required)
--output-dir — Output directory (defaults to input file directory)
--states 2 Number of HMM states, or auto to pick via BIC (see Model selection)
--guesses None Comma-separated initial signal guesses; count must match --states (ignored with --states auto)
--max-iter 500 Maximum Baum-Welch iterations
--tol 1e-4 Convergence tolerance
--workers 1 Number of parallel workers (>1 enables multiprocessing)
--mode auto Data mode: auto / paired_channel / single_channel
--signal-column 1 1-based signal column index after Time for single_channel mode
--low-state-tail-trim-seconds None Lowest-state window duration in seconds; the option name is retained for compatibility (see Data Filtering)
--n-init 10 Number of deterministic multi-start Baum-Welch runs; best log-likelihood wins (use 1 to reproduce the legacy single-fit)
--min-states 2 Minimum state count for BIC selection (only with --states auto)
--max-states 6 Maximum state count for BIC selection (only with --states auto)
--remove-spikes off Enable outlier spike detection and cleaning via rolling median and robust MAD scale estimation
--spike-threshold-sigma 5.0 Spike detection threshold in units of robust sigma
--trim-initial-artifacts off Automatically detect and trim initial acquisition/shutter artifacts (e.g. Frame 0 surge)
--max-initial-artifact-frames 5 Maximum initial frames allowed to be trimmed as artifacts
--smooth-window None Edge-preserving median smoothing window (odd integer) to suppress shot noise before HMM fitting
--min-dwell-frames 1 Minimum state dwell duration in frames; shorter transient noise flickers are merged into neighbors
--merge-state-threshold None Absolute difference threshold below which adjacent fitted states are merged
--merge-state-sigma-factor None Relative scale threshold (k * sigma) below which adjacent fitted states are merged
--classified-only off Output only *_classified.csv, skip summary/report/path/dwell
-v / --verbose off Verbose output, show all warnings and fit quality diagnostics (SNR, RMSE, R^2)

Batch processing notes:

  • --input-dir scans all .csv, .dat, .txt, .tsv files, automatically skipping output files (*report.dat, *path.dat, *dwell.dat, *_classified.csv, *_summary.json)
  • --workers N enables multi-process parallelism; N should not exceed CPU core count
  • Individual file errors do not interrupt the overall batch; errors are printed to the terminal

review-grid — Batch Visual Review

# Basic: generate a 4×4 grid of 2-state traces
frethmm review-grid --input-dir ./traces/ --output review.png --states 2

# Custom grid layout
frethmm review-grid --input-dir ./traces/ --output review.png --states 3 --rows 5 --cols 6

# With initial guesses and output directory for classified CSVs
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
    --guesses "0.2,0.8" --output-dir ./classified/

# Combined with lowest-state window filtering
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
    --low-state-tail-trim-seconds 5.0

# Accelerate with 4 parallel workers
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
    --workers 4 --rows 4 --cols 8

review-grid subcommand parameters:

Parameter Default Description
--input-dir — Input trace file directory (either --input-dir or --files is required)
--files — Specify one or more trace files (either --files or --input-dir is required)
--output — Output PNG path, e.g. review.png (required)
--output-dir None Optional directory for classified CSV side outputs
--states 2 Number of HMM states, or auto to pick via BIC
--guesses None Comma-separated initial signal guesses (ignored with --states auto)
--max-iter 500 Maximum Baum-Welch iterations
--tol 1e-4 Convergence tolerance
--workers 1 Number of parallel workers
--mode auto Data mode: auto / paired_channel / single_channel
--signal-column 1 Signal column index for single_channel mode
--low-state-tail-trim-seconds None Lowest-state window duration in seconds (legacy option name)
--n-init 10 Deterministic multi-start count (use 1 for legacy single-fit)
--min-states 2 Minimum state count for BIC selection (only with --states auto)
--max-states 6 Maximum state count for BIC selection (only with --states auto)
--rows 4 Panel rows per page
--cols 4 Panels per row

Pagination: When the number of traces exceeds rows × cols, multiple page images are generated automatically (e.g., review_page_01.png, review_page_02.png). Each panel overlays the raw signal (gray) with the HMM classified signal (red). The title shows filename, log-likelihood, and state means. Traces with fitting warnings are highlighted with orange borders.

tdp — Transition Density Plot

# Generate TDP from report files (interactive window)
frethmm tdp --input-dir ./results/ --exposure 0.1

# Save to file
frethmm tdp --input-dir ./results/ --exposure 0.1 --output tdp.png

# Show only top N states (sorted by transition frequency)
frethmm tdp --input-dir ./results/ --exposure 0.1 --states 3 --output tdp.png

tdp subcommand parameters:

Parameter Default Description
--input-dir — Directory containing *report.dat files (required)
--exposure 0.1 Frame exposure time in seconds, used for rate calculations
--states None Show only top N states (sorted by transition frequency)
--output None Output image path (e.g., tdp.png); opens interactive window if not specified

events — ON/OFF Event Analysis

Extract discrete ON/OFF events from *_classified.csv files (the primary output of run). In a 2-state trace, high fluorescence is ON and a low segment is OFF only when high fluorescence later recovers (high → low → high); a terminal low segment is permanent loss of activity and is omitted. For 3+ states, each non-lowest state is analysed as an independent stage, with a descent counted as OFF only when that original stage recovers.

# Batch: scan a directory of *_classified.csv files
frethmm events --input-dir ./results/ --output-dir ./events/

# Process specific files
frethmm events --files trace1_classified.csv trace2_classified.csv --output-dir ./events/

# Legacy compatibility option; terminal low segments are omitted, not counted as OFF
frethmm events --input-dir ./results/ --tail-off-threshold-seconds 250 --output-dir ./events/

events subcommand parameters:

Parameter Default Description
--input-dir — Directory of *_classified.csv files (mutually exclusive with --files, required)
--files — One or more *_classified.csv paths (mutually exclusive with --input-dir, required)
--output-dir — Output directory (required)
--tail-off-threshold-seconds 100.0 Legacy compatibility option; terminal low segments are omitted rather than recorded as OFF

Four CSV tables are written per run:

File Description
event_details.csv One row per event: source file, type (ON/OFF), index, state value, start/end time and frame, duration, excluded flag, plus state_value_range — the file-level fluorescence amplitude: max − min over included events; for a file whose only included event is a single ON (e.g. ON followed by an omitted photobleach tail), ON level − minimum classified value (signal above the bleached baseline)
event_summary.csv One row per source file: ON/OFF counts, total and mean dwell times
event_stats_overall.csv Aggregate across all files: event counts, total/mean ON and OFF times
input_plot.csv One row per source file for downstream visualisation: source_file, ON_events, OFF_events, Fluorescence_strength (= state_value_range), Duration_time (= last event's end_time; for ON-then-bleach traces this is the observable duration before photobleaching)

dwell-stats — Dwell-Time Statistics + Rate-Constant Fit

Consumes the event_details.csv produced by events and computes the deeper descriptive statistics single-molecule analysis needs: median, standard deviation, min/max, and 25th/75th percentiles for ON and OFF dwell times, pooled across all molecules. Optionally fits a single exponential A·exp(-k·t) to each dwell-time distribution (histogram + scipy.optimize.curve_fit, bounded so k ≥ 0) and reports the rate constant k and its implied mean dwell time 1/k.

Physical interpretation of the rate constants: on_rate_constant is the rate of leaving the ON state (≈ k_off in binding/unbinding kinetics), and off_rate_constant is the rate of leaving the OFF state (≈ k_on).

# Default: descriptive stats + exponential fit, consuming events output
frethmm dwell-stats --input ./events/event_details.csv --output-dir ./stats/

# Descriptive statistics only (skip the fit)
frethmm dwell-stats --input ./events/event_details.csv --output-dir ./stats/ --no-fit

# Custom histogram bin count for the fit
frethmm dwell-stats --input ./events/event_details.csv --output-dir ./stats/ --bins 30

dwell-stats subcommand parameters:

Parameter Default Description
--input — Path to event_details.csv (the output of frethmm events), required
--output-dir — Output directory (required)
--bins None Histogram bin count for the exponential fit (default: max(10, n_events // 3))
--no-fit off Skip the exponential fit; emit descriptive statistics only

Two CSV tables are written:

File Description
dwell_stats_summary.csv Single row: pooled ON/OFF counts, mean/median/std/min/max/p25/p75/total, plus fit columns (rate constant, std, mean time, amplitude, n_bins, converged) — blank under --no-fit or when the fit fails
dwell_stats_per_file.csv One row per source file: the extended descriptive block per molecule (for inspecting single-molecule variability)

When the fit is blank: the exponential fit requires ≥ 5 dwell samples per type and a decay-shaped histogram. Constant dwell times, too few events, or a non-decaying distribution yield blank rate columns — the descriptive statistics remain valid.

gui — Graphical Interface

frethmm gui

GUI screenshot (v1.0.0, with batch review grid panel):

FretHMM GUI v1.0.0

GUI Guide

  • Menu bar:
    • File: Add files, add folder, clear all, exit
    • Settings: HMM parameters dialog, language switch (English / Chinese), appearance mode (Light / Dark / System)
    • Help: About dialog
  • File selection: Add .csv / .dat trace files via buttons or menu, or specify an input directory for batch processing
  • State folder batches: Add multiple raw-trace folders, each with its own state count, data mode, and signal column; each folder writes to its adjacent <folder>_output directory by default
  • Parameters panel: States, initial guesses, max iterations, tolerance, workers, data mode, signal column (displayed alongside output panel)
  • GUI workers: The Workers field defaults to 2 in the GUI only and accepts 1–4; the CLI --workers default remains 1. Use 1 for low-memory systems, 2 for a typical laptop (recommended), and 3 or 4 only when higher CPU and memory use is acceptable. Parallelism is between files only: each file is one task.
  • Output options: Checkboxes to select output files — classified.csv / summary.json / report.dat / path.dat / dwell.dat
  • Review Grid section: Click "Generate Review Grid" to classify the selected raw folders and save their paginated review images alongside the corresponding classified CSV files in each <folder>_output directory
  • Manual-review workflow: Inspect each review image, delete unsuitable *_classified.csv files, then select the reviewed <folder>_output directory in the ON/OFF section; results are written to <folder>_output_ONOFF
  • Output collision handling: Before a folder run, choose to overwrite existing output, cancel the whole batch, or create an independent _v2, _v3, … output directory
  • Runtime panel: Collapsible right sidebar (Show/Hide Runtime) showing real-time status, progress, run summary, and last output path
  • Result details: Select a row in the results table to display full fitting metrics (states, log_prob, state means, sigma) and warnings in the right panel
  • Progress bar: Real-time analysis task progress
  • Results table: Fitting results for each file with color coding (green = OK, orange = warning, red = error)
  • Theme toggle: Switch between Light / Dark / System via Settings menu or 🌓 button in the title bar
  • Bilingual support: Real-time English / Chinese UI switching via Settings → Language
  • Threaded processing: All analysis runs in a background thread with cancel support (Cancel button). After cancellation, no further files are scheduled; files already active are allowed to complete and their *_classified.csv outputs are kept. A cancelled Review Grid run does not write a partial review image.
  • Log panel: Color-coded log output (blue headers, orange warnings, red errors, green completion)
  • Status bar: Current status and version number at the bottom

Visualization

Review Grid

Review Grid is a batch visualization tool for manual quality inspection. It classifies all traces in a directory via HMM and renders the results as a paginated multi-panel image.

How it works:

  1. Scan all trace files in the input directory
  2. Run HMM state classification on each file
  3. Arrange results in a rows × cols grid of panels
  4. Each panel overlays raw signal (gray thin line) with classified signal (red thick line)
  5. Panel titles show filename, log-likelihood, and state means
  6. Traces with fitting warnings are marked with orange borders for quick identification

Output examples:

review.png                     # Single page (trace count ≤ rows × cols)
review_page_01.png             # Auto-numbered pages when traces exceed one page
review_page_02.png

Typical workflow:

# 1. Quick quality review of all traces
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 --rows 4 --cols 8

# 2. Re-process problematic files individually
frethmm run --files traces/bad_trace.csv --states 3 --guesses "0.1,0.5,0.9" -v

# 3. After review passes, batch export full results
frethmm run --input-dir ./traces/ --states 2 --workers 4 --output-dir ./results/

Transition Density Plot (TDP)

TDP aggregates transition information from all molecules (via *report.dat files) and renders a scatter density plot.

Chart composition:

  • X-axis: Start state mean
  • Y-axis: Stop state mean
  • Point size and color: Encode transition count (using hot colormap; warmer = higher frequency)
  • Diagonal dashed line: Self-transition reference

--states N filtering: When mixing datasets with different state counts, this parameter keeps only the top-N states per molecule (by total transition frequency) for cross-dataset comparison.

Rate analysis: Beyond visualization, FretHMM provides a fit_gaussian_to_rates() programming interface for Gaussian fitting on transition rate distributions between specific state pairs, extracting mean rate and standard deviation.

Algorithm Hardening (Multi-start + BIC Model Selection)

Baum-Welch is sensitive to the initial state means, so a single fit can land in a poor local optimum. FretHMM ships two algorithm-hardening features to make results more stable and less reliant on manual tuning.

Multi-start fitting (--n-init)

For every fit, FretHMM runs Baum-Welch --n-init times (default 10) from deterministic initial means and keeps the result with the highest log-likelihood.

  • Start 0 always uses the legacy evenly-spaced default means, so --n-init 1 reproduces the historical single-fit output exactly.
  • Starts 1..n-1 perturb the default means with fixed-seed jitter (the seed depends only on the configuration, never on wall-clock time), so repeated runs on the same input are byte-for-byte reproducible.
  • Use --n-init 1 to opt out of multi-start entirely (fastest, legacy behaviour).
# Default 10 starts (recommended for stability)
frethmm run --files trace.csv --states 3

# Reproduce the legacy single-fit
frethmm run --files trace.csv --states 3 --n-init 1

BIC model selection (--states auto)

When you don't know the state count, pass --states auto and FretHMM will scan the range [--min-states, --max-states] (default 2..6), fit each candidate with the full multi-start procedure, and pick the one with the lowest Bayesian Information Criterion (BIC = k·ln(n) − 2·log_prob, where k is the free-parameter count of a tied-covariance Gaussian HMM and n is the number of frames).

# Auto-select the state count via BIC over 2..5 states
frethmm run --input-dir ./traces/ --states auto --min-states 2 --max-states 5 --workers 4

When auto-selection runs, *_summary.json records the chosen BIC, AIC, and a model_candidates table listing every candidate's n_states / log_prob / bic / aic so you can audit the decision.

{
  "n_states": 3,
  "bic": -12259.40,
  "aic": -12330.18,
  "n_init": 10,
  "model_candidates": [
    {"n_states": 2, "log_prob": 5000.1, "bic": -9800.2, "aic": -9870.5},
    {"n_states": 3, "log_prob": 6193.8, "bic": -12259.4, "aic": -12330.2},
    {"n_states": 4, "log_prob": 6195.0, "bic": -12240.8, "aic": -12320.1}
  ]
}

Note on --guesses: initial guesses are ignored with --states auto because the per-candidate state count varies; multi-start provides the initialization diversity instead.

Data Filtering

Lowest-State Window Filtering

Background: In single-molecule fluorescence experiments, a persistent lowest-signal period can indicate photobleaching or fluorophore inactivation. Retaining data after that period can distort classification of the biologically relevant states.

Two-pass fitting workflow:

┌──────────────┐     ┌──────────────┐     ┌──────────────┐
│  1st Pass    │ ──→ │  Locate      │ ──→ │  Trim data   │
│  HMM on full │     │  lowest      │     │  at cutoff   │
│  trace       │     │  state run   │     │  point       │
└──────────────┘     └──────────────┘     └──────────────┘
                                                 │
                                                 ▼
                                          ┌──────────────┐
                                          │  2nd Pass    │
                                          │  HMM on      │
                                          │  trimmed data│
                                          └──────────────┘
  1. First pass classification: Fit HMM on the full trace to obtain the Viterbi state path.
  2. Identify lowest state: Find the state with the lowest mean value.
  3. Forward scan: Starting at 0 s, locate the first uninterrupted lowest-state run that reaches --low-state-tail-trim-seconds.
  4. Trim data: Retain samples through run start + configured seconds (inclusive) and discard every later sample.
  5. Second pass classification: Re-run HMM on the retained trace for the final classification.

Note: If the lowest state never persists beyond the threshold duration, no trimming is performed and the first-pass result is retained.

CLI examples:

# Single file: retain through the first lowest-state window that reaches 5 seconds
frethmm run --files trace.csv --states 2 --low-state-tail-trim-seconds 5.0

# Batch: 3-second threshold, 4 parallel workers
frethmm run --input-dir ./traces/ --states 3 --low-state-tail-trim-seconds 3.0 --workers 4

# Combined with Review Grid: trim then review
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
    --low-state-tail-trim-seconds 5.0 --rows 4 --cols 8

Output metadata: When trimming is active, *_summary.json records these additional fields:

{
  "low_state_tail_trim_seconds": 5.0,
  "low_state_tail_cutoff_time": 47.3,
  "low_state_tail_kept_frames": 473
}
  • low_state_tail_trim_seconds: The configured trim threshold
  • low_state_tail_cutoff_time: The actual cutoff time point (null if trimming was not triggered)
  • low_state_tail_kept_frames: Number of frames retained after trimming

GUI usage: The GUI's Lowest-state window (sec) field defaults to 250. Leave it blank to disable filtering. The setting applies to all direct runs and folder batches, including Review Grid generation.

Input Format

The program auto-detects file format (header presence, delimiter type, column count). Two data modes are supported:

Single-channel mode (CSV, with header):

Time,channel1
0,2820
1,2884
2,2570

For multi-column signals, use --signal-column to select a specific column:

Time,channel1,channel2
0,2884,-5096
1,2884,1289

--signal-column 1 uses the channel1 column, --signal-column 2 uses channel2.

Dual-channel Donor/Acceptor mode (whitespace/tab delimited, 3 columns, no header):

<time>  <donor>  <acceptor>

In this mode, FRET efficiency A/(D+A) is automatically computed as the HMM input signal.

Output Files

Each input file generates the following outputs:

File Format Description
*_classified.csv CSV Primary output: time, classified_mean — the idealized trace
*_summary.json JSON State means, frame fractions, transition matrix, dwell statistics, trim metadata, warnings
*report.dat Text Model parameters (state count, means, sigma, transition probability matrix)
*path.dat TSV Raw signal channels + FRET signal + classified signal per frame
*dwell.dat TSV Dwell time table: <start_mean> <stop_mean> <frames_lasted> per dwell segment

Project Structure

FretHMM/
├── frethmm/
│   ├── __init__.py              # Version info
│   ├── app/
│   │   ├── cli.py               # CLI entry point (run / tdp / review-grid / gui)
│   │   ├── gui.py               # CustomTkinter GUI
│   │   └── i18n.py              # Internationalization (English / Chinese, 138 keys)
│   ├── assets/
│   │   ├── frethmm.ico          # Application icon
│   │   └── frethmm_logo.png     # Application logo
│   ├── core/
│   │   ├── io.py                # File I/O (trace reading + report output)
│   │   ├── model.py             # HMM engine (Baum-Welch + Viterbi + lowest-state window filtering)
│   │   ├── batch.py             # Multi-process batch processor
│   │   └── postprocess.py       # Classified signal + dwell segments + transition stats
│   ├── domain/
│   │   └── models.py            # Data models (Config / Trace / Result / ExportOptions)
│   ├── formats/
│   │   └── report_parser.py     # report.dat parser
│   ├── legacy/
│   │   └── report_parser.py     # Legacy report format parser
│   └── viz/
│       ├── review_grid.py       # Review Grid batch visualization (paginated layout)
│       └── tdp.py               # Transition Density Plot + Gaussian rate fitting
├── tests/
│   ├── fixtures/                # Regression test reference data
│   ├── test_io.py               # I/O and report parsing tests
│   ├── test_review_grid.py      # Review Grid visualization tests
│   └── test_golden.py           # CLI regression tests
├── docs/
│   ├── images/                  # Screenshots
│   └── FretHMM-refactor-plan.md # Development roadmap
├── pyproject.toml               # Project configuration
├── build_exe.py                 # PyInstaller build script
├── frethmm.spec                 # PyInstaller spec file
├── LICENSE                      # MIT License
└── README.md

Testing

pip install -e ".[dev]"
pytest tests/ -v

Reproducibility

Each successful CLI analysis writes a timestamped frethmm_run_manifest_*.json next to its primary outputs. The manifest records the command, fit parameters, input/output file metadata, FretHMM version, Python version, and runtime dependency versions without copying experimental data. This also applies to --classified-only, which suppresses auxiliary classification exports but retains the run manifest.

The committed fixtures under tests/data/ are small, synthetic or de-identified CSV and legacy report samples. They are sufficient for the core regression suite and do not include raw images or ND2 files. FretHMM currently starts from exported trajectory files (.csv, .dat, .txt, .tsv); ND2 image-to-trajectory processing remains an upstream workflow.

Packaging as Executable

# Install the explicit build dependency first
pip install -e ".[gui]"

# Directory mode (default, produces dist/FretHMM/ directory)
python build_exe.py

# Single-file mode (produces dist/FretHMM.exe, easy to distribute)
python build_exe.py --onefile

The build produces a standalone Windows GUI executable — no Python installation required. Directory-mode builds also create dist/FretHMM.zip, a SHA-256 sidecar, and a JSON release manifest. The single-file build produces dist/FretHMM.exe; validate it without launching the UI with dist/FretHMM.exe --version.

Changelog

v1.8.0 (2026-09-29)

Data preprocessing, post-fit state consolidation, diagnostics, and GUI binding:

  • Preprocessing module (frethmm.core.preprocess):
    • Robust scale estimation using Median Absolute Deviation (estimate_robust_sigma).
    • Outlier spike detection and cleaning (--remove-spikes, --spike-threshold-sigma).
    • Initial acquisition/shutter artifact trimming (--trim-initial-artifacts, --max-initial-artifact-frames).
    • Edge-preserving median smoothing window (--smooth-window).
  • Postprocessing module (frethmm.core.postprocess):
    • Minimum dwell frame segment merging (--min-dwell-frames) to suppress transient noise flickers.
    • Near-identical state merging (--merge-state-threshold, --merge-state-sigma-factor).
    • End-to-end post-fit classification cleanup (cleanup_classification_result).
  • Fit quality diagnostics (frethmm.core.metrics):
    • Quantitative metrics: minimum and mean SNR, RMSE, MAE, $R^2$, and state occupancies recorded in *_summary.json.
  • Review grid enhancement:
    • Added --files parameter to review-grid to allow visual grid generation for specific files without folder restructuring.
  • GUI controls & bilingual i18n:
    • Main parameters panel quick controls (Trim Start Artifacts, Clean Spikes, Min Dwell frames).
    • Dedicated Clean & Filter Options dialog accessible from main panel, settings menu, and parameters dialog.
    • Full English/Chinese bilingual support with dynamic switching.

v1.7.1 (2026-09-07)

Plot-ready per-molecule table from events:

  • New input_plot.csv written alongside the other three event tables (CLI and GUI): one row per source file with source_file, ON_events, OFF_events, Fluorescence_strength (the file-level state_value_range), and Duration_time (the last event's end_time — for ON-then-bleach traces, the observable duration before photobleaching). Designed as a direct input for downstream visualisation.
  • New summarize_plot_input + PLOT_INPUT_FIELDS in frethmm.core.events, shared by CLI and GUI; the run manifest now lists input_plot.csv among the outputs.

v1.7.0 (2026-09-07)

Per-event fluorescence amplitude in event_details.csv:

  • New state_value_range column (file-level value repeated on every event row of a source file):
    • Files with multiple included events: max(state_value) − min(state_value) over included events — the real ON/OFF fluorescence contrast (excluded/omitted photobleach tails do not participate).
    • Files whose only included event is a single ON event (e.g. molecule stays ON then photobleaches): ON level − minimum classified value in the file — the signal amplitude above the bleached baseline.
    • Degenerate cases (constant-ON trace, single OFF, no included events): 0.0.
  • Backward compatible: read_event_details parses pre-v1.7 event files without the column (defaults to 0.0), so dwell-stats keeps consuming older outputs.
  • CLI and GUI events outputs pick up the column automatically (both share DETAIL_FIELDS).

v1.6.0 (2026-08-01)

Windows GUI process-pool parallelism with safe defaults and cancel semantics:

  • GUI workers: GUI Workers default to 2 and accept 1–4; CLI --workers remains default 1. Values 3 and 4 prompt about higher CPU/memory use. Parallelism is between files only.
  • True Windows multi-process pool: classification and review-grid jobs run through a spawn-compatible process pool so Workers=2/4 actually accelerate batch work on Windows.
  • Cancel / close safety: after cancel, no further files are scheduled; in-flight files may finish and keep their *_classified.csv outputs. A cancelled review-grid run does not publish a partial PNG. Close after a successful submit keeps completed status.
  • Release artifact: Windows single-file FretHMM.exe is published with a SHA-256 checksum and JSON version manifest.

v1.5.0 (2026-07-29)

Batch review workflow, multi-state event statistics, and lowest-state window filtering:

  • GUI workflow: supports multiple raw input folders with automatic per-folder <folder>_output directories, per-folder review grids, collision choices (overwrite / cancel / versioned output), and a dedicated reviewed-classification folder input for ON/OFF analysis. ON/OFF results are written to <folder>_output_ONOFF by default.
  • Multi-state ON/OFF statistics: 2-state traces report high-state ON and recovered low-state OFF events. For 3 or more states, every non-lowest activity stage is analysed independently; a descent is an OFF event only if that same stage recovers. Terminal non-recovering low segments are omitted.
  • Lowest-state window filtering: replaces terminal-only trimming with a forward scan from 0 s for the first continuous lowest-state window meeting the configured duration (default 250 s), then refits only the retained data.
  • Release artifact: Windows single-file FretHMM.exe is published with a SHA-256 checksum and JSON version manifest.

v1.4.0 (release candidate)

Dwell-time statistics and rate-constant fitting:

  • New dwell-stats subcommand: consumes event_details.csv (output of events) and writes dwell_stats_summary.csv + dwell_stats_per_file.csv with extended descriptive statistics (median, std, min/max, p25/p75) for ON and OFF dwell times.
  • Exponential rate-constant fit: single-exponential A·exp(-k·t) fit of the pooled dwell-time histogram (histogram + scipy.optimize.curve_fit, bounded k ≥ 0), reporting the rate constant and implied mean dwell time for ON and OFF. Skippable via --no-fit.
  • New modules: frethmm/core/dwell_stats.py (describe_durations, fit_exponential_dwell, extended summaries) and frethmm/formats/event_details_parser.py (reverse-parse event_details.csv).
  • No regression: events command and events.py are untouched; dwell-stats is a pure downstream consumer.
  • Tests: new test_dwell_stats.py (12 tests) and test_dwell_stats_cli.py (3 end-to-end tests), all self-contained.
  • Release reproducibility: CLI analysis commands now emit a timestamped run manifest; committed de-identified fixtures replace workspace-dependent skipped regression tests; Windows CI builds and version-smoke-tests the GUI bundle.
  • Tail trimming correction: low-state trimming applies only to a persistent terminal low state, avoiding accidental removal of valid low-state segments in the middle of a trace.

v1.3.0 (2026-06-15)

ON/OFF event analysis brought in-package:

  • New events subcommand: extract ON/OFF events from *_classified.csv files and write event_details.csv, event_summary.csv, and event_stats_overall.csv.
  • N-state generalisation: the highest-mean state is ON, all others are OFF (2-state traces behave exactly like the legacy "high = ON" rule).
  • Terminal-OFF exclusion: a long final OFF run (default ≥ 100 s) is flagged excluded and omitted from dwell statistics while still listed.
  • New modules: frethmm/core/events.py (event detection + summaries), frethmm/formats/classified_parser.py (reverse-parse *_classified.csv), and find_classified_files in frethmm/core/io.py.
  • Tests: new test_events.py, test_classified_parser.py, and test_events_cli.py with self-contained synthetic traces (no external sample dependency).

v1.2.0 (2026-06-15)

Algorithm hardening — multi-start fitting and BIC-based state-count selection:

  • Multi-start fitting (--n-init, default 10): deterministic multi-start Baum-Welch that keeps the best log-likelihood. Start 0 reproduces the legacy single-fit, so --n-init 1 is byte-compatible with v1.1.
  • BIC model selection (--states auto with --min-states/--max-states): scan a state-count range, fit each candidate with multi-start, and pick the lowest BIC.
  • New metrics module (frethmm.core.metrics): compute_aic, compute_bic, count_gaussian_hmm_params.
  • Summary JSON: records n_init, best_start_index, bic, aic, and model_candidates when multi-start or auto-selection is active (omitted for single-start legacy runs).
  • GUI: new "Auto-select states (BIC)" checkbox with min/max range, plus an n_init field in the parameters panel and dialog; folder-batch jobs support auto-selection.
  • Tests: new test_multistart.py and test_model_selection.py with synthetic-trace fixtures; golden tests pinned to --n-init 1 to preserve byte-exact regression.

v1.1.0 (2026-06-09)

Documentation and release infrastructure update:

  • Enhanced README with detailed visualization docs (Review Grid, TDP) and data filtering workflow (low-state tail trimming)
  • Added LICENSE file (MIT)
  • Version bump to 1.1.0 in __init__.py and pyproject.toml
  • Updated .gitignore with additional exclusion rules

v1.0.0 (2026-06-04)

Batch visual review release:

  • Batch review grid CLI: New review-grid subcommand for batch classification with paginated grid overview
  • GUI review grid: Dedicated Review Grid section in GUI for generating paginated review images
  • Paginated layout: Configurable rows × cols grid, suitable for quick visual screening of 2-state, 3-state batches
  • Visual review enhancements: Each panel overlays raw signal with classified trace; shows filename, log_prob, and state means

v0.6.0 (2026-06-01)

GUI layout optimization and packaging improvements:

  • Layout restructure: Removed ScrollableFrame; parameters and output panels side-by-side
  • Collapsible runtime panel: Right sidebar defaults hidden; toggle via Show/Hide Runtime button
  • Application icon: Added frethmm.ico and frethmm_logo.png
  • Window sizing: Default 1280×720, minimum 1180×660
  • PyInstaller slimming: Reduced EXE size by excluding unused packages
  • --onefile mode: Added single-file EXE build support

v0.5.0 (2026-06-01)

GUI stability and export options:

  • ExportOptions: Fine-grained control over output file types
  • GUI output checkboxes: Select which files to generate
  • Worker error handling: Full traceback logging
  • Global exception hook: Error dialogs even in console=False EXE mode

v0.4.0 (2026-06-01)

CustomTkinter migration and CLI enhancements:

  • CustomTkinter migration: Full dark/light/system theme support
  • --classified-only: CLI flag to output only *_classified.csv
  • Folder batch panel: Per-folder state count, data mode, and signal column
  • Runtime panel: Real-time status, progress, and result details

v0.3.0 (2026-06-01)

Project restructured as FretHMM with modular architecture:

  • Modular packages: core / domain / app / formats / legacy / viz
  • CLI with single-file and directory batch modes, multiprocessing support
  • Full GUI with menu bar, parameters dialog, bilingual support, threaded analysis
  • Default outputs: *_classified.csv + *_summary.json
  • Additional outputs: report.dat, path.dat, dwell.dat
  • TDP visualization + Gaussian rate fitting
  • PyInstaller Windows executable build
  • Regression test coverage (I/O, report parsing, end-to-end CLI)

v0.2.0 (2026-06-01)

GUI major update:

  • Menu bar (File / Settings / Help) and parameters dialog
  • Bilingual support (i18n): English/Chinese real-time switching
  • Modern UI: platform-adaptive fonts, custom ttk theme, color-coded logs
  • Startup optimization: lazy loading of heavy libraries
  • Warning handling: captured and displayed with orange highlighting

v0.1.0 (2026-05-30)

Initial release:

  • Complete HMM analysis pipeline (Baum-Welch training + Viterbi decoding)
  • CLI tool (run / tdp / gui subcommands)
  • tkinter GUI (file selection, parameters panel, progress bar, results table, log panel)
  • Multi-process batch processing
  • TDP visualization
  • PyInstaller GUI packaging script

License

MIT License

About

Single-molecule time series Hidden Markov Model (HMM) state classification tool. Inspired by HaMMy, rewritten from scratch in Python, supporting cross-platform operation, batch processing, and GUI interaction.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages