A single-molecule time-series Hidden Markov Model (HMM) state classification tool. Inspired by HaMMy, rewritten from scratch in Python with cross-platform support, batch processing, and a full GUI.
| Feature | Description |
|---|---|
| HMM Engine | Baum-Welch training + Viterbi decoding (via hmmlearn), with customizable initial guesses |
| Data Modes | Auto-detect / Single-channel signal / Dual-channel Donor-Acceptor (auto-computes FRET efficiency) |
| Batch Processing | Multi-file parallel processing (ProcessPoolExecutor), directory scanning with multi-worker support |
| Review Grid | Batch classification + paginated multi-panel PNG visual review for quick quality screening |
| Lowest-State Window Filtering | Two-pass HMM fitting that scans forward for the first persistent lowest-state window and discards later signal |
| CLI | Six subcommands: run, tdp, review-grid, events, dwell-stats, gui |
| GUI | CustomTkinter interface with dark/light themes, English/Chinese switching, threaded background analysis, and batch review grid export |
| Output Formats | *_classified.csv, *_summary.json, *report.dat, *path.dat, *dwell.dat (selectable in GUI) |
| TDP | Transition Density Plot visualization + Gaussian rate fitting |
| Packaging | PyInstaller one-click build for Windows executables (directory mode / --onefile mode) |
git clone https://github.com/Caizhaohui/FretHMM.git
cd FretHMM
pip install -e .Requirements:
- Python >= 3.10
- NumPy >= 1.24
- SciPy >= 1.10
- hmmlearn >= 0.3.0
- matplotlib >= 3.7 (required for TDP and Review Grid visualization)
- customtkinter >= 5.2.0 (required for GUI)
Optional dependencies:
pip install -e ".[dev]" # Install pytest testing framework
pip install -e ".[gui]" # Install PyInstaller packaging toolFretHMM provides six subcommands: run (HMM analysis), review-grid (visual review), tdp (transition density plot), events (ON/OFF event analysis), dwell-stats (dwell-time statistics + rate-constant fit), and gui (graphical interface).
# Single file analysis (2 states, auto-detect data format)
frethmm run --files trace.csv --states 2 --output-dir ./results/
# Batch process all trace files in a directory (4 parallel workers)
frethmm run --input-dir ./traces/ --states 5 --workers 4 --output-dir ./results/
# Process multiple files at once
frethmm run --files trace1.csv trace2.csv trace3.csv --states 3 --output-dir ./results/
# Provide initial guesses (useful when state spacing is small)
frethmm run --files data.csv --states 2 --guesses "0.3,0.7"
# Specify single-channel mode and signal column
frethmm run --files data.csv --states 2 --mode single_channel --signal-column 1
# Use lowest-state window filtering (keep through the first 5-second lowest-state window, then re-classify)
frethmm run --files trace.csv --states 2 --low-state-tail-trim-seconds 5.0
# Output only classified.csv
frethmm run --files data.csv --states 2 --classified-only
# Verbose mode (show all warnings)
frethmm run --files data.csv --states 3 -vrun subcommand parameters:
| Parameter | Default | Description |
|---|---|---|
--files |
— | One or more trace file paths (mutually exclusive with --input-dir, required) |
--input-dir |
— | Input directory to scan for trace files (mutually exclusive with --files, required) |
--output-dir |
— | Output directory (defaults to input file directory) |
--states |
2 | Number of HMM states, or auto to pick via BIC (see Model selection) |
--guesses |
None | Comma-separated initial signal guesses; count must match --states (ignored with --states auto) |
--max-iter |
500 | Maximum Baum-Welch iterations |
--tol |
1e-4 | Convergence tolerance |
--workers |
1 | Number of parallel workers (>1 enables multiprocessing) |
--mode |
auto | Data mode: auto / paired_channel / single_channel |
--signal-column |
1 | 1-based signal column index after Time for single_channel mode |
--low-state-tail-trim-seconds |
None | Lowest-state window duration in seconds; the option name is retained for compatibility (see Data Filtering) |
--n-init |
10 | Number of deterministic multi-start Baum-Welch runs; best log-likelihood wins (use 1 to reproduce the legacy single-fit) |
--min-states |
2 | Minimum state count for BIC selection (only with --states auto) |
--max-states |
6 | Maximum state count for BIC selection (only with --states auto) |
--remove-spikes |
off | Enable outlier spike detection and cleaning via rolling median and robust MAD scale estimation |
--spike-threshold-sigma |
5.0 | Spike detection threshold in units of robust sigma |
--trim-initial-artifacts |
off | Automatically detect and trim initial acquisition/shutter artifacts (e.g. Frame 0 surge) |
--max-initial-artifact-frames |
5 | Maximum initial frames allowed to be trimmed as artifacts |
--smooth-window |
None | Edge-preserving median smoothing window (odd integer) to suppress shot noise before HMM fitting |
--min-dwell-frames |
1 | Minimum state dwell duration in frames; shorter transient noise flickers are merged into neighbors |
--merge-state-threshold |
None | Absolute difference threshold below which adjacent fitted states are merged |
--merge-state-sigma-factor |
None | Relative scale threshold (k * sigma) below which adjacent fitted states are merged |
--classified-only |
off | Output only *_classified.csv, skip summary/report/path/dwell |
-v / --verbose |
off | Verbose output, show all warnings and fit quality diagnostics (SNR, RMSE, R^2) |
Batch processing notes:
--input-dirscans all.csv,.dat,.txt,.tsvfiles, automatically skipping output files (*report.dat,*path.dat,*dwell.dat,*_classified.csv,*_summary.json)--workers Nenables multi-process parallelism; N should not exceed CPU core count- Individual file errors do not interrupt the overall batch; errors are printed to the terminal
# Basic: generate a 4×4 grid of 2-state traces
frethmm review-grid --input-dir ./traces/ --output review.png --states 2
# Custom grid layout
frethmm review-grid --input-dir ./traces/ --output review.png --states 3 --rows 5 --cols 6
# With initial guesses and output directory for classified CSVs
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
--guesses "0.2,0.8" --output-dir ./classified/
# Combined with lowest-state window filtering
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
--low-state-tail-trim-seconds 5.0
# Accelerate with 4 parallel workers
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
--workers 4 --rows 4 --cols 8review-grid subcommand parameters:
| Parameter | Default | Description |
|---|---|---|
--input-dir |
— | Input trace file directory (either --input-dir or --files is required) |
--files |
— | Specify one or more trace files (either --files or --input-dir is required) |
--output |
— | Output PNG path, e.g. review.png (required) |
--output-dir |
None | Optional directory for classified CSV side outputs |
--states |
2 | Number of HMM states, or auto to pick via BIC |
--guesses |
None | Comma-separated initial signal guesses (ignored with --states auto) |
--max-iter |
500 | Maximum Baum-Welch iterations |
--tol |
1e-4 | Convergence tolerance |
--workers |
1 | Number of parallel workers |
--mode |
auto | Data mode: auto / paired_channel / single_channel |
--signal-column |
1 | Signal column index for single_channel mode |
--low-state-tail-trim-seconds |
None | Lowest-state window duration in seconds (legacy option name) |
--n-init |
10 | Deterministic multi-start count (use 1 for legacy single-fit) |
--min-states |
2 | Minimum state count for BIC selection (only with --states auto) |
--max-states |
6 | Maximum state count for BIC selection (only with --states auto) |
--rows |
4 | Panel rows per page |
--cols |
4 | Panels per row |
Pagination: When the number of traces exceeds rows × cols, multiple page images are generated automatically (e.g., review_page_01.png, review_page_02.png). Each panel overlays the raw signal (gray) with the HMM classified signal (red). The title shows filename, log-likelihood, and state means. Traces with fitting warnings are highlighted with orange borders.
# Generate TDP from report files (interactive window)
frethmm tdp --input-dir ./results/ --exposure 0.1
# Save to file
frethmm tdp --input-dir ./results/ --exposure 0.1 --output tdp.png
# Show only top N states (sorted by transition frequency)
frethmm tdp --input-dir ./results/ --exposure 0.1 --states 3 --output tdp.pngtdp subcommand parameters:
| Parameter | Default | Description |
|---|---|---|
--input-dir |
— | Directory containing *report.dat files (required) |
--exposure |
0.1 | Frame exposure time in seconds, used for rate calculations |
--states |
None | Show only top N states (sorted by transition frequency) |
--output |
None | Output image path (e.g., tdp.png); opens interactive window if not specified |
Extract discrete ON/OFF events from *_classified.csv files (the primary output of run). In a 2-state trace, high fluorescence is ON and a low segment is OFF only when high fluorescence later recovers (high → low → high); a terminal low segment is permanent loss of activity and is omitted. For 3+ states, each non-lowest state is analysed as an independent stage, with a descent counted as OFF only when that original stage recovers.
# Batch: scan a directory of *_classified.csv files
frethmm events --input-dir ./results/ --output-dir ./events/
# Process specific files
frethmm events --files trace1_classified.csv trace2_classified.csv --output-dir ./events/
# Legacy compatibility option; terminal low segments are omitted, not counted as OFF
frethmm events --input-dir ./results/ --tail-off-threshold-seconds 250 --output-dir ./events/events subcommand parameters:
| Parameter | Default | Description |
|---|---|---|
--input-dir |
— | Directory of *_classified.csv files (mutually exclusive with --files, required) |
--files |
— | One or more *_classified.csv paths (mutually exclusive with --input-dir, required) |
--output-dir |
— | Output directory (required) |
--tail-off-threshold-seconds |
100.0 | Legacy compatibility option; terminal low segments are omitted rather than recorded as OFF |
Four CSV tables are written per run:
| File | Description |
|---|---|
event_details.csv |
One row per event: source file, type (ON/OFF), index, state value, start/end time and frame, duration, excluded flag, plus state_value_range — the file-level fluorescence amplitude: max − min over included events; for a file whose only included event is a single ON (e.g. ON followed by an omitted photobleach tail), ON level − minimum classified value (signal above the bleached baseline) |
event_summary.csv |
One row per source file: ON/OFF counts, total and mean dwell times |
event_stats_overall.csv |
Aggregate across all files: event counts, total/mean ON and OFF times |
input_plot.csv |
One row per source file for downstream visualisation: source_file, ON_events, OFF_events, Fluorescence_strength (= state_value_range), Duration_time (= last event's end_time; for ON-then-bleach traces this is the observable duration before photobleaching) |
Consumes the event_details.csv produced by events and computes the deeper descriptive statistics single-molecule analysis needs: median, standard deviation, min/max, and 25th/75th percentiles for ON and OFF dwell times, pooled across all molecules. Optionally fits a single exponential A·exp(-k·t) to each dwell-time distribution (histogram + scipy.optimize.curve_fit, bounded so k ≥ 0) and reports the rate constant k and its implied mean dwell time 1/k.
Physical interpretation of the rate constants: on_rate_constant is the rate of leaving the ON state (≈ k_off in binding/unbinding kinetics), and off_rate_constant is the rate of leaving the OFF state (≈ k_on).
# Default: descriptive stats + exponential fit, consuming events output
frethmm dwell-stats --input ./events/event_details.csv --output-dir ./stats/
# Descriptive statistics only (skip the fit)
frethmm dwell-stats --input ./events/event_details.csv --output-dir ./stats/ --no-fit
# Custom histogram bin count for the fit
frethmm dwell-stats --input ./events/event_details.csv --output-dir ./stats/ --bins 30dwell-stats subcommand parameters:
| Parameter | Default | Description |
|---|---|---|
--input |
— | Path to event_details.csv (the output of frethmm events), required |
--output-dir |
— | Output directory (required) |
--bins |
None | Histogram bin count for the exponential fit (default: max(10, n_events // 3)) |
--no-fit |
off | Skip the exponential fit; emit descriptive statistics only |
Two CSV tables are written:
| File | Description |
|---|---|
dwell_stats_summary.csv |
Single row: pooled ON/OFF counts, mean/median/std/min/max/p25/p75/total, plus fit columns (rate constant, std, mean time, amplitude, n_bins, converged) — blank under --no-fit or when the fit fails |
dwell_stats_per_file.csv |
One row per source file: the extended descriptive block per molecule (for inspecting single-molecule variability) |
When the fit is blank: the exponential fit requires ≥ 5 dwell samples per type and a decay-shaped histogram. Constant dwell times, too few events, or a non-decaying distribution yield blank rate columns — the descriptive statistics remain valid.
frethmm guiGUI screenshot (v1.0.0, with batch review grid panel):
- Menu bar:
- File: Add files, add folder, clear all, exit
- Settings: HMM parameters dialog, language switch (English / Chinese), appearance mode (Light / Dark / System)
- Help: About dialog
- File selection: Add
.csv/.dattrace files via buttons or menu, or specify an input directory for batch processing - State folder batches: Add multiple raw-trace folders, each with its own state count, data mode, and signal column; each folder writes to its adjacent
<folder>_outputdirectory by default - Parameters panel: States, initial guesses, max iterations, tolerance, workers, data mode, signal column (displayed alongside output panel)
- GUI workers: The Workers field defaults to
2in the GUI only and accepts1–4; the CLI--workersdefault remains1. Use1for low-memory systems,2for a typical laptop (recommended), and3or4only when higher CPU and memory use is acceptable. Parallelism is between files only: each file is one task. - Output options: Checkboxes to select output files — classified.csv / summary.json / report.dat / path.dat / dwell.dat
- Review Grid section: Click "Generate Review Grid" to classify the selected raw folders and save their paginated review images alongside the corresponding classified CSV files in each
<folder>_outputdirectory - Manual-review workflow: Inspect each review image, delete unsuitable
*_classified.csvfiles, then select the reviewed<folder>_outputdirectory in the ON/OFF section; results are written to<folder>_output_ONOFF - Output collision handling: Before a folder run, choose to overwrite existing output, cancel the whole batch, or create an independent
_v2,_v3, … output directory - Runtime panel: Collapsible right sidebar (Show/Hide Runtime) showing real-time status, progress, run summary, and last output path
- Result details: Select a row in the results table to display full fitting metrics (states, log_prob, state means, sigma) and warnings in the right panel
- Progress bar: Real-time analysis task progress
- Results table: Fitting results for each file with color coding (green = OK, orange = warning, red = error)
- Theme toggle: Switch between Light / Dark / System via Settings menu or 🌓 button in the title bar
- Bilingual support: Real-time English / Chinese UI switching via Settings → Language
- Threaded processing: All analysis runs in a background thread with cancel support (Cancel button). After cancellation, no further files are scheduled; files already active are allowed to complete and their
*_classified.csvoutputs are kept. A cancelled Review Grid run does not write a partial review image. - Log panel: Color-coded log output (blue headers, orange warnings, red errors, green completion)
- Status bar: Current status and version number at the bottom
Review Grid is a batch visualization tool for manual quality inspection. It classifies all traces in a directory via HMM and renders the results as a paginated multi-panel image.
How it works:
- Scan all trace files in the input directory
- Run HMM state classification on each file
- Arrange results in a
rows × colsgrid of panels - Each panel overlays raw signal (gray thin line) with classified signal (red thick line)
- Panel titles show filename, log-likelihood, and state means
- Traces with fitting warnings are marked with orange borders for quick identification
Output examples:
review.png # Single page (trace count ≤ rows × cols)
review_page_01.png # Auto-numbered pages when traces exceed one page
review_page_02.png
Typical workflow:
# 1. Quick quality review of all traces
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 --rows 4 --cols 8
# 2. Re-process problematic files individually
frethmm run --files traces/bad_trace.csv --states 3 --guesses "0.1,0.5,0.9" -v
# 3. After review passes, batch export full results
frethmm run --input-dir ./traces/ --states 2 --workers 4 --output-dir ./results/TDP aggregates transition information from all molecules (via *report.dat files) and renders a scatter density plot.
Chart composition:
- X-axis: Start state mean
- Y-axis: Stop state mean
- Point size and color: Encode transition count (using
hotcolormap; warmer = higher frequency) - Diagonal dashed line: Self-transition reference
--states N filtering: When mixing datasets with different state counts, this parameter keeps only the top-N states per molecule (by total transition frequency) for cross-dataset comparison.
Rate analysis: Beyond visualization, FretHMM provides a fit_gaussian_to_rates() programming interface for Gaussian fitting on transition rate distributions between specific state pairs, extracting mean rate and standard deviation.
Baum-Welch is sensitive to the initial state means, so a single fit can land in a poor local optimum. FretHMM ships two algorithm-hardening features to make results more stable and less reliant on manual tuning.
For every fit, FretHMM runs Baum-Welch --n-init times (default 10) from deterministic initial means and keeps the result with the highest log-likelihood.
- Start 0 always uses the legacy evenly-spaced default means, so
--n-init 1reproduces the historical single-fit output exactly. - Starts 1..n-1 perturb the default means with fixed-seed jitter (the seed depends only on the configuration, never on wall-clock time), so repeated runs on the same input are byte-for-byte reproducible.
- Use
--n-init 1to opt out of multi-start entirely (fastest, legacy behaviour).
# Default 10 starts (recommended for stability)
frethmm run --files trace.csv --states 3
# Reproduce the legacy single-fit
frethmm run --files trace.csv --states 3 --n-init 1When you don't know the state count, pass --states auto and FretHMM will scan the range [--min-states, --max-states] (default 2..6), fit each candidate with the full multi-start procedure, and pick the one with the lowest Bayesian Information Criterion (BIC = k·ln(n) − 2·log_prob, where k is the free-parameter count of a tied-covariance Gaussian HMM and n is the number of frames).
# Auto-select the state count via BIC over 2..5 states
frethmm run --input-dir ./traces/ --states auto --min-states 2 --max-states 5 --workers 4When auto-selection runs, *_summary.json records the chosen BIC, AIC, and a model_candidates table listing every candidate's n_states / log_prob / bic / aic so you can audit the decision.
{
"n_states": 3,
"bic": -12259.40,
"aic": -12330.18,
"n_init": 10,
"model_candidates": [
{"n_states": 2, "log_prob": 5000.1, "bic": -9800.2, "aic": -9870.5},
{"n_states": 3, "log_prob": 6193.8, "bic": -12259.4, "aic": -12330.2},
{"n_states": 4, "log_prob": 6195.0, "bic": -12240.8, "aic": -12320.1}
]
}Note on
--guesses: initial guesses are ignored with--states autobecause the per-candidate state count varies; multi-start provides the initialization diversity instead.
Background: In single-molecule fluorescence experiments, a persistent lowest-signal period can indicate photobleaching or fluorophore inactivation. Retaining data after that period can distort classification of the biologically relevant states.
Two-pass fitting workflow:
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ 1st Pass │ ──→ │ Locate │ ──→ │ Trim data │
│ HMM on full │ │ lowest │ │ at cutoff │
│ trace │ │ state run │ │ point │
└──────────────┘ └──────────────┘ └──────────────┘
│
▼
┌──────────────┐
│ 2nd Pass │
│ HMM on │
│ trimmed data│
└──────────────┘
- First pass classification: Fit HMM on the full trace to obtain the Viterbi state path.
- Identify lowest state: Find the state with the lowest mean value.
- Forward scan: Starting at
0 s, locate the first uninterrupted lowest-state run that reaches--low-state-tail-trim-seconds. - Trim data: Retain samples through
run start + configured seconds(inclusive) and discard every later sample. - Second pass classification: Re-run HMM on the retained trace for the final classification.
Note: If the lowest state never persists beyond the threshold duration, no trimming is performed and the first-pass result is retained.
CLI examples:
# Single file: retain through the first lowest-state window that reaches 5 seconds
frethmm run --files trace.csv --states 2 --low-state-tail-trim-seconds 5.0
# Batch: 3-second threshold, 4 parallel workers
frethmm run --input-dir ./traces/ --states 3 --low-state-tail-trim-seconds 3.0 --workers 4
# Combined with Review Grid: trim then review
frethmm review-grid --input-dir ./traces/ --output review.png --states 2 \
--low-state-tail-trim-seconds 5.0 --rows 4 --cols 8Output metadata: When trimming is active, *_summary.json records these additional fields:
{
"low_state_tail_trim_seconds": 5.0,
"low_state_tail_cutoff_time": 47.3,
"low_state_tail_kept_frames": 473
}low_state_tail_trim_seconds: The configured trim thresholdlow_state_tail_cutoff_time: The actual cutoff time point (nullif trimming was not triggered)low_state_tail_kept_frames: Number of frames retained after trimming
GUI usage: The GUI's Lowest-state window (sec) field defaults to 250. Leave it blank to disable filtering. The setting applies to all direct runs and folder batches, including Review Grid generation.
The program auto-detects file format (header presence, delimiter type, column count). Two data modes are supported:
Single-channel mode (CSV, with header):
Time,channel1
0,2820
1,2884
2,2570For multi-column signals, use --signal-column to select a specific column:
Time,channel1,channel2
0,2884,-5096
1,2884,1289--signal-column 1 uses the channel1 column, --signal-column 2 uses channel2.
Dual-channel Donor/Acceptor mode (whitespace/tab delimited, 3 columns, no header):
<time> <donor> <acceptor>
In this mode, FRET efficiency A/(D+A) is automatically computed as the HMM input signal.
Each input file generates the following outputs:
| File | Format | Description |
|---|---|---|
*_classified.csv |
CSV | Primary output: time, classified_mean — the idealized trace |
*_summary.json |
JSON | State means, frame fractions, transition matrix, dwell statistics, trim metadata, warnings |
*report.dat |
Text | Model parameters (state count, means, sigma, transition probability matrix) |
*path.dat |
TSV | Raw signal channels + FRET signal + classified signal per frame |
*dwell.dat |
TSV | Dwell time table: <start_mean> <stop_mean> <frames_lasted> per dwell segment |
FretHMM/
├── frethmm/
│ ├── __init__.py # Version info
│ ├── app/
│ │ ├── cli.py # CLI entry point (run / tdp / review-grid / gui)
│ │ ├── gui.py # CustomTkinter GUI
│ │ └── i18n.py # Internationalization (English / Chinese, 138 keys)
│ ├── assets/
│ │ ├── frethmm.ico # Application icon
│ │ └── frethmm_logo.png # Application logo
│ ├── core/
│ │ ├── io.py # File I/O (trace reading + report output)
│ │ ├── model.py # HMM engine (Baum-Welch + Viterbi + lowest-state window filtering)
│ │ ├── batch.py # Multi-process batch processor
│ │ └── postprocess.py # Classified signal + dwell segments + transition stats
│ ├── domain/
│ │ └── models.py # Data models (Config / Trace / Result / ExportOptions)
│ ├── formats/
│ │ └── report_parser.py # report.dat parser
│ ├── legacy/
│ │ └── report_parser.py # Legacy report format parser
│ └── viz/
│ ├── review_grid.py # Review Grid batch visualization (paginated layout)
│ └── tdp.py # Transition Density Plot + Gaussian rate fitting
├── tests/
│ ├── fixtures/ # Regression test reference data
│ ├── test_io.py # I/O and report parsing tests
│ ├── test_review_grid.py # Review Grid visualization tests
│ └── test_golden.py # CLI regression tests
├── docs/
│ ├── images/ # Screenshots
│ └── FretHMM-refactor-plan.md # Development roadmap
├── pyproject.toml # Project configuration
├── build_exe.py # PyInstaller build script
├── frethmm.spec # PyInstaller spec file
├── LICENSE # MIT License
└── README.md
pip install -e ".[dev]"
pytest tests/ -vEach successful CLI analysis writes a timestamped
frethmm_run_manifest_*.json next to its primary outputs. The manifest records
the command, fit parameters, input/output file metadata, FretHMM version,
Python version, and runtime dependency versions without copying experimental
data. This also applies to --classified-only, which suppresses auxiliary
classification exports but retains the run manifest.
The committed fixtures under tests/data/ are small, synthetic or
de-identified CSV and legacy report samples. They are sufficient for the core
regression suite and do not include raw images or ND2 files. FretHMM currently
starts from exported trajectory files (.csv, .dat, .txt, .tsv); ND2
image-to-trajectory processing remains an upstream workflow.
# Install the explicit build dependency first
pip install -e ".[gui]"
# Directory mode (default, produces dist/FretHMM/ directory)
python build_exe.py
# Single-file mode (produces dist/FretHMM.exe, easy to distribute)
python build_exe.py --onefileThe build produces a standalone Windows GUI executable — no Python installation
required. Directory-mode builds also create dist/FretHMM.zip, a SHA-256
sidecar, and a JSON release manifest. The single-file build produces
dist/FretHMM.exe; validate it without launching the UI with
dist/FretHMM.exe --version.
Data preprocessing, post-fit state consolidation, diagnostics, and GUI binding:
-
Preprocessing module (
frethmm.core.preprocess):- Robust scale estimation using Median Absolute Deviation (
estimate_robust_sigma). - Outlier spike detection and cleaning (
--remove-spikes,--spike-threshold-sigma). - Initial acquisition/shutter artifact trimming (
--trim-initial-artifacts,--max-initial-artifact-frames). - Edge-preserving median smoothing window (
--smooth-window).
- Robust scale estimation using Median Absolute Deviation (
-
Postprocessing module (
frethmm.core.postprocess):- Minimum dwell frame segment merging (
--min-dwell-frames) to suppress transient noise flickers. - Near-identical state merging (
--merge-state-threshold,--merge-state-sigma-factor). - End-to-end post-fit classification cleanup (
cleanup_classification_result).
- Minimum dwell frame segment merging (
-
Fit quality diagnostics (
frethmm.core.metrics):- Quantitative metrics: minimum and mean SNR, RMSE, MAE,
$R^2$ , and state occupancies recorded in*_summary.json.
- Quantitative metrics: minimum and mean SNR, RMSE, MAE,
-
Review grid enhancement:
- Added
--filesparameter toreview-gridto allow visual grid generation for specific files without folder restructuring.
- Added
-
GUI controls & bilingual i18n:
- Main parameters panel quick controls (Trim Start Artifacts, Clean Spikes, Min Dwell frames).
- Dedicated Clean & Filter Options dialog accessible from main panel, settings menu, and parameters dialog.
- Full English/Chinese bilingual support with dynamic switching.
Plot-ready per-molecule table from events:
- New
input_plot.csvwritten alongside the other three event tables (CLI and GUI): one row per source file withsource_file,ON_events,OFF_events,Fluorescence_strength(the file-levelstate_value_range), andDuration_time(the last event'send_time— for ON-then-bleach traces, the observable duration before photobleaching). Designed as a direct input for downstream visualisation. - New
summarize_plot_input+PLOT_INPUT_FIELDSinfrethmm.core.events, shared by CLI and GUI; the run manifest now listsinput_plot.csvamong the outputs.
Per-event fluorescence amplitude in event_details.csv:
- New
state_value_rangecolumn (file-level value repeated on every event row of a source file):- Files with multiple included events:
max(state_value) − min(state_value)over included events — the real ON/OFF fluorescence contrast (excluded/omitted photobleach tails do not participate). - Files whose only included event is a single ON event (e.g. molecule stays ON then photobleaches):
ON level − minimum classified value in the file— the signal amplitude above the bleached baseline. - Degenerate cases (constant-ON trace, single OFF, no included events):
0.0.
- Files with multiple included events:
- Backward compatible:
read_event_detailsparses pre-v1.7 event files without the column (defaults to0.0), so dwell-stats keeps consuming older outputs. - CLI and GUI events outputs pick up the column automatically (both share
DETAIL_FIELDS).
Windows GUI process-pool parallelism with safe defaults and cancel semantics:
- GUI workers: GUI Workers default to
2and accept1–4; CLI--workersremains default1. Values3and4prompt about higher CPU/memory use. Parallelism is between files only. - True Windows multi-process pool: classification and review-grid jobs run through a
spawn-compatible process pool so Workers=2/4 actually accelerate batch work on Windows. - Cancel / close safety: after cancel, no further files are scheduled; in-flight files may finish and keep their
*_classified.csvoutputs. A cancelled review-grid run does not publish a partial PNG. Close after a successful submit keeps completed status. - Release artifact: Windows single-file
FretHMM.exeis published with a SHA-256 checksum and JSON version manifest.
Batch review workflow, multi-state event statistics, and lowest-state window filtering:
- GUI workflow: supports multiple raw input folders with automatic per-folder
<folder>_outputdirectories, per-folder review grids, collision choices (overwrite / cancel / versioned output), and a dedicated reviewed-classification folder input for ON/OFF analysis. ON/OFF results are written to<folder>_output_ONOFFby default. - Multi-state ON/OFF statistics: 2-state traces report high-state ON and recovered low-state OFF events. For 3 or more states, every non-lowest activity stage is analysed independently; a descent is an OFF event only if that same stage recovers. Terminal non-recovering low segments are omitted.
- Lowest-state window filtering: replaces terminal-only trimming with a forward scan from
0 sfor the first continuous lowest-state window meeting the configured duration (default250 s), then refits only the retained data. - Release artifact: Windows single-file
FretHMM.exeis published with a SHA-256 checksum and JSON version manifest.
Dwell-time statistics and rate-constant fitting:
- New
dwell-statssubcommand: consumesevent_details.csv(output ofevents) and writesdwell_stats_summary.csv+dwell_stats_per_file.csvwith extended descriptive statistics (median, std, min/max, p25/p75) for ON and OFF dwell times. - Exponential rate-constant fit: single-exponential
A·exp(-k·t)fit of the pooled dwell-time histogram (histogram +scipy.optimize.curve_fit, boundedk ≥ 0), reporting the rate constant and implied mean dwell time for ON and OFF. Skippable via--no-fit. - New modules:
frethmm/core/dwell_stats.py(describe_durations,fit_exponential_dwell, extended summaries) andfrethmm/formats/event_details_parser.py(reverse-parseevent_details.csv). - No regression:
eventscommand andevents.pyare untouched;dwell-statsis a pure downstream consumer. - Tests: new
test_dwell_stats.py(12 tests) andtest_dwell_stats_cli.py(3 end-to-end tests), all self-contained. - Release reproducibility: CLI analysis commands now emit a timestamped run manifest; committed de-identified fixtures replace workspace-dependent skipped regression tests; Windows CI builds and version-smoke-tests the GUI bundle.
- Tail trimming correction: low-state trimming applies only to a persistent terminal low state, avoiding accidental removal of valid low-state segments in the middle of a trace.
ON/OFF event analysis brought in-package:
- New
eventssubcommand: extract ON/OFF events from*_classified.csvfiles and writeevent_details.csv,event_summary.csv, andevent_stats_overall.csv. - N-state generalisation: the highest-mean state is ON, all others are OFF (2-state traces behave exactly like the legacy "high = ON" rule).
- Terminal-OFF exclusion: a long final OFF run (default ≥ 100 s) is flagged
excludedand omitted from dwell statistics while still listed. - New modules:
frethmm/core/events.py(event detection + summaries),frethmm/formats/classified_parser.py(reverse-parse*_classified.csv), andfind_classified_filesinfrethmm/core/io.py. - Tests: new
test_events.py,test_classified_parser.py, andtest_events_cli.pywith self-contained synthetic traces (no external sample dependency).
Algorithm hardening — multi-start fitting and BIC-based state-count selection:
- Multi-start fitting (
--n-init, default 10): deterministic multi-start Baum-Welch that keeps the best log-likelihood. Start 0 reproduces the legacy single-fit, so--n-init 1is byte-compatible with v1.1. - BIC model selection (
--states autowith--min-states/--max-states): scan a state-count range, fit each candidate with multi-start, and pick the lowest BIC. - New metrics module (
frethmm.core.metrics):compute_aic,compute_bic,count_gaussian_hmm_params. - Summary JSON: records
n_init,best_start_index,bic,aic, andmodel_candidateswhen multi-start or auto-selection is active (omitted for single-start legacy runs). - GUI: new "Auto-select states (BIC)" checkbox with min/max range, plus an
n_initfield in the parameters panel and dialog; folder-batch jobs support auto-selection. - Tests: new
test_multistart.pyandtest_model_selection.pywith synthetic-trace fixtures; golden tests pinned to--n-init 1to preserve byte-exact regression.
Documentation and release infrastructure update:
- Enhanced README with detailed visualization docs (Review Grid, TDP) and data filtering workflow (low-state tail trimming)
- Added LICENSE file (MIT)
- Version bump to 1.1.0 in
__init__.pyandpyproject.toml - Updated
.gitignorewith additional exclusion rules
Batch visual review release:
- Batch review grid CLI: New
review-gridsubcommand for batch classification with paginated grid overview - GUI review grid: Dedicated Review Grid section in GUI for generating paginated review images
- Paginated layout: Configurable
rows × colsgrid, suitable for quick visual screening of 2-state, 3-state batches - Visual review enhancements: Each panel overlays raw signal with classified trace; shows filename,
log_prob, andstate means
GUI layout optimization and packaging improvements:
- Layout restructure: Removed ScrollableFrame; parameters and output panels side-by-side
- Collapsible runtime panel: Right sidebar defaults hidden; toggle via Show/Hide Runtime button
- Application icon: Added
frethmm.icoandfrethmm_logo.png - Window sizing: Default 1280×720, minimum 1180×660
- PyInstaller slimming: Reduced EXE size by excluding unused packages
--onefilemode: Added single-file EXE build support
GUI stability and export options:
- ExportOptions: Fine-grained control over output file types
- GUI output checkboxes: Select which files to generate
- Worker error handling: Full traceback logging
- Global exception hook: Error dialogs even in
console=FalseEXE mode
CustomTkinter migration and CLI enhancements:
- CustomTkinter migration: Full dark/light/system theme support
--classified-only: CLI flag to output only*_classified.csv- Folder batch panel: Per-folder state count, data mode, and signal column
- Runtime panel: Real-time status, progress, and result details
Project restructured as FretHMM with modular architecture:
- Modular packages:
core/domain/app/formats/legacy/viz - CLI with single-file and directory batch modes, multiprocessing support
- Full GUI with menu bar, parameters dialog, bilingual support, threaded analysis
- Default outputs:
*_classified.csv+*_summary.json - Additional outputs:
report.dat,path.dat,dwell.dat - TDP visualization + Gaussian rate fitting
- PyInstaller Windows executable build
- Regression test coverage (I/O, report parsing, end-to-end CLI)
GUI major update:
- Menu bar (File / Settings / Help) and parameters dialog
- Bilingual support (i18n): English/Chinese real-time switching
- Modern UI: platform-adaptive fonts, custom ttk theme, color-coded logs
- Startup optimization: lazy loading of heavy libraries
- Warning handling: captured and displayed with orange highlighting
Initial release:
- Complete HMM analysis pipeline (Baum-Welch training + Viterbi decoding)
- CLI tool (
run/tdp/guisubcommands) - tkinter GUI (file selection, parameters panel, progress bar, results table, log panel)
- Multi-process batch processing
- TDP visualization
- PyInstaller GUI packaging script
