Skip to content

feat: normalized [0,1] score by default + batch-first API (v0.1.1) - #2

Merged
kohbanye merged 5 commits into
mainfrom
docs/readme-update
Jul 23, 2026
Merged

kohbanye merged 5 commits into
mainfrom
docs/readme-update

Conversation

@kohbanye

@kohbanye kohbanye commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Makes the normalized [0, 1] score the default and reworks score() to be batch-first. Also trims the README. Bumps to v0.1.1.

Scoring API

  • CLaMNPScorer.score(smiles, raw=False) now accepts a single SMILES or a list:
    • a list is scored in one batched forward pass (much faster than one-at-a-time) and returns a list;
    • a string returns a single score.
    • raw=True returns the raw log-likelihood ratio.
  • This replaces score() / score_normalized() / batch_score() / batch_score_normalized() with one method.
  • The default is the sigmoid-normalized score in [0, 1] (0.5 = neutral, higher = more natural-product-like), calibrated with sigmoid_k=2 to match the paper (DEFAULT_SIGMOID_K).
scorer = CLaMNPScorer.from_pretrained()
scorer.score(["CC(=O)Oc1ccccc1C(=O)O", "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"])
# -> [0.44, 0.59]   (aspirin synthetic-like, caffeine natural-like)
scorer.score("CCO")             # single -> one score
scorer.score(["CCO"], raw=True) # raw log-likelihood ratio

CLI

  • Default output is now the normalized [0, 1] score; --normalized is replaced by --raw-score for the raw ratio.
  • Inputs are always scored in a single batch.

README

  • Drop the "How it works" section; fold essentials into the intro (models are GPT-2, score is [0, 1]).
  • Show real values (aspirin 0.44, caffeine 0.59) and the batch list-in/list-out form as primary.

Verification

  • ruff check + ruff format --check: clean.
  • ty check clamnp: clean (the overloaded score() adds no diagnostics).
  • Functional test (no models): score("x") → float, score([...]) → list, raw=True returns raw, batch_score removed. Normalized values match sigmoid(2·raw) (0.4395 / 0.5859).

Numbers are computed from the published models' verified raw scores.

🤖 Generated with Claude Code

https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs

kohbanye and others added 3 commits July 23, 2026 17:31
- Drop the "How it works" section; fold the essentials into the intro
  (the models are GPT-2, the score is a log-likelihood ratio where positive
  is more natural-product-like).
- Replace the qualitative "low/high" example comments with real default
  scores: aspirin -0.12 (synthetic-like), caffeine +0.17 (natural-like).
- Note that the CLI/`score()` default is the raw ratio and `--normalized`
  gives the sigmoid-squashed [0, 1] value.
- Drop the stale "(once published)" note now that clamnp is on PyPI.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
The public score is now the sigmoid-normalized value in [0, 1] (0.5 = neutral,
higher = more natural-product-like), calibrated with sigmoid_k=2 to match the
paper; the raw log-likelihood ratio is now opt-in.

- CLaMNPScorer.score(smiles, raw=False) accepts a single SMILES OR a list.
  A list is scored in one batched forward pass (much faster than one at a time)
  and returns a list; a string returns a single score. raw=True returns the raw
  log-likelihood ratio. This replaces the separate score()/score_normalized()/
  batch_score()/batch_score_normalized() methods.
- sigmoid_k now defaults to DEFAULT_SIGMOID_K (2.0).
- CLI: --normalized is replaced by --raw-score; the default output is the
  normalized score and inputs are always scored in a single batch.

Bump version to 0.1.1.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
Show the [0, 1] score (aspirin 0.44, caffeine 0.59) and the list-in/list-out
batch form as the primary usage; note --raw-score / score(raw=True).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
@kohbanye kohbanye changed the title docs: trim README and use real example scores feat: normalized [0,1] score by default + batch-first API (v0.1.1) Jul 23, 2026
kohbanye and others added 2 commits July 23, 2026 17:57
- Intro: just name the models "(GPT-2)"; drop the [0, 1]/raw explanation.
- Drop the "score a list in one batched pass" comment and the "By default
  clamnp prints ... --raw-score" note (both already clear from the examples).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
Restore "A molecule that a natural-product model finds likely but a synthetic
model finds unlikely gets a high score."

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
@kohbanye
kohbanye merged commit 3c01d06 into main Jul 23, 2026
2 checks passed
@kohbanye
kohbanye deleted the docs/readme-update branch July 23, 2026 09:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant