feat: normalized [0,1] score by default + batch-first API (v0.1.1) - #2
Merged
Merged
Conversation
- Drop the "How it works" section; fold the essentials into the intro (the models are GPT-2, the score is a log-likelihood ratio where positive is more natural-product-like). - Replace the qualitative "low/high" example comments with real default scores: aspirin -0.12 (synthetic-like), caffeine +0.17 (natural-like). - Note that the CLI/`score()` default is the raw ratio and `--normalized` gives the sigmoid-squashed [0, 1] value. - Drop the stale "(once published)" note now that clamnp is on PyPI. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
The public score is now the sigmoid-normalized value in [0, 1] (0.5 = neutral, higher = more natural-product-like), calibrated with sigmoid_k=2 to match the paper; the raw log-likelihood ratio is now opt-in. - CLaMNPScorer.score(smiles, raw=False) accepts a single SMILES OR a list. A list is scored in one batched forward pass (much faster than one at a time) and returns a list; a string returns a single score. raw=True returns the raw log-likelihood ratio. This replaces the separate score()/score_normalized()/ batch_score()/batch_score_normalized() methods. - sigmoid_k now defaults to DEFAULT_SIGMOID_K (2.0). - CLI: --normalized is replaced by --raw-score; the default output is the normalized score and inputs are always scored in a single batch. Bump version to 0.1.1. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
Show the [0, 1] score (aspirin 0.44, caffeine 0.59) and the list-in/list-out batch form as the primary usage; note --raw-score / score(raw=True). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
- Intro: just name the models "(GPT-2)"; drop the [0, 1]/raw explanation. - Drop the "score a list in one batched pass" comment and the "By default clamnp prints ... --raw-score" note (both already clear from the examples). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
Restore "A molecule that a natural-product model finds likely but a synthetic model finds unlikely gets a high score." Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes the normalized
[0, 1]score the default and reworksscore()to be batch-first. Also trims the README. Bumps to v0.1.1.Scoring API
CLaMNPScorer.score(smiles, raw=False)now accepts a single SMILES or a list:raw=Truereturns the raw log-likelihood ratio.score()/score_normalized()/batch_score()/batch_score_normalized()with one method.[0, 1](0.5 = neutral, higher = more natural-product-like), calibrated withsigmoid_k=2to match the paper (DEFAULT_SIGMOID_K).CLI
[0, 1]score;--normalizedis replaced by--raw-scorefor the raw ratio.README
[0, 1]).0.44, caffeine0.59) and the batch list-in/list-out form as primary.Verification
ruff check+ruff format --check: clean.ty check clamnp: clean (the overloadedscore()adds no diagnostics).score("x")→ float,score([...])→ list,raw=Truereturns raw,batch_scoreremoved. Normalized values matchsigmoid(2·raw)(0.4395 / 0.5859).Numbers are computed from the published models' verified raw scores.
🤖 Generated with Claude Code
https://claude.ai/code/session_01C4cq7P9sF3VuHXFkDozKRs