Skip to content

Qanat is an MCP server: scopes, HTTP, and the console leaves - #8

Merged
fidetolabs merged 11 commits into
mainfrom
mcp-scopes
Sep 30, 2026
Merged

fidetolabs merged 11 commits into
mainfrom
mcp-scopes

Conversation

@fidetolabs

@fidetolabs fidetolabs commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

What this is

Point Qanat at any source that carries a timestamp. It processes that into factors, stores them, replays the chain over history one date at a time with no lookahead, prices what the portfolio held after fees and slippage, and serves all of it to your agent over MCP.

claude mcp add qanat -- qanat mcp --scope research

That was always what Qanat did. This is the release where the package says it and says nothing else.

Three scopes, and each is a promise

--read-only split the tools 23 and 10, which describes how the code is written and promises nothing anybody could build against. Replaced by data (6 tools), research (19) and full (31), nested. data is the door an institution connects at; research is the door a hosted service connects at.

scope has no default on the tool decorator, so a tool cannot be added without naming who may see it. set_bar sits in full while record_trial sits in research, so a caller can write down what it tried and cannot lower the line those attempts are held to.

It no longer has to be on your machine

qanat mcp --http speaks MCP's Streamable HTTP transport on /mcp, with sessions, and the scope fixed by the command that starts the server. qanat serve is that server with a clock: it opens the store, runs the scheduler, and mounts the same endpoint.

What you install is the engine and the server

The console, the terminal app and the HTTP API behind them are gone. Your agent client is the screen. About 14,000 lines, and that is the point rather than the cost: someone opening this repo used to see a chat window and a dashboard and decide what Qanat was before reading a word.

The unattended falsification pass got narrower with it. It used to reach the project through a dozen curl endpoints with --allowedTools Bash; it now goes over MCP at research, so the scope is what stops it editing the strategy it is attacking. That used to be a sentence in a prompt.

One guard got stronger

A negative fee pays you to trade, so turnover becomes profit: -9999 bps once put +1938% into the strategy book with nothing marking it. That was refused by a ge=0 on the console's request model, which held for one caller. It is in the engine now.

Upgrading

--read-only is gone, use --scope. qanat serve serves /mcp rather than /api/* and opens no console. qanat tui is gone. Everything else on the CLI is unchanged.

180 tests pass, ruff clean. Verified after the cut: stdio MCP at 31 tools, qanat serve answering MCP with the old /api returning 404, and the hosted client reaching the project through it.

fidetolabs and others added 11 commits September 25, 2026 00:29
A backtest could report +12.31% and nothing on the page could tell you whether
that meant anything. It now reports that the same run is +9.31% over doing
nothing, has a t-stat of 0.64 against a bar of 2.73, and was picked from eight
attempts -- which is to say, it is probably noise.

  * `totals` carries vol, Sharpe and max drawdown, annualised from the rebalance
    gap rather than a 252 that would call a weekly rule five times more volatile
    than it is. `backtest.benchmark` takes a symbol or `equal_weight` and the
    report prints excess, priced on the strategy's own periods.

  * Every replay is a row in a ledger with what it was an attempt at, what it
    varied, and what was decided about it. A trial is a distinct digest, not a
    run: re-running one configuration is the same question asked twice.

  * `backtest.bar` raises the threshold with the trial count. Off by default and
    reporting-only by default; with `gate: true` the one thing it refuses is
    calling a run `live`. Changes to it are logged with the actor, and loosening
    is logged as a warning -- whoever proposes strategies should not be able to
    quietly lower the line that judges them.

  * The console keeps its conversations. Sessions carry their transcript, cost
    and replays, write themselves a summary on close, and resume through
    `--session-id` / `--resume` so one id names the same thing on both sides.

  * `research:` runs an unattended pass that tries to *break* the least-tested
    strategy rather than improve it -- the one goal that cannot overfit. Budget
    is measured from what the CLI reports, not estimated.

Searching for new strategies is deliberately absent, and the window is not
sealed: the point-in-time views stop a step seeing the future, not the agent.
Both are prerequisites for a search loop, and both are named in the README
rather than implied.

Dropped `cursor-agent`. It was given only `--print`, whose default is plain
text, while the parser reads JSON lines -- so every answer was discarded and the
console showed a finished question with nothing in it. It also never got the
tool fence, and refuses to start until Workspace Trust is granted by hand.

One commit rather than five: the phases interleave inside the same files, and
splitting them afterwards would mean rewriting history that is easier to read
whole.
A run that holds four semiconductors out of a priced list of five hundred was
being scored against all five hundred, because `equal_weight` averages every
column of the price table. So the sector call and the ranking arrived as one
number, and the easier achievement -- riding a sector that doubled -- read as the
harder one. `benchmark: universe_equal_weight` holds the run's own universe
instead, and the two floors together say which half of the work earned anything.

  * Membership is read as of the day the portfolio was decided, not the day it
    was priced. The floor has to be the pool the step was choosing from, so it
    moves as the pool moved -- and a universe file with no `from`/`to` dates makes
    the floor carry the same survivorship bias the strategy does, which the run
    says in a note rather than leaving the comparison looking clean.

  * Which universe is the run's: the override if there was one, otherwise the
    alphas' own. Two alphas priced together that name different pools get the
    refusal `_own` already gives for disagreeing about `rebalance`, since a book
    spanning two pools has no single pool to be measured against. A run with no
    universe at all gets no benchmark lines and a note saying why; the market is
    not a stand-in for a pool nobody chose.

  * `equal_weight` keeps its behaviour. Redefining it to follow the universe
    would have changed the meaning of every number already recorded under that
    name. What was wrong was its description -- README and `models.py` both
    claimed the whole universe, and it has always been the whole price table.

  * The report says which floor it read, and `conditions` carries `benchmark`,
    so a run records what it was measured against rather than leaving it to be
    inferred from a yaml file that has since moved on.

`docs/attribution.md` is what the pair is for: the universe carries the claim
about the world, the step carries the arithmetic, and because a universe is
already a swappable argument both can be switched off independently. Four runs,
and the four numbers say whether the idea or the formula was worth having.
Where a hypothesis comes from was the one part of this project that lived in
somebody's head. `docs/narratives.md` is the design for the page that holds it --
deliberately outside the alpha pipeline: no narrative table feeds a weights
table, and no alpha reads a probability. News here is not a signal. It is what a
person reads before deciding what to test.

  * A narrative is a question the user wrote, ingesting on a schedule, and its
    three stances are fixed by that question -- the answer yes, unchanged, no.
    Nothing authors them and nothing replaces them, so a strategy saved in March
    cannot later show reasoning nobody had read when they saved it. An earlier
    draft had the agent superseding scenario text when the world moved; fixed
    stances make that unrepresentable rather than discouraged.

  * A `reading` is one date's text and, per stance, a probability with its
    reason. Appended, never updated, three every time, whole numbers summing to
    100 -- three floats do not sum to 1.0, and three probabilities summing to 1.2
    are not close to anything.

  * A strategy built from a reading is in-sample by construction, all of it. The
    reading was written at T from news up to T, and the replay prices history
    before T. The as-of views stop a step seeing the future; they cannot stop the
    agent, which read the reading. So a checkpoint's date is the frontier, the
    backtest is an examination rather than a test, and a live narrative
    manufactures its own out-of-sample by staying live.

  * The probabilities are the allocation. One alpha per stance, priced together
    as one book, and `_shares` normalises 55/30/15 raw -- nothing to build. But
    `combine` nets opposite sides and spreads the proceeds, so the blend cancels
    the hedge the negative stance was holding. Blending gives the expected
    portfolio; surviving whichever future arrives wants a second step that clears
    a floor, and two visible steps rather than one solver that explains nothing.

  * What cannot be scored is written down beside what can. A stance's later
    probability is not its answer -- that grades the agent against its own output,
    and self-consistency is not accuracy. A `resolution` rule against an
    ingestable number is, and no accuracy figure prints without the count of
    resolved readings beside it.

`topic` and `deck` are dropped; a narrative is the container and one word is
enough. `snapshot` is not used -- `plan.snapshot()` already holds it -- so the
object a strategy is built from is a `reading`.

Nothing here is built. The doc says so in its first line, because every other
file in `docs/` describes what ships.
A cron line assumes a machine that stays on, and most people running this have a
laptop that closes. So `schedule:` on a narrative is optional, which the
scheduler already handles correctly -- it collects `[j for j in project.jobs if
j.schedule]`, so an unscheduled narrative is left alone by design rather than by
accident, and `fire(job_id, actor=...)` already records whether a person or a
clock asked.

What the second mode costs is one row of a table, and it is invisible on a chart:

  * Scheduled, a flat line means the reading did not change. Asked for, it means
    nobody looked. Two readings three weeks apart draw the same shape as two a
    day apart, and both shapes read as `stable`. So a reading carries the window
    it read -- `covers_from`, `covers_to`, `n_articles` -- because 340 articles
    across 21 days being flat is arithmetic, and 12 across one day being flat is
    evidence. The chart draws what it observed and does not interpolate between
    two readings it did not watch.

  * Prices can be backfilled and news cannot. Skipping three weeks of price
    history loses nothing: a replay run later prices that window exactly as a
    nightly pass would have, so forward evidence on a laptop is discovered late
    rather than lost. A gap in news ingestion is a hole, because the APIs stop
    looking back and what a search still surfaces months later is what turned out
    to matter. So a manual trigger ingests the whole gap since the last reading,
    and says so when the gap runs past what the source can still answer for.

  * `as_of` is a moment rather than a date on a grid. Two readings in one
    afternoon are two readings -- somebody who asks again after fresh news has
    asked a second question, and only a re-run of one firing is the same question
    asked twice.

  * Readings that were asked for are a biased sample, since people check in when
    something is happening. Calibration figures carry their mode beside their
    count and are not pooled across the two.

Floating allocation is also less unruly here: readings arrive a handful of times
a quarter rather than nightly, so the count of distinct strategies stays small
enough to mean something. The mode that makes floating dangerous is the cron.
…g is

The previous version worried that a flat stretch could mean either `the reading
did not change` or `nobody looked`, and answered it with a stored window. The
worry was wrong: nothing draws a line through unwatched time, because a dot
exists where a reading exists and nowhere else. Three weeks of silence is three
weeks of empty chart.

  * Every dot opens, and shows that reading -- the text, the three probabilities,
    and the reason given for each. Which is what the append-only rule was for:
    nothing is overwritten, so every point stays answerable months later and a
    person can see what it looked like at the time rather than being told.

  * Dots that produced a checkpoint are marked apart from dots that did not, so
    the chart shows what was believed and where a belief became a position.

  * `covers_from` / `covers_to` survives, narrowed to the one thing dots cannot
    show: an ingest that could not reach back to the previous reading. That is a
    hole *inside* a dot, and it is the only gap here that hides. `n_articles`
    stays too -- three articles and three hundred draw dots of the same height.
Three things this file had wrong or missing. The probabilities were called the
allocation, which invents a second three-number vector reading like a second set
of beliefs -- there is one, and it scores candidate portfolios rather than being
one. A stance was described as a list of names, when `tech keeps spending on AI`
reaches semis, utilities, REITs, cooling and networking, and cuts the other way
for the hyperscalers paying. And nothing said which page any of it happens on.

  * A stance is several industries, so a book is built in two levels: how much
    each industry gets, from `exposure`; which names inside it, by margin of
    safety ranked **within that industry**. Ranking across industries compares
    measuring sticks, and the loosest method wins -- looser assumptions make
    larger discounts, so the shakiest fair value looks like the best bargain.
    Still one universe per stance, carrying the `sector` column `qanat init`
    already ships.

  * The probabilities appear in one line: `score(w) = Σ p_s · R_s(w)`. Splitting
    the money 55/30/15 across the books is the candidate you get by ignoring the
    floor, not the answer. Worked through with $1,000, four candidates and a
    -10% floor: the belief scores best and is infeasible, and what the user is
    shown is the belief, the weights, and what the floor cost.

  * The search is an exhaustive grid over a two-dimensional space -- ~5,000
    candidates, three dot products each, numpy, no new dependency, and
    deterministic, which the replay rules require. Exhaustive also makes the
    infeasible case honest: *no combination of these three books survives your
    limit* says the negative book is not hedging anything, where a solver would
    have returned the least-bad vector and said nothing.

  * `combine` cannot do this job. It nets opposite sides of a name and spreads
    the freed money over what is left, so a hedge in the negative book is
    cancelled and reinvested into the long it was hedging. It is the one function
    in the engine actively hostile to surviving a scenario, so the blend happens
    in a step and only finished books reach the weights stage.

  * Two alphas over the same features, which is the condition for comparing them:
    `blend` is the probabilities without the floor, `target` clears it. With
    `--universe u_base` that is three commands for three questions -- the
    strategy, what the floor cost, and whether the narrative was worth anything.

The line the design turns on: about fifteen rows of judgement -- which industries
a stance moves, and which way -- and five hundred numbers that were measured. An
agent that asserts the five hundred produces weights that look identical and mean
nothing.

`driver`, `exposure`, `book` and `floor` join the words. Nine now.
--read-only kept 23 tools and dropped 10, which describes how the code is
written and promises nothing anybody could build against. D-20260929-04
replaces it with --scope: data reads the tables, research adds replays and
their results, full adds authoring, ingest and scheduling. Each contains the
one before it.

scope has no default on the tool() decorator. A new tool now answers two
questions before it exists -- rule 3, does it need Qanat's data or engine, and
this one, which caller may see it -- and leaving the second to a default would
quietly put every new tool in front of an institution that connected at data.

set_bar sits in full and record_trial in research, so a hosted caller can
record what it tried and cannot lower the line those trials are held to.
backtest and record_trial are the only tools in research that write, and
neither touches qanat.yaml.

A tool above your scope names its scope instead of claiming not to exist. Each
scope is a contract, so if that message keeps coming back for the same tool,
the line is in the wrong place.
qanat mcp speaks JSON-RPC on a pipe. That is right for one person on one machine
and cannot be the hosted one: a request arriving at a server has nowhere to keep
a child process, and the client is somewhere else. --http serves the same tools
over Streamable HTTP on /mcp, with sessions, DELETE, and a 404 on an idle hour
that means call initialize again. GET answers 405, because this server sends
nothing on its own.

Three things are fixed by the command that starts it rather than by a request.
The scope, because a caller that names its own scope has no scope. One project
per process, because the store takes one writer. And loopback unless a token is
set, because binding somewhere reachable and answering anybody is not a default
worth having.

The README led with "an agent-first backtesting engine" and opened on qanat
serve, which describes what the console was the front door to. Line one is now
the MCP server, the quick start connects it to an agent, and the scopes and the
HTTP transport are on the page. GitHub description, topics and PyPI keywords
moved with it; dag came out of the topics. 518 lines to 235, no em dashes.

The console is still in the package. The README says so.

267 tests pass.
Someone opened the repo, saw a chat window and a dashboard, and decided what
Qanat was before reading a word. That question was being asked by the code, not
the README. D-20260929-01 rule 2 answers it by taking the screen out.

Gone: src/qanat/console (13 files), the twelve UI probes, the CSS test, and with
them GET /, the static mount, /api/ask and its stream, /api/trace, the five
/api/sessions routes and PUT /api/agent. api.py 2,010 lines to 1,592.

agent.py is headless.py. The filename made a claim this package should not make.
The module is unchanged and the unattended falsification pass still uses it.

open_console and console_status go too, so the surface is 31 tools and a Session
no longer holds a page. Scopes are 6 / 19 / 31.

qanat serve is the scheduler and says so, in its help, in what it prints, in the
README a new project is created with, and in the busy-store error, which used to
offer two doors that were the console and now points at qanat mcp --http. qanat
tui stays: a terminal program is not a screen this package has to host.

The release check verified the wheel carried console assets. It now verifies the
wheel carries the modules and carries no console and no agent.py.

Ten API tests covered the routes that went, so they went. The store's session
machinery survives because the pass uses it, so it got its own test instead of
losing coverage along with the routes.

255 tests pass. Verified after: stdio MCP at 31 tools, HTTP MCP at research, and
margiana answering through it.
There is no way to sit and operate qanat, so the parts that existed to be
operated are gone. tui.py and chart.py go with the web console: a terminal app
you watch is still a screen, and this package does not ship one.

api.py goes too, and that is the larger half. 1,592 lines of read model for a
console that no longer exists, kept alive by one caller: the unattended pass,
handed a dozen curl endpoints and --allowedTools Bash. That is a wide door.
Asked once to reshape some ideas, the agent walked out of the project into
qanat's own installed source and then into ~/.claude/projects.

The pass now reaches the project the way everything else does, over MCP at
--scope research, with the tools that scope offers and nothing besides. The
scope is also what stops it editing the strategy it is meant to be attacking.
That used to be a sentence in a prompt.

qanat serve is the MCP server with a clock: it opens the store, starts the
scheduler, and mounts the same /mcp endpoint, so one process holds the store and
everything reaches the project through it.

AppState moved to runtime.py, build_graph to graph.py next to the thing that
draws it, the Host and Origin guard to mcp_http.py.

One guard was nearly lost with the API. ge=0 on the request model was the only
thing refusing a negative fee, and a negative cost pays you to trade: -9999 bps
once put +1938% into the strategy book with nothing marking it. It is in the
engine now, so it holds for every caller.

180 tests pass. Verified: qanat serve answers MCP, the old /api is 404, and
margiana reaches the project through it.
The changelog led with what left, which reads as a smaller product. It leads
with what this is instead: point it at any source with a timestamp, it makes
factors, replays them over history with no lookahead, prices what was held after
costs, and serves all of it to an agent over MCP.

The removals are still in there, under what you install. 14,000 lines out is the
point rather than the cost: someone opening the repo used to see a chat window
and a dashboard and decide what Qanat was before reading a word.
@fidetolabs
fidetolabs merged commit 6395bc9 into main Sep 30, 2026
3 checks passed
@fidetolabs
fidetolabs deleted the mcp-scopes branch September 30, 2026 00:39
fidetolabs added a commit that referenced this pull request Sep 30, 2026
Qanat is an MCP server: scopes, HTTP, and the console leaves
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants