Page 02/05
Technology
Language models read. They do not trade.
The research pipeline uses agentic tooling heavily at the front, where the work is reading, triaging, and proposing. Behind a hard boundary in the middle, everything is deterministic Python with no model in the loop. That boundary is the most important architectural decision in the whole system.
- Store
- Point-in-time, as-of
- Evaluation
- Deterministic Python
- Models
- Front of pipeline only
- Seeds
- Fixed, no network
- Point-in-time
- As-of queries
- Corporate actions
- Override table
- Retrieval
- Graveyard
- Fixed seeds
- No network at eval
- Forward logs
- Reproducible from state
Six stages, one boundary.
Stages one through three are where models are allowed to operate: reading unstructured text, proposing candidates, and summarising what a filing said. Everything after the boundary has to be reproducible from state, so nothing stochastic is permitted to touch it.
The rule is not about model quality. It is that a position you cannot re-derive is a position you cannot audit when it goes wrong.
- 01
Ingest
Filings, concall transcripts, exchange bhavcopy and corporate actions land in a point-in-time store. Every row is stamped with when it became knowable, not when it was fetched.
- 02
Scout
A retrieval layer over that corpus proposes candidate hypotheses and triages them against a graveyard of things already tried, so the same dead idea is not re-tested twice.
- 03
Reader
Survivors get read closely. A hypothesis only leaves this stage with a written economic mechanism attached. No mechanism, no test.
- Model boundary04
Harness
Deterministic Python. The signal is computed, costed, and evaluated on a development window, then once - and only once - on a holdout that is spent on contact.
- 05
Warden
A veto layer, not a ranking layer. Liquidity, concentration, corporate-action integrity, and regime constraints hold a hard veto over anything the model wants to do.
- 06
Book
Position sizing and execution. Fully reproducible from the signal state and the constraint set, with no model in the loop.
The parts that are boring on purpose.
None of this is novel. All of it is the part that quietly decides whether a backtest was ever real, which is why it gets the attention a model would get somewhere else.
A point-in-time store, not a database
Two timestamps per row: when the fact was true, and when it became knowable. Every query the research harness makes is an as-of query. This is more expensive to build than a conventional store and it is the single highest-return thing in the stack, because it makes an entire category of backtest fantasy structurally impossible rather than merely discouraged.
Corporate actions handled adversarially
Exchange close prices are not adjusted for splits and bonuses, and third-party adjustment data on Indian listings is wrong often enough that it cannot be trusted unchecked. We keep an override table on top of the vendor feed and alert on any single-day move that lacks a matching entry in the action calendar, because the alternative is a reversal signal that reliably buys stock splits.
Retrieval over filings and transcripts
Annual reports, quarterly filings, and concall transcripts are chunked and indexed so a hypothesis can be checked against what companies actually said, rather than against a summary of it. This is a reading aid for researchers. Nothing it produces reaches a position.
A graveyard of dead hypotheses
Every idea that has been tested and killed is recorded with its mechanism, its evaluation window, and why it failed. New candidates are checked against it before anything is run. Without this, a research programme quietly re-tests the same idea every few months and eventually gets a false positive.
Deterministic evaluation
The backtest harness is plain Python with fixed seeds and no network calls at evaluation time. Given the same state, it produces the same numbers, on any machine, in any order. Reproducibility is not a nice-to-have here: it is what makes a disagreement about a result resolvable.
Forward logs, not just backtests
Live signals are logged as they fire, before the outcome is known, so predicted and realised behaviour can be compared honestly later. A backtest tells you about the past you fitted; the forward log tells you whether the thing works.
Next
The failures this catches are documented.
A split read as a 79.5% drawdown, an overlay that reversed sign out of sample, a strategy that matched the benchmark it was supposed to beat. The research notes are where the infrastructure earns its keep.