Skip to content

Page 02/05

Technology

Language models read. They do not trade.

The research pipeline uses agentic tooling heavily at the front, where the work is reading, triaging, and proposing. Behind a hard boundary in the middle, everything is deterministic Python with no model in the loop. That boundary is the most important architectural decision in the whole system.

Store
Point-in-time, as-of
Evaluation
Deterministic Python
Models
Front of pipeline only
Seeds
Fixed, no network
  • Point-in-time
  • As-of queries
  • Corporate actions
  • Override table
  • Retrieval
  • Graveyard
  • Fixed seeds
  • No network at eval
  • Forward logs
  • Reproducible from state
01The pipeline

Six stages, one boundary.

Stages one through three are where models are allowed to operate: reading unstructured text, proposing candidates, and summarising what a filing said. Everything after the boundary has to be reproducible from state, so nothing stochastic is permitted to touch it.

The rule is not about model quality. It is that a position you cannot re-derive is a position you cannot audit when it goes wrong.

  1. 01

    Ingest

    Filings, concall transcripts, exchange bhavcopy and corporate actions land in a point-in-time store. Every row is stamped with when it became knowable, not when it was fetched.

  2. 02

    Scout

    A retrieval layer over that corpus proposes candidate hypotheses and triages them against a graveyard of things already tried, so the same dead idea is not re-tested twice.

  3. 03

    Reader

    Survivors get read closely. A hypothesis only leaves this stage with a written economic mechanism attached. No mechanism, no test.

  4. Model boundary
    04

    Harness

    Deterministic Python. The signal is computed, costed, and evaluated on a development window, then once - and only once - on a holdout that is spent on contact.

  5. 05

    Warden

    A veto layer, not a ranking layer. Liquidity, concentration, corporate-action integrity, and regime constraints hold a hard veto over anything the model wants to do.

  6. 06

    Book

    Position sizing and execution. Fully reproducible from the signal state and the constraint set, with no model in the loop.

02Infrastructure

The parts that are boring on purpose.

None of this is novel. All of it is the part that quietly decides whether a backtest was ever real, which is why it gets the attention a model would get somewhere else.

01

A point-in-time store, not a database

Two timestamps per row: when the fact was true, and when it became knowable. Every query the research harness makes is an as-of query. This is more expensive to build than a conventional store and it is the single highest-return thing in the stack, because it makes an entire category of backtest fantasy structurally impossible rather than merely discouraged.

02

Corporate actions handled adversarially

Exchange close prices are not adjusted for splits and bonuses, and third-party adjustment data on Indian listings is wrong often enough that it cannot be trusted unchecked. We keep an override table on top of the vendor feed and alert on any single-day move that lacks a matching entry in the action calendar, because the alternative is a reversal signal that reliably buys stock splits.

03

Retrieval over filings and transcripts

Annual reports, quarterly filings, and concall transcripts are chunked and indexed so a hypothesis can be checked against what companies actually said, rather than against a summary of it. This is a reading aid for researchers. Nothing it produces reaches a position.

04

A graveyard of dead hypotheses

Every idea that has been tested and killed is recorded with its mechanism, its evaluation window, and why it failed. New candidates are checked against it before anything is run. Without this, a research programme quietly re-tests the same idea every few months and eventually gets a false positive.

05

Deterministic evaluation

The backtest harness is plain Python with fixed seeds and no network calls at evaluation time. Given the same state, it produces the same numbers, on any machine, in any order. Reproducibility is not a nice-to-have here: it is what makes a disagreement about a result resolvable.

06

Forward logs, not just backtests

Live signals are logged as they fire, before the outcome is known, so predicted and realised behaviour can be compared honestly later. A backtest tells you about the past you fitted; the forward log tells you whether the thing works.

Next

The failures this catches are documented.

A split read as a 79.5% drawdown, an overlay that reversed sign out of sample, a strategy that matched the benchmark it was supposed to beat. The research notes are where the infrastructure earns its keep.