Catalyst.

Under the hood

Sources and Methodology

Catalyst is a research tool that combines market data, your own portfolio information and AI analysis. This page explains exactly what data we use, how it is processed, and how the scores and calls are produced.

Pre-registered · written while this record was empty

What we would have to earn

A model that can only be judged by the people selling it is not being judged. So what it would have to demonstrate was written down before the record contained anything: We are not claiming this model is predictive. It ships as a clearly labelled second opinion, free, with its record published in full. If the 95% confidence band on win rate against the sector benchmark sits entirely above 50% over at least 200 sector-gradeable calls, we will have earned the right to describe it as predictive. Until then we won't.

Deciding this now, with no data, is the same discipline as fixing the grading windows before the outcome is known. Deciding it later, with subscribers attached, is when people start relaxing the test.

Supersedes a stopping rule published earlier, withdrawn on 2026-08-21. Its original wording and the reason it changed are kept in full on the record.

Every call is graded against both the S&P 500 and, where the company sits in a sector we track, a liquid sector ETF. The pre-registered stopping rule is tested against the sector benchmark, which is the harder of the two.

The honest prior is that scheduled events are the most efficiently priced moments in a market: the date is public weeks ahead and the consensus view is already in the price. That is why the surrounding product — your own event concentration, the calendar for what you actually hold, and the empirical distribution of past moves — is built to be useful whether or not the model clears this bar.

See the record

Editorial, not modelled

How macro releases are tiered

Macro releases are universal: they move every portfolio at once, so interleaving two dozen of them with the events attached to individual holdings buries the ones you can act on. The next five Tier 1 releases are shown as a compact lookahead above the calendar; everything else stays in the full “all events” view. Nothing is deleted.

The split is a published editorial list of release types, not a market-moving score. A score would be another unvalidated composite, and it would be arguable. This is a line we can state plainly:

  • Tier 1 — central bank rate decisions (Bank of England, Federal Reserve, ECB); inflation (CPI, PCE and HICP — PCE is the US inflation release and is Tier 1 despite its formal title, “Personal Income and Outlays”); labour market and employment; GDP in any estimate, first, second, third, monthly or quarterly.
  • Tier 2 — retail sales, trade balance, industrial production and the index of production, producer prices, house prices, public finances, business investment, balance of payments, and any other statistical bulletin.

A release type matching neither list defaults to Tier 2 and is logged, because defaulting the other way would let the lookahead fill with unclassified noise. Which of the upcoming Tier 1 releases make the five slots is weighted by the listing venues in the book; the survivors are then shown in date order, because a calendar out of date order is confusing. The minority venue is never excluded outright.

This is the same registry rendered on the legal page, so adding or changing a provider updates both disclosures together.

22 active feeds in the shared registry.

Market feeds are batched and cached rather than streamed tick by tick. Earnings dates and prices are cross-checked against the independent providers named above; event cards show whether the pair confirms, differs, or remains single-source. Nasdaq date confirmation covers US-style symbols, so LSE lines may remain single-source until an exchange or issuer record is available.

Two independent layers are combined into a single rating card on the calendar.

Quantitative layer

Historical earnings

For ticker-linked events, Catalyst fetches the company's last several earnings surprises from Finnhub and computes:

  • Beat rate — the share of recent quarters where reported EPS beat the consensus estimate.
  • Average surprise % — the mean percentage by which the company has beaten or missed.
  • Sample size — how many quarters are included in the calculation.

These are pure historical facts. They tell you how reliably the company has delivered versus expectations, not what will happen this time.

AI layer

Gemini

The AI is asked to produce:

  • Direction — positive, negative or mixed expected reaction.
  • Confidence — a 1 to 5 score, where 1 is highly speculative and 5 is strongly supported by evidence.
  • Magnitude — a realistic expected percentage move, capped at 50%.
  • Swing factors — at most four short factors that could move the stock either way.
  • Rationale — a concise explanation of why the call is being made.

The model is explicitly told not to claim certainty and to flag thin evidence. Its output is clamped to valid ranges before it is stored.

When the app declines to rate

Abstention

A low confidence score is a poor way of saying "we don't know" — it still reads as a call. So an event is returned as insufficient evidence, with no direction and no confidence number, whenever any of the following is true:

  • The lead model declines to call the event on the evidence given.
  • The two independent models reach directly opposite conclusions — one positive, one negative — from the same evidence. A second model that only hedges caps the confidence instead of voiding the call.
  • The two data providers materially disagree on that company's price or reporting date.
  • Fewer than three usable inputs are behind the score.
  • There is no reporting history and no recent coverage to reason from.

The reasons are shown on the card itself. An abstention makes no claim, so it is never written to the graded record and cannot flatter the hit rate.

Citations

Auditable

Every score and position call stores the exact inputs it was generated from, listed underneath it as a numbered source list: the calendar event and its date, the earnings periods behind the beat rate, the price snapshot taken at the moment of scoring, any insider filings considered, and each news headline with its publisher and publication date.

Headlines link straight to the original article. Where an input comes from an API rather than a page — the earnings history, the quote, the calendar entry — the citation links to a public page for that company so the figure can be checked independently, and the underlying provider is named. Sources are frozen at the time of scoring: re-rating an event produces a fresh score with a fresh source list.

You can attach your own links — a filing, a company release, a report the feeds missed — when requesting a re-rating. Up to five pages are fetched, their text is passed to the model as evidence for that rating only, and each link joins the source list with its own confidence tier. Pages behind a login or paywall cannot be read and are marked as such rather than silently dropped.

Audit drawer

Replayable

Every rated event has an "Audit" drawer that replays the score from the values stored at scoring time: the T+1, T+5 and T+30 windows with their closing dates and the exact move, benchmark and excess arithmetic; then each metric step — event anchor, base rate, price anchor, qualitative read, expected move band, grading rule — with the numbered citations that fed it and the weakest source tier among them.

Nothing in the drawer is recomputed from live data, so it cannot drift away from the score it explains. Where a step has no source, it says so.

Source confidence

Per citation

Each citation carries a tier showing how directly that input supports the calculation. Tiers are assigned from the input type and, for news, from the publisher.

  • Primary — the record itself: an SEC filing, an issuer release carried verbatim on a regulatory wire, or the exchange/vendor data point the maths ran on (price snapshot, reported earnings history).
  • Secondary — an established outlet reporting on the record, or a scheduled calendar entry sourced from a data vendor rather than the company. Used for context and timing, one step removed from the numbers.
  • Tertiary — commentary, aggregators, and unattributed feeds. Background only; a tertiary source never moves a score on its own.

A score leaning mostly on secondary and tertiary citations should be read as weaker evidence than one built on filings and reported figures, even where the headline number looks the same. You can filter any source list to a single tier, or hide low-confidence sources entirely, to see what a score rests on.

Evidence strength meter

0–100

Every score shows one overall evidence meter above its source list. Each citation contributes a weight set by its tier — primary 1.0, secondary 0.6, tertiary 0.25 — multiplied by what the last link check found: a live unchanged page counts in full, a redirected or unchecked link 0.8, an edited page 0.55, an unreachable one 0.5, and a dead link only 0.15.

The weighted total is mapped through a saturating curve, so a wall of aggregator links can never read as strongly as two filings. Bands are Thin (under 35), Moderate (35–59), Solid (60–79) and Strong (80+). Because the discounts come from the link checker, the meter falls on its own when a cited page is edited or removed — nothing has to be re-rated for the weakening to show.

Link verification

Re-checked daily

Every citation link carries a "last verified" badge. Links are re-fetched automatically each morning, and again on the spot when a score is opened and its last check is more than a day old.

  • Verified — the page loads and its text is unchanged since we cited it.
  • Page changed — still live, but the article text has been edited since it fed the score. Read it before relying on the score.
  • Redirected — the publisher now points the link elsewhere; we follow it through to the new destination.
  • Link dead — the page has been removed. The citation is kept and struck through rather than deleted, so the record of what the score was based on stays intact.
  • Unreachable — we could not read it, usually a paywall or a block on automated checks. This is not evidence the page is gone.

Change detection compares the page's visible text with scripts, styling and clock times stripped out, so routine page furniture does not register as an edit.

Every transition between those statuses is written to an append-only audit log with the previous verdict, the new one, the HTTP response and the timestamp — as is any change to a source's confidence tier. Repeat checks that return the same verdict are not logged, so the log reads as a history of what actually changed. Open it under any source list ("Source change log"), or read the full record on the source change log. Entries are never edited or removed.

A generic read on the listed security and its upcoming events, separate from your cost basis, position size and risk preference.

Inputs

Per ticker
  • The next upcoming catalyst for that ticker, if one exists.
  • Recent company news and insider activity from the same data feeds.
  • A cached current price (not a live quote) for market context.

Output

Outlook

The model returns an outlook on the security itself — never an instruction about your holding:

  • Constructive — the evidence leans towards a positive reaction to the coming event.
  • Two-sided — evidence genuinely points both ways.
  • Cautious — the evidence leans towards a negative reaction.
  • No view — the evidence is too thin to lean either way. The model is allowed to abstain rather than manufacture a verdict.

This is deliberate. A view on your specific holding, weighted by the size of the position you actually hold, would edge towards a personal recommendation under Article 53 of the FCA’s Regulated Activities Order — and a disclaimer does not cure that. So the analysis describes the security; what it means for your portfolio is your decision.

Each outlook is paired with a confidence label (1–5), a risk level, the key uncertainty behind the rating and timing guidance. The confidence and risk labels have explicit definitions in the UI — hover or expand the “What do these mean?” key to read them.

The track record is the honesty check: every scored catalyst is tracked against what actually happened.

When a score is created

Snapshot

At the moment an impact rating is generated, Catalyst stores the predicted direction, predicted magnitude, the price before the event, and the catalyst metadata. This snapshot is what makes the track record verifiable.

Fixed forward windows

1d / 5d / 30d

Every call is measured at three fixed points after the event date, decided in advance and never moved: 1 trading session, 5 trading sessions and 30 calendar days. The 5-day window is counted in sessions, not calendar days, so a holiday week is graded on the same rule as any other week — the grader simply waits until at least 7 calendar days have passed before attempting it, which guarantees five sessions exist. The window is fixed before the outcome is known, so a call cannot be quietly re-graded at a more flattering moment.

The headline hit rate on the track record is the 5-day window. The 30-day window is published alongside it so you can see whether the call held up or simply caught a short-lived reaction.

How a call is graded

Daily job

A scheduled job runs at 22:00 UTC each day and grades any call whose window has just closed. It compares the last close before the announcement — not the price at the moment of the call — to the close of the window, so a call made weeks early does not book unrelated drift as event reaction. Drift between the call and the event is recorded separately as pre-event drift and classifies the result as:

  • Hit — the predicted direction matched the actual move.
  • Miss — the predicted direction was wrong.
  • No clear move — a directional call where the move was smaller than 1% in either direction. Excluded from the hit rate.
  • A "mixed" call is falsifiable both ways. It is a claim that the move stays inside the 1% band: a hit inside the band, a miss outside it. It is never excluded. Grading it any other way would make "mixed" a free win and reward a model for never committing.

Hit rate is calculated over hits and misses only; no-clear-move results are reported separately rather than being quietly counted as wins.

Measured against the market and its sector

S&P 500 + sector ETF

The S&P 500 is measured from the same anchor close and over the same window as the share, so each result shows both the raw move and the excess return over the market — a call that gained 3% while the whole market gained 4% was not a good call.

Every call is graded against both the S&P 500 and, where the company sits in a sector we track, a liquid sector ETF. The pre-registered stopping rule is tested against the sector benchmark, which is the harder of the two.

Where the company sits in a sector we track, the move is also compared with a liquid sector ETF (for example XBI for biotech, SOXX for semiconductors). Beating the market can just mean the sector rallied; beating the sector is what suggests the call was about the company.

Companies that report after the close first trade on the following session, so their anchor is the close on the event date. Companies reporting before the open are anchored to the previous close. This is recorded per result and shown in the audit trail.

The same job asks the AI to write a short debrief explaining what happened, why the call worked or failed, and what it means for the position now.

Nothing is removed

Public record

The track record is public and requires no account. Every non-abstained live call appears in it, including losses and calls still awaiting their window. Beta account and feature data may be reset or corrected as methods change. This never applies to published live calls or settled record outcomes: those are never re-graded, reset or deleted, and any correction is added as a dated note beside the original rather than replacing it.

Win or loss against the benchmark

Relative

Direction alone is a weak test, so every benchmarked window also records a win or a loss against the S&P 500 over the identical period. A call counts as a win only if the share beat the index; a 3% gain in a 4% market is recorded as a loss. The track record publishes that win rate next to the raw hit rate.

Sample size and confidence bands

Wilson 95%

A rate without a sample size is close to meaningless: 7 hits out of 10 and 700 out of 1,000 are very different claims. Every rate on the track record is therefore published with the number of calls behind it and a 95% confidence band:

  • Rates (hit rate, win rate versus the benchmark) use a Wilson score interval, which stays well behaved at small samples and at rates near 0% or 100%.
  • Averages (average move, average excess return) use a normal 95% interval around the mean. When that band crosses zero, the record does not yet demonstrate an edge, and the track record says so explicitly.
  • Any rate computed on fewer than 20 settled calls is labelled as indicative only.

These bands describe sampling uncertainty in this record — how much the measured rate could move with more data. They are not a prediction of the risk in any individual trade.

Is the confidence number real?

Calibration

A confidence score is only useful if higher confidence actually means a higher chance of being right. The curve below plots stated confidence (1/5 to 5/5) against the realised hit rate over the historical replay. The thin line is what each level implies; a bar shorter than the line means the model is overconfident there.

The same chart, alongside naive baselines and the Brier score, is on the track record.

Sliced by prediction

Bias check

A model can look accurate simply by being persistently bullish while the market rises. The track record breaks the hit rate down by what was predicted — positive, negative, mixed — so a directional bias shows up rather than hiding inside a single headline number.

Sliced by confidence, evidence and lead time

Self-audit

The 1–5 confidence score and the evidence band are quality claims, and a quality claim that is never tested is decoration. The track record therefore groups the same hit rate by stated confidence, by evidence band and by how many days before the event the call was published. If a higher confidence score or a stronger evidence band does not produce a higher hit rate, that is visible on the public page rather than buried.

Broker key handling

Encrypted at rest

A Trading 212 API key is encrypted before it is stored, using AES-GCM with a server-held secret that never reaches the browser. Only the ciphertext and a short hint (the last few characters, so you can tell one key from another) are kept. The plaintext key is decrypted in memory for the duration of a sync and is never logged, never returned to the browser, and never included in an AI prompt.

Row-level security means a key row is readable only by the account that created it. You can revoke access at any time from Integrations — removing the connection deletes the stored ciphertext — and you can additionally revoke the key inside Trading 212, which invalidates it everywhere at once. Keys are read-only by scope: they cannot place, amend or cancel an order.

Portfolio weight

Calculation

For each ticker, exposure is calculated as market value divided by total portfolio market value across all imported holdings. A 5% badge on a catalyst means 5% of your portfolio value is in that stock at the last cached price. Watchlist tickers do not contribute to exposure because no capital is allocated to them.

Reminders

7-day window

The reminders panel shows any catalyst in the next 7 days for a ticker you hold or watch. Dismissing a reminder stores a dismissal in your account so it does not reappear, but it does not affect the underlying event or score.

  • Prices are delayed and batched. They are refreshed once daily, not live. Intraday entries and exits are not captured by the quote cache.
  • Coverage is US-market heavy. Earnings history, surprise data and insider filings are most reliable for US-listed stocks. UK and EU tickers may have thinner data or none at all.
  • Earnings dates move. Companies reschedule and the feed updates the date in place, but there may be a lag between an announced change and the next sync.
  • AI can be wrong. The model is instructed to be conservative and cite evidence, but it has no access to non-public information and cannot predict market shocks. The track record exists precisely because the calls are sometimes wrong.
  • Currency is not normalised. Holdings and quotes are stored in each stock's listing currency, so a mixed-currency portfolio total is indicative rather than exact.

Catalyst is a research and information tool. It is not a financial adviser, broker, investment manager or regulated advice service. Nothing on this site is a recommendation, solicitation or offer to buy, sell or hold any security, financial product or investment. All AI-generated ratings, ideas, position calls and track record entries are based on public data and statistical patterns, not personalised advice. Always do your own research, consider your own risk tolerance and financial circumstances, and consult a qualified professional before making investment decisions. Past performance and historical accuracy shown on the track record are not a reliable guide to future results. Data displayed is obtained from third-party sources (Finnhub, SEC EDGAR, Trading212 exports) and may be delayed, incomplete or inaccurate. Prices shown are delayed snapshots, not live tick data, and are shown in each security's listing currency.

Read the full disclaimer

Research and information only. Nothing here is financial advice, and past accuracy is not a guide to future results. See the user guide for the daily workflow.