Every number, defined before you read it.
This page is the metric contract behind everything we publish: what each accuracy figure means, exactly how it is computed, and every time a definition has changed. Definitions here are frozen: a change appends to the changelog, it never silently redefines a number.
The live values themselves are on the proof page, recomputed continuously. Numbers quoted on this page are snapshots, each dated as of July 2026.
1 · What a liquidation magnet is
A magnet (also rendered as a level, band, or bright bucket) is a price zone where our model estimates a high density of pending forced liquidations, given current market positioning. It is built from real, aggregated cross-venue inputs: open interest and long/short skew, allocated across a distribution of leverage bands (3× through 100×) whose weights are continuously re-calibrated against the observed liquidation tape, with entry prices modeled as a distribution around recent volume-weighted price. Each leverage band maps deterministically to the price at which those positions would be forcibly closed; summing that mass per price bucket produces the liquidation-density ladder. The densest buckets are the magnets.
What a magnet is not. A magnet is descriptive: it says "given the positioning that exists right now, a large pool of forced-liquidation exposure sits near this price." It is not a statement that price will travel there, not a directional signal, and not trading advice. When positioning changes (open interest unwinds, positions close, leverage rotates), magnets legitimately move, fade, or disappear, and the interface says so rather than hiding it.
Model inputs and calibration corpus are described at the aggregate level only; the exact weighting recipe is proprietary. Results, not recipes, are what we publish and what /proof scores.2 · Metric definitions: the frozen contract
Seven definitions cover every accuracy number on the site. Each is computed walk-forward: the model is scored on data that arrived after the inputs it was given, and each carries its sample size and confidence bound wherever it is displayed.
2·1precision@N: the headline
At each sample time we reconstruct the model's inputs exactly as they stood one hour earlier: open interest, long/short ratio, entry-price distribution, strictly as-of that moment, with no look-ahead. The model ranks all price buckets by estimated liquidation density and the top ~10% are marked "bright" (a small number of symbols use a wider fraction, disclosed per symbol). precision@N is the dollar-weighted fraction of the liquidation events that actually printed in the following window which landed inside a bright bucket. Every published precision@N figure carries its sample size (n) and its Wilson 95% lower bound.
Picking 10% of buckets at random would score roughly 0.10. That baseline is shown for context but never headlined; see changelog entry #1 for why the multiple-of-baseline framing was retired.
The ruler mix, disclosed. Most symbols are graded on the top-10% ruler described above; ETH, SOL and XRP are graded on a wider top-20% ruler (a volatility-aware override of 24 July 2026, disclosed per symbol on /proof). The pooled out-of-sample precision@N therefore mixes two ruler difficulties by design: a random pick scores ≈0.10 on the narrow ruler but ≈0.20 on the wide one, so the honest baseline for the mixed pool is the dollar-weighted mix (≈0.146 at the current composition), not the flat 0.10. Every pooled baseline comparison on this site uses the mixed figure; the flat 0.10 is kept for reference only. Alongside the pooled number we publish a same-ruler shadow: the identical served configurations re-graded solely on the uniform top-10% ruler over the same held-out days (n = 13 independent day-clusters, with its own Wilson 95% lower bound), labeled descriptive and never headlined. A walk-forward placement decomposition (BTC + ETH, n = 13 held-out days) measured the ruler mix contributing ≈+19.75pp to the pooled point. The mix is stated here so the number survives scrutiny; the shadow shows the same evidence on one ruler.
As of July 2026: global sample-weighted precision@N ≈ 0.34 (n = 254,590 scored events across 25 symbols; per-symbol Wilson 95% lower bounds published alongside). Live per-symbol values on /proof.2·2The near-spot tautology: why capture metrics exclude the zone at spot
Roughly 88% of liquidation dollars land within about 1.5% of the current price. That means any model, including a trivial guess that paints the bands just off spot, scores well on a metric that counts near-spot events. A magnet sitting next to price gets "reached" almost by definition. Reading a near-spot-inclusive score as model skill would flatter us, so we publish companion metrics that exclude that zone entirely and grade only the harder, off-spot calls a trader actually cares about.
2·3Off-spot magnet capture (>1.5% from spot)
Of the liquidation dollars landing more than 1.5% from spot, the fraction caught by a top-ranked predicted band, scored walk-forward (out-of-sample), counted only on the (symbol, horizon) cells that cleared their own validation gate, averaged with equal weight across cells.
As of July 2026: ≈ 0.44 across 90 validated cells. Live on /proof.2·4Deep-magnet capture (>3% from spot): the strictest cut
The number we consider most defensible under diligence. Of the liquidation events landing more than 3% from spot (deep enough that a trivial near-boundary guess collapses), the fraction caught by a top-ranked predicted band. Computed walk-forward, pooled across every (symbol, horizon) cell with an adequate sample (at least 100 deep events), event-weighted, and deliberately not filtered to the model's winning cells: filtering to winners would cherry-pick the model up and collapse the comparison baseline, overstating the edge. Losing cells drag the number down, truthfully.
It is always published side-by-side with the trivial boundary baseline computed over the same unselected events, plus a dollar-at-risk-weighted variant, so you can see exactly what a no-skill guess achieves on the identical event set.
As of July 2026: model ≈ 0.51 vs trivial baseline ≈ 0.06, over ≈ 1.6M deep events (dollar-weighted ≈ 0.48). Live on /proof.2·5Wilson 95% lower bound: why we headline it
Wherever a per-symbol rate is published, the number we treat as load-bearing is not the raw rate but the lower edge of its 95% Wilson score interval, computed on the de-overlapped effective sample count (2·6 below). A rate on a small or dependent sample overstates certainty; the Wilson lower bound is the version of the number that remains defensible if the sample was lucky. A symbol grades "validated" only when its Wilson lower bound clears the random baseline on at least three independent windows. A strong-looking point estimate with a weak lower bound is labeled low-confidence instead, and is barred from carrying strong claims anywhere downstream.
2·6De-overlapped effective n: overlapping windows are not independent samples
Calibration runs are logged far more often than their evaluation windows elapse: a run lands every hour or few hours, but each one scores a multi-day trailing window, so consecutive measurements share most of their underlying events. Counting every logged run as an independent trial would inflate n and dishonestly shrink the confidence interval. For the published interval we count a measurement only when its own evaluation window has fully elapsed since the previously counted one, and the effective sample count scales the raw observation total by that non-overlapping fraction. Effective n is always ≤ raw n, and both are published per symbol so nothing is hidden in the aggregation.
2·7Walk-forward / out-of-sample protocol
Parameter selection and scoring never touch the same data. Any tunable parameter is chosen on the earlier slice of a chronological split (roughly the first 60%) and scored on the later, held-out slice: a genuine forward test, never a shuffle. Two additional guards: a split where one slice contains essentially no scored liquidation mass is refused the out-of-sample label rather than reported as one; and changes to the served model are gated on the held-out score, never the in-sample one. Every scored window reconstructs inputs strictly as-of their timestamps; no future information reaches any input.
2·8Per-magnet labels: n / reach% / horizon
Every magnet rendered in the product carries three labels: n, the sample count behind its statistics; reach%, the empirical share of past cases in which price reached a magnet at that distance within the stated window; and horizon, the window that reach% refers to. Reach rates are descriptive history, n-weighted across tracked symbols, and refreshed continuously. For scale, at a ~1% distance from spot, as of July 2026:
| Horizon | Reach rate | Sample n |
|---|---|---|
| 15m | ≈ 5% | ≈ 224k |
| 1h | ≈ 18% | ≈ 224k |
| 4h | ≈ 42% | ≈ 224k |
| 1D | ≈ 73% | ≈ 223k |
3 · The auto-disable rule
Any symbol whose walk-forward precision@N falls below 0.25 is switched off: the model stops serving magnets on it, and the symbol is displayed as auto-disabled everywhere it appears, including the public proof page. Re-enabling requires recovering above 0.30 across consecutive clean calibration cycles (hysteresis, so a symbol cannot flap on and off around the line). The validation target the model is held to is 0.40. The verdict is binding only on an adequate sample: a symbol below the line on fewer than n = 20 scored windows is also switched off, but labeled insufficient-data rather than carrying a precision verdict, and each symbol’s n and Wilson 95% lower bound are published beside its state.
This is honesty enforcement, not decoration: a dashboard that only ever shows its winners is not publishing accuracy. Disabled symbols stay visible, labeled as off, with the record of why.
As of July 2026: 10 of 25 tracked symbols are auto-disabled, shown as off, not hidden. Current list and per-symbol status live on /proof.4 · Trim-quality metric (added 2026-07-03)
The position-alert surface grades its "trim" (partial de-risk) alerts under a separate, pre-registered rule, distinct from the headline alert accuracy: a trim is vindicated if price moves adversely for the position beyond ±0.3% within 4 hours of the alert (same horizon and same price source as the headline grading); a favorable move beyond that line counts against it as given-up gains; smaller moves are neutral and skipped, and every skip is accounted under its own reason, never silent. Trim results publish with their own sample size (n) and Wilson 95% lower bound once the sample matures.
Separate from headline accuracy by construction: trim outcomes are recorded under their own distinct event type, so the headline alert-accuracy roll-up structurally cannot ingest them; the separation is enforced by the data model, not by policy. And because the rule was registered before results accumulated, it cannot be re-fit after the fact to flatter the outcome.
Registered 2026-07-03. Track record accrues from that date; results surface with n and confidence bounds once the sample matures.5 · Changelog: every revision, on the record
Metric history is a trust asset. Every change to how a published number is defined or presented is disclosed here, including the ones that made our numbers look smaller.
Baseline re-derivation: the multiple-of-random-baseline headline retired
Early copy headlined accuracy as a multiple of a 10% random baseline. Internal review in mid-June 2026 flagged that multiple as a baseline artifact: because most liquidation dollars land near spot, a uniform 10% random line understates the real bar a model must clear, which flatters the multiple. The multiple was retired from every headline in favor of walk-forward precision@N, always published with its sample size (n) and Wilson 95% lower bound, plus off-spot and deep-magnet capture shown against the trivial boundary baseline (definitions 2·1-2·4). The raw multiples remain in the public stats payload for transparency, explicitly marked do-not-headline.
Track-record span, disclosed: the per-symbol public ledger begins 2026-06-22, the date per-symbol snapshots started being recorded on every calibration run; before it, only run-level global figures were journaled. The tracked universe also expanded from 15 to 25 symbols in late June 2026. Both facts mean the per-symbol spans shown on the proof page are short (≈11-12 days as of early July 2026), even though the run-level global walk-forward record extends back to 2026-04-30 (~64 days, 1,700+ calibration runs) and independent cross-exchange feed verification has run continuously since 2026-05-10. Nothing was deleted or reset. The underlying measurement journal is append-only; the finer-grained per-symbol instrumentation simply began later. Both spans are published side by side rather than quoting only the longer one.
Random-baseline uplift fields formally deprecated in the public payload
The uplift-vs-random-baseline fields in the public stats payload were moved under a deprecated_baselines block that carries the do-not-headline caveat inline, so the number can no longer be machine-read or quoted without shipping its warning. No consumer of the payload headlines a baseline multiple.
Trim-quality metric pre-registered and added
The trim-vindication rule (section 4) was pre-registered and switched on: trim alerts are graded under their own frozen rule, recorded under a distinct event type, and structurally excluded from headline alert accuracy. Added as a new metric; no existing definition changed.
Ruler-mix basis disclosed; pooled baseline corrected to the dollar-weighted mix
Since 24 July 2026 three symbols (ETH, SOL, XRP) are graded on a top-20% ruler while the rest use top-10%, so the pooled out-of-sample precision@N (published with n = 13 independent day-clusters and its Wilson 95% lower bound) has mixed ruler difficulties by design. The pooled random baseline served with it was the flat 0.10, which is wrong for the mix (the dollar-weighted mix is ≈0.146 at current composition). All pooled baseline comparisons now use the mixed figure (the flat 0.10 remains published for reference), and a same-ruler shadow (the identical served configurations re-graded on the uniform top-10% ruler over the same held-out days) is published alongside the pooled number, labeled descriptive and never headlined. No accuracy definition, gate, or target changed; the headline stays on the same governed basis it was set on. This entry discloses the basis, it does not move a number.
6 · Commitments
The rules this page operates under. They bind us, not you.
Definitions are frozen at launch
The definitions on this page are frozen as of Methodology v1.0 (July 2026). Any change appends a dated entry to the changelog above; a metric is never silently redefined under the same name.
The proof ledger is never reset
The measurement journal behind the proof page is append-only. Losing streaks, auto-disabled symbols, and downgrades stay on the record permanently; the track record is never restarted to look better.
Every number carries its n
Published rates ship with their sample size and Wilson 95% lower bound, computed on the de-overlapped effective n. A number without its n is not one of our numbers.
Descriptive, dated, checkable
Everything here is descriptive history, not a forecast. Figures on this page are dated as of July 2026; the live, continuously recomputed values are always on /proof.
7 · For technical diligence
Four disclosures for readers who grade the process, not the pitch. Each is a mechanism this system operates, stated with its dated measurement; the live values recompute continuously on /accuracy and /proof.
7·1Kill-gate doctrine: the promotion rule is code, not policy
Every strategy promotion in this system clears one gate, implemented once in the shared calibration library and imported by the evaluation loops: kill_gate(net_usd, ci95_low) — an edge is real only if net_usd > 0 AND ci95_low > 0: positive net outcome and a day-clustered 95% lower confidence bound still above zero, measured walk-forward on out-of-sample data. A strong point estimate whose cautious bound sits at or below zero does not graduate. The gate is symmetric — a previously promoted component that stops clearing is demoted on the same evidence rule — and it is never weakened to let a favorite through.
The gate has fired against the house’s own models, on the record. On 2026-07-29 the liq-shadow v5mk2 paper experiment measured to its verdict — 142 fills, fill rate 0.8235, net −$434.63, gate=False — and was retired. On 2026-07-27 the near-band placement model was retired from serving after its walk-forward capture graded 0.2723 against a trivial no-model boundary reference’s 0.9666 on identical windows, with all 36 candidate variants failing the same kill-gate. Both retirements are journaled with their full numbers; nothing was reset or hidden.
Gate as implemented:kill_gate(net_usd, ci95_low) in the shared calibration library every evaluation loop imports. Retirements journaled 2026-07-27 and 2026-07-29 in the append-only engineering journal.
7·2Trust-plane architecture: receipts, not renderings
Every published number is served from a measurement artifact, byte-equivalent: the rate, its sample size (n), its Wilson 95% lower bound, and the stamp of when it was measured are written by the measurement loop and read back verbatim — never re-derived at render time, so what you see is what was measured. When an artifact is missing or stale, the surface fails closed — “accuracy data temporarily unavailable — we don’t estimate” — rather than interpolating a plausible-looking figure. Benched symbols, losing cells, and downgrades stay visible in the same plane. And every outbound message this system sends (alerts, posts, emails) passes an owner-review gate before it ships.
Inspect the plane directly: /accuracy for the engine’s scored record, /proof for per-symbol liquidation-capture receipts.7·3The base-rate disclosure: we publish the measurement that questions us
The diligence differentiator is not a number that flatters us; it is the number that could incriminate us. The engine’s 3-class read (up / down / sideways) is continuously scored against the do-nothing base rate — the realized sideways share of the same resolved calls — bucketed per realized-volatility regime, and the residual (served rate minus base rate) is published with its 95% floor. A read that merely echoes the base rate adds nothing, and the page says so in plain language when that is what the measurement shows.
As of 2026-07-29 (n = 8,593 resolved calls, three regime buckets): the served 3-class rate matches the realized sideways base rate exactly — residual +0.0pp in all three buckets, verdict “matches the base rate”. That verdict is on /accuracy right now, recomputed on every loop round. No competitor in this category publishes the check that questions its own engine; the accuracy page is built around it. (The 3-class rate is a sideways-base-rate number — descriptive market context, never a directional-edge claim.)
Measured 2026-07-29 per regime (buckets of n ≈ 2.9k each); live residual and floor on /accuracy.7·4What we deliberately don’t claim
No forecasts. Everything published here is descriptive history — measurements of what happened against calls already made — not a forecast of what comes next, and the marketing surfaces are pinned that way in the test suite.
No blended performance. Hypothetical reference-desk figures (paper units, backtests, shadow streams) are labeled hypothetical wherever they appear and are never blended with, or presented as, real-capital results. There is no fund and no managed money behind these numbers; the only live capital this system touches is the operator’s own manual desk, whose results are not a product claim.
The reference desk itself is public end-to-end: /reference-desk publishes the full graded record of that stream — a hypothetical desk with no capital behind it, measured daily, every count shown with its sample size and its own floor, never hidden at zero.
No winners-only display. Symbols get benched when their grade slips (section 3): as of 2026-07-29, 11 of 37 tracked symbols are auto-disabled — shown as off, with the record of why, never hidden.
Benched count from the served calibration artifact, 2026-07-29; the current per-symbol status is always live on /proof.Now check the live numbers.
A methodology page is only as good as the numbers it points at. The proof page recomputes every figure above continuously: per symbol, with n, effective n, and Wilson bounds, disabled symbols shown as disabled.
Questions about a definition? support@hunterkiller.io, we answer methodology questions directly.