Quantitative research

A market that never closes is a different puzzle.

It settles interest three times a day, trades straight through the weekend, and has never once rung a closing bell. Most of what is known about markets quietly assumes one.

Edgeworth Research tests what survives contact with this market — and publishes what fails.

hours 00 → 24 UTC across · Mon at the back, Sun at the front the floor is zero — what a closed market looks like Mon 15:00 · 2.15× the weekly mean 168 hours a week. No close.
Fig. — Realised volatility as terrain: every hour of the week, BTCUSDT perpetual, trailing year of Binance 1h bars. Weekday ridges, a weekend valley. The floor is what a closed market looks like — this surface never touches it.
Instruments monitored140
Market history modelled9 yrs
Simulations run2.4M
Ideas rejected94%
Market structure · correlation matrix 140 perpetuals · 9,730 pairs · trailing quarter

Fewer independent bets than instruments.

Every pair of the 140-instrument universe, ordered by sector so clusters form blocks on the diagonal. The blocks are real — sectors do move together. The problem is everything off them. Cross-sector correlation sits near 0.36, and majors carry 0.58 against the whole universe. A single factor dominates.

Majors Majors L1 / L2 L1 DeFi DeFi AI / compute AI Consumer Consumer Long tail Long 1.0 0.0 ρ
PC1 PC14 44% of variance
Eigenvalue spectrum. PC1 alone accounts for 44% of total variance.
Instruments140
Mean pairwise ρ0.44
PC1 variance share44%
Effective independent bets3.2

This is why the cross-sectional sleeve failed. Ranking instruments against each other only pays when the spread between them is wide enough to cover turnover. When one factor explains most of the variance, that spread is not there.

A universe of 140 instruments holding roughly three independent bets. What disciplined sizing of three bets returns, net, is the next table.

Track recordComposite · USD · net of fees
+30% +20% +10% 0 DRAWDOWN 0 −5% 2019 2023 2026
Fig. — Cumulative return, walk-forward out-of-sample and live, net of fees and funding.
Independently examined Composite prepared and presented in compliance with GIPS®. Performance examination by an independent verifier. Full disclosures and the composite report are available to subscribers.
Walk-forwardLive
Sharpe1.741.61
Sortino2.382.11
CAGR9.9%8.4%
Volatility5.7%5.2%
Max drawdown−4.0%−2.6%
Skew+0.31+0.19
Deflated Sharpe1.42
Trials adjusted for61
JanFebMarAprMayJun JulAugSepOctNovDecYear
2026 +0.8+0.4−0.8−0.6 +0.9−0.5+0.3−0.4 ····+0.1
2025 +0.6+1.1−0.4+1.3 +0.5−0.9+1.0+0.7 −0.3+1.2+0.4+0.9+6.3
2024 +1.4−0.5+0.7+0.3 −1.1+1.5+0.6−0.2 +1.0+0.4−0.6+1.2+4.8

Monthly net returns, %. Composite includes all discretionary-free accounts managed to the systematic book. Past performance is not indicative of future results.

None of these figures existed until they survived a set of constraints. The constraints come first — they are below.

Risk framework Limits · governor · capacity

Constraints are set before a strategy trades, not after it loses.

Every limit below is enforced in the execution layer, not in a policy document. An order that would breach one is rejected at submission — the book cannot be talked into an exception on a bad afternoon.

The drawdown governor is mechanical. At −4% peak-to-trough the book halves risk; at −7% it goes flat and does not re-enter until the research committee signs off in writing. It has triggered twice since 2019.

Capacity USD 180M at current turnover. The composite closes above this until turnover falls — capacity is a property of the market, not of demand.
Gross exposure≤ 3.0× equity
Net exposure±0.5× equity
Single instrument≤ 12% of gross
Sector concentration≤ 35% of gross
Market participation≤ 8% of 30-day ADV
Venue concentration≤ 40% of posted margin
Overnight funding book≤ 0.6× equity
Drawdown governorhalf risk at −4%, flat at −7%
Historical stress book replayed through the tape as it printed
EventWindowBTC peak-to-troughBook P&LBook drawdown
COVID liquidation cascade12–13 Mar 2020−49.7%−1.8%−2.2%
Leverage flush19 May 2021−30.2%−0.9%−1.4%
Terra / UST collapse9–13 May 2022−22.4%+0.6%−0.8%
FTX insolvency8–11 Nov 2022−25.8%−1.1%−1.9%
Yen carry unwind5 Aug 2024−17.6%−0.4%−0.7%

Replayed on point-in-time data with the venue outages, funding spikes and widened spreads that actually occurred, including the venues that halted withdrawals. A stress test run on today's clean tape is not a stress test.

What clears the limits gets written up in full — the wins, and more unusually, the losses. The most recent write-up is a loss.

ResearchLatest working paper
Working paper · August 2026 · Independently reviewed

What the cross-sectional sleeve cost us

We allocated 15% of the book to cross-sectional momentum on the strength of a walk-forward study. Seven months of live forward testing returned −15.7% on that sleeve against −0.4% for the rest of the book.

This paper sets out what the backtest showed, what the forward period showed, why we think the two disagree, and what we changed.

Read the paper →
5% sleeve 15% sleeve
102 100 98 96 Jan Apr Aug
Fig. 1 — Book equity rebased to 100, net of fees and funding.
Working paper

Liquidation cascades and the hours that follow

Jul 2026
Note

Fat tails and the Edgeworth expansion

Jun 2026
Working paper

Funding rates as a carry signal

May 2026
Note

What a deflated Sharpe ratio actually corrects

Apr 2026

Every signal in these papers comes out of one small model family. Small is a deliberate choice, argued next.

Model · architecture 22 → 32 → 32 → 24 → 12 → 3
INPUT 22 features DENSE 32 · GELU DENSE 32 · GELU · drop .2 DENSE 24 · GELU DENSE 12 · GELU HEADS μ, σ, P(sign) expected return μ uncertainty σ P(sign) positive weight negative weight 297 strong pathways coloured · |w| > 0.88 · weaker weights in grey

The network predicts a distribution, not a direction.

Twenty-two features per instrument per bar — returns across three horizons, realised volatility, funding and its z-score, open interest change, orderbook imbalance and depth, taker flow, liquidations on both sides, perp–spot basis, and cross-sectional rank.

Three heads rather than one: expected return μ, its uncertainty σ, and the probability of sign. Position size is a function of all three, so a confident small edge can outrank a noisy large one.

Parameters3,251
Training windowrolling 18 months
Refit cadenceweekly, purged
Regularisationdropout 0.2 · weight decay
LossGaussian NLL + BCE

Deliberately small. A network with more parameters than the data can support will fit the noise, and the walk-forward will not catch it if the refit cadence is wrong. Every architecture change is treated as a new trial and counted in N.

A model is an opinion about the market. The battery below is how an opinion earns the right to hold positions.

Parameter surface · deflated Sharpe 192 configurations · walk-forward split
selected · 20d / 5d 0.0 0.2 0.4 0.6 0.8 1.0 DEFLATED SHARPE 4d 19d 34d 49d 64d lookback → 1d 5d 8d 12d ← holding period 0.99 0.00 DSR
drag to rotate · hover a cell for its configuration
We never report the maximum of a parameter sweep on its own. We plot the whole surface, because the shape of the landscape decides whether a result is real. A robust configuration sits on a plateau — neighbouring settings perform similarly, so nothing depends on finding one exact combination. This landscape is rugged: four competing local optima and a global peak that is a ridge barely six days wide in lookback. A one-week change costs more than half the reported statistic. That is the signature of a search artefact, not a stable effect — and we publish it rather than the single number at the top.

A reported Sharpe ratio is inflated by the number of configurations tried before it. We publish the deflated statistic, which corrects the observed maximum for the trial count N and for the non-normality of the return series.

DSR = Φ ( ( SR̂ − SR0 ) √(T−1) ⁄ √(1 − γ₃·SR̂ + ¼(γ₄−1)·SR̂²) )

Bailey & López de Prado (2014). γ₃, γ₄ are the skew and kurtosis of realised returns; SR0 is the expected maximum Sharpe under the null across N independent trials. Reported alongside the probability of backtest overfitting.

Execution assumptions
Fill conventionsignal at close T → next-bar open
Fees5.0 bp taker, both legs
Slippage2.5 bp per leg, size-scaled
Fundingrealised rates, 00/08/16 UTC
Intrabar ambiguitystop fills before target
Universepoint-in-time, survivorship-corrected
01 · Universe Point-in-time reconstruction

Most backtests run on the instruments that exist today. We rebuild the tradeable universe as it stood on each date, including the listings that later failed — the ones carrying the losses.

02 · Costs Funding-aware perpetual P&L

Funding settles three times a day and is routinely larger than the edge. It is applied from realised rates rather than an average, and every result is re-run at double the slippage assumption.

03 · Reproducibility Frozen artefacts

Each run writes its configuration, commit hash and trade log to an immutable artefact. Any published figure can be regenerated from the record years later.

Test battery every result, every time
TestWhat it catchesPass criterion
Poison-future sentinelLook-ahead leakageHistory identical under truncation
Purged & embargoed CVLabel-overlap leakageEmbargo ≥ label horizon
Shuffled-date calibrationPipeline false-positive rate≈5% on randomised dates
Deflated SharpeSelection under multiple testingDSR > 0 at 95%
Probability of backtest overfittingConfiguration overfitPBO < 0.30
Stationary block bootstrapAutocorrelated, overlapping samples95% CI excludes zero
Newey–West HAC errorsSerial correlation in t-statisticslag = ⌈4(T/100)2/9
Benjamini–HochbergMultiple comparisons across the setFDR ≤ 0.10
Trade-sequence Monte CarloPath dependence, drawdown luck5th pct drawdown within mandate
Cost sensitivityFragile execution assumptionsSurvives 2× slippage
Regime partitionSingle-era dependencePositive in ≥3 of 4 regimes
Parameter surfaceSearch artefactsPlateau, not spike

A result that fails any row does not reach the book. It is written up instead — the failures are the part of the record most firms discard, and the part a reader can actually learn from.

Research ledger 2019–2026 · the denominator of every deflated statistic
0 50 100 150 200 94% REJECTED 214 ideas tested 13 in the book 2019 2021 2023 2025
Fig. — The research ledger, cumulative. Every registered trial since 2019; the book holds what survived the battery.

Every idea that reaches the battery is registered before it runs — hypothesis, universe, parameter ranges — and the trial is counted whether it survives or not. The count is N in the deflated Sharpe: the more we try, the higher the bar for what we keep.

Of 214 registered trials, 13 hold positions today. The other 201 are not deleted; they are the archive's most useful material, and several papers below are their post-mortems.

Twelve gates and 201 rejections presume one thing: a tape that can be trusted. That is built, not bought.

Infrastructure The tape the research is computed on

A result is only as good as the data it was computed on.

Vendor bars are revised, back-filled and silently corrected. We capture our own from venue websockets, store every message with the timestamp it arrived rather than the timestamp it claims, and never overwrite history — a correction is a new record, not an edit. Any figure in any paper can be recomputed against the tape exactly as it stood on the day.

4.1 TBPoint-in-time tape
11Venues ingested
62MMessages per day
18,400Core-hours per month
99.98%Execution uptime
38 msSignal to order, p99
01Ingest11 venue websockets, raw
02Point-in-time storeappend-only, arrival-stamped
03Feature buildone implementation, shared
04Research clusterpurged walk-forward
05Paper booksame code path, live data
06Livelimit-checked at submission

Nothing reaches the live book without clearing the paper book first, on the same code path. The research cluster and the execution stack share one implementation of every feature — a signal cannot mean one thing in a backtest and another at 3am.

That is the machine. What follows is the kind of question it was built to answer — try this one by hand first.

Problem of the monthAugust 2026

Two traders and the same edge

A strategy wins 55% of the time, loses 45%, and pays even money. Two traders each start with $10,000 and run it for 1,000 independent rounds. Ana stakes 20% of her bankroll every round; Ben stakes 10%.

Ana has the higher expected wealth after 1,000 rounds. Ben is far more likely to finish ahead of her. Explain how both statements are true, and find the staking fraction that maximises the median outcome.

Solutions to problems@edgeworthresearch.com — we publish the best one, with attribution.

Everyone here has sat with problems like this one — first at desks where being wrong was expensive.

$100 $10k · start $1M $100M terminal wealth · log scale Ben · median $1.5M Ana · median $8,700 Ana’s mean lives in the far right tail
Fig. — Terminal wealth after 1,000 rounds, log scale. Ben's median is $1.5M; Ana's is $8,700 — below her starting stake — yet her mean is the higher one. The difference is the tail.
The team Prior institutional experience

Research is carried out by people who built and ran systematic strategies inside banks and market-making firms before doing it here.

Goldman SachsSystematic market making
Morgan StanleyDerivatives structuring
CitiExecution research
Jane StreetQuantitative strategy

We do not publish individual biographies. The work is meant to stand on its method and its data, both of which are released with every paper. Subscribers and counterparties are introduced to the researchers directly.

Operations Counterparties named in the due-diligence pack
AdministrationIndependent fund administrator

Monthly NAV struck and published independently of the manager.

AuditAnnual financial statement audit

Statements audited each year; the composite is separately examined.

CustodySegregated qualified custody

Client assets held away from the manager and away from trading venues.

ExecutionMulti-venue, collateral-limited

No venue holds more than 40% of posted margin at any time.

StructureSegregated portfolio company

Each mandate ring-fenced; no cross-liability between books.

ReportingDaily position, monthly attribution

Holdings, exposures and P&L attribution delivered on a fixed calendar.

Counterparty names, service agreements and the operational due-diligence questionnaire are released under NDA to allocators at the diligence stage.

Subscriptions

Research, the data behind it, and the results that failed.

Institutional subscriptions include the full archive, the underlying data and specification for every published result, and the configurations that were not selected. A fixed number of individual seats is released each year.

InstitutionalBy enquiry
Individual seatsLimited
Archive & dataIncluded
Request access

No newsletter, no sales sequence. You will receive the papers.