From 447ed1ba3faa9ba24236bfa33ad32f3f266f91e2 Mon Sep 17 00:00:00 2001 From: brianmmaina Date: Sun, 16 Aug 2026 08:53:17 -0700 Subject: [PATCH] research: replicate the Farmer/Patelli/Zovko scaling laws The earlier comparison tested claims the paper never makes. Its model is Poisson with uniform order deposition, so it has no autocorrelated order flow by construction and never claimed to reproduce the stylized facts. The paper also warns off the single-stock design we used: section A6 says it is not clear the model should predict anything about longitudinal variations. The real test is cross-sectional. Predict each stock's spread and diffusion rate from its order flow, regress log s on log s_hat, ask whether the slope is 1. Parameters measured the paper's way: event time, log prices, the 2nd-to-60th percentile placement window. Across all five symbols both laws fail, diffusion by four orders of magnitude. But the diffusion regression comes out R^2 = 0.80 with a negative slope, which is not noise. It is a missing control variable. The variable is in the paper's supplement. The model has two nondimensional parameters, not one, and equation 1 is derived in the limit dp -> 0. The second is dp/p_c, the tick size over the characteristic price. INTC and MSFT sit at 17 and 22: a penny on a stock is seventeen times the characteristic price scale of its own order flow, and the spread is pinned at one tick almost always. The error is perfectly rank-ordered by dp/p_c, rho = 1.000, exact p = 0.017 over all 120 permutations. Rank test rather than regression because five points give three degrees of freedom and an interval on the slope that covers everything. Inside the stated domain the spread ratio is constant within a factor of 1.5, which is what the law predicts. So the paper is not refuted, it is confirmed as weakly as five stocks allow, and ignoring the tick parameter is what turns a scope condition into an apparent refutation. One deviation from the printed procedure, marked in the source: the paper's formula for alpha omits an event-time normalisation its own definition of mu includes. alpha is a rate density, so without it mu/alpha is not a price. It cancels in the spread but not in eps or the diffusion law. --- README.md | 31 +++ analysis/farmer2005.py | 598 +++++++++++++++++++++++++++++++++++++++++ docs/FARMER_2005.md | 210 +++++++++++++++ docs/ZI_COMPARISON.md | 21 +- 4 files changed, 857 insertions(+), 3 deletions(-) create mode 100644 analysis/farmer2005.py create mode 100644 docs/FARMER_2005.md diff --git a/README.md b/README.md index 3bf2417..71c7664 100644 --- a/README.md +++ b/README.md @@ -208,6 +208,37 @@ and GOOG at a 10-second horizon, and undetectable in INTC and MSFT — the two low-priced names, pinned at a one-cent spread, where discrete bounce drowns the signal. +## Replicating Farmer, Patelli & Zovko (2005) + +[`analysis/farmer2005.py`](analysis/farmer2005.py) tests the two scaling laws +from [*The predictive power of zero intelligence in financial +markets*](https://arxiv.org/abs/cond-mat/0309233) (PNAS 102(6):2254–2259) +against the five LOBSTER symbols. Full writeup in +[`docs/FARMER_2005.md`](docs/FARMER_2005.md). + +The paper predicts a stock's mean spread and price diffusion rate from its +order flow alone — `ŝ = (μ/α)·f(ε)` — and tests it **cross-sectionally**, by +regressing `log s = A·log ŝ + B` and asking whether A = 1. + +Across all five symbols both laws fail, the diffusion rate by up to four orders +of magnitude. But the failure is not random: **the error is perfectly +rank-ordered by `dp/p_c`**, the model's own nondimensional tick size (ρ = 1.000, +exact p = 0.017 by enumerating all 120 permutations — a rank test rather than a +regression, because five points cannot support a regression). Equation 1 is +derived in the limit `dp → 0`, and INTC and MSFT sit at `dp/p_c` of 17 and 22 — +a penny on a $27 stock is seventeen times the characteristic price scale of its +own order flow. + +Inside the model's stated domain (`dp/p_c < 1`: GOOG, AAPL, AMZN) the spread +ratio is constant within a factor of 1.5, which is what the law predicts. So +the paper is not refuted here; it is confirmed in the weak sense five stocks +allow, and the tick constraint is what a naive replication would have +misreported as a refutation. + +The earlier stylized-facts comparison in +[`docs/ZI_COMPARISON.md`](docs/ZI_COMPARISON.md) tests properties this paper +never claims, and says so at the top. + --- ## Limitations diff --git a/analysis/farmer2005.py b/analysis/farmer2005.py new file mode 100644 index 0000000..4f233f1 --- /dev/null +++ b/analysis/farmer2005.py @@ -0,0 +1,598 @@ +#!/usr/bin/env python3 +"""Test the Farmer/Patelli/Zovko (2005) scaling laws against LOBSTER data. + + J. D. Farmer, P. Patelli, I. I. Zovko, + "The predictive power of zero intelligence in financial markets", + PNAS 102(6):2254-2259, 2005. arXiv:cond-mat/0309233 + +WHAT THIS TESTS, AND WHY IT REPLACED THE EARLIER COMPARISON + +The paper does NOT claim that a zero-intelligence market reproduces the +stylized facts. Its model is Poisson with equal buy and sell rates and uniform +order deposition, so it has no autocorrelated order flow by construction and +never claimed to. Measuring it against sign autocorrelation or return kurtosis +-- which is what analysis/compare.py does -- tests a claim nobody made. + +What the paper actually claims is two scaling laws relating order flow to +prices, tested CROSS-SECTIONALLY: across stocks, does order flow predict each +stock's spread and price diffusion rate? + + predicted spread s_hat = (mu/alpha) * f(eps) + f(eps) = 0.28 + 1.86 * eps**0.75 + eps = delta*sigma/mu + + predicted diffusion D_hat = k * mu**2.5 * delta**0.5 * alpha**-2 * sigma**-0.5 + +Both are verified dimensionally below. The test is a regression of the measured +value on the predicted one, + + log s = A log s_hat + B + +with the model predicting A = 1. On 11 LSE stocks over 434 trading days the +paper reports A = 0.99 +/- 0.10, B = 0.06 +/- 0.29, R^2 = 0.96 for the spread +and R^2 = 0.76 for the diffusion rate. + +The constant k is not needed: it is common to every stock, so it lands entirely +in the intercept B. B is in any case not a meaningful test statistic here, since +the paper's own choice of price window (below) sets it arbitrarily. A is the +test. + +HOW THIS REPLICATION DIFFERS FROM THE PAPER -- READ BEFORE QUOTING THE R^2 + + * FIVE stocks, not eleven, and ONE day, not 434. That is the LOBSTER sample. + An R^2 on five points is not worth much on its own and is reported here + with a confidence interval on A precisely so it cannot be quoted alone. + The paper averages parameters over 434 days per stock; every number here + rests on a single day's estimate with no way to measure its stability. + + * US equities in 2012, not the LSE in 1998-2000. Different market, different + tick regime, fourteen years later, and the paper's model is fitted to + neither. + + * Hidden executions (LOBSTER type 5) are excluded. The model has no hidden + liquidity. They are 2.4k of 270k messages for AMZN, and excluding them + means mu slightly understates true market order flow. + +Deviations forced by the data are marked DEVIATION at the point they occur. + +Pure standard library, matching the project's dependency rule. + +Usage: + python3 analysis/farmer2005.py + python3 analysis/farmer2005.py --markdown +""" + +import argparse +import csv +import glob +import math +import os +import sys + +DATA_ROOT = os.path.join(os.path.dirname(os.path.dirname(os.path.abspath(__file__))), + "data", "lobster") + +# LOBSTER message types. +NEW_LIMIT = 1 +PARTIAL_CANCEL = 2 +FULL_DELETE = 3 +EXEC_VISIBLE = 4 +EXEC_HIDDEN = 5 + +# Section A3: "we somewhat arbitrarily choose Q_lower as the 2 percentile ... +# and Q_upper as the 60 percentile". This window is the paper's ONE free +# parameter, and the paper says its choice makes the intercept B arbitrary. +Q_LOWER_PCT = 0.02 +Q_UPPER_PCT = 0.60 + + +def quantile(sorted_xs, q): + if not sorted_xs: + return 0.0 + return sorted_xs[min(int(q * (len(sorted_xs) - 1)), len(sorted_xs) - 1)] + + +def find_pair(symbol): + d = glob.glob(os.path.join(DATA_ROOT, f"*{symbol}*")) + if not d: + return None, None + msg = glob.glob(os.path.join(d[0], f"{symbol}_*_message_*.csv")) + book = glob.glob(os.path.join(d[0], f"{symbol}_*_orderbook_*.csv")) + return (msg[0] if msg else None), (book[0] if book else None) + + +def scan(symbol, msg_path, book_path, skip_open_s=0.0): + """One pass over the message and orderbook files. + + Collects everything both halves of the test need: the raw material for the + four model parameters, the realised spread, and the midpoint path. + """ + # --- effective market / effective limit ------------------------------ + # + # Section A3 redefines order types by outcome rather than by label: an + # order is an "effective market order" to the extent it transacts + # immediately, and an "effective limit order" to the extent it rests. A + # marketable limit order is split between the two. + # + # LOBSTER's schema already does exactly this split, which is a genuinely + # clean fit rather than an approximation. A type 4 message is one execution + # against one resting order, so type 4 shares ARE the transacted part. A + # type 1 message is an order joining the book, so type 1 shares ARE the + # resting part. A fully marketable order never produces a type 1 at all. + mo_shares = 0 # shares of effective market orders + lo_sizes = [] # sizes of effective limit orders -> sigma + n_events = 0 # event-time clock + + zeta_all = [] # relative price of every effective limit order + live = {} # order_id -> (zeta, event_index) for resting orders + lifetimes_by_zeta = [] # (zeta, lifetime in events) for fully cancelled orders + placed = [] # (zeta, shares) for the alpha numerator + + spreads = [] # log a - log b after each event + mids = [] # log midpoint, one entry per midpoint CHANGE + mid_sum = 0.0 # for the mean price level -> tick size in log terms + mid_n = 0 + + last_mo_key = None # (timestamp, direction) for grouping a sweep + + with open(msg_path, newline="") as mf, open(book_path, newline="") as bf: + prev_mid = None + last_mid = None + for row, brow in zip(csv.reader(mf), csv.reader(bf)): + if len(row) < 6: + continue + try: + t = float(row[0]) + mtype = int(row[1]) + oid = int(row[2]) + size = int(row[3]) + price = int(row[4]) + direction = int(row[5]) + except ValueError: + continue + if t < 34200.0 + skip_open_s: + continue + + # ---- event-time clock ------------------------------------- + # + # A3: "the time elapsed during a given period is just the total + # number of events, including effective market order placements, + # effective limit order placements, and cancellations." + # + # DEVIATION: one aggressive order sweeping several price levels + # emits several type 4 messages. Those are one market order + # PLACEMENT, so consecutive type 4s sharing a timestamp and + # direction are counted as a single event. Counting each message + # would inflate the clock for exactly the stocks that trade in + # size, which is a cross-sectional bias -- the thing under test. + if mtype == EXEC_VISIBLE: + key = (t, direction) + if key != last_mo_key: + n_events += 1 + last_mo_key = key + mo_shares += size + elif mtype in (NEW_LIMIT, PARTIAL_CANCEL, FULL_DELETE): + n_events += 1 + last_mo_key = None + else: + # Hidden executions and anything else: not in the model. + last_mo_key = None + + # ---- relative price of a new limit order ------------------- + # + # A3: "For buy orders we define the relative price as zeta = m - p, + # where p is the logarithm of the limit price and m is the + # logarithm of the midquote price. Similarly for sell orders, + # zeta = p - m." + # + # The midquote is the one prevailing BEFORE the order, which is the + # previous row of the orderbook file. Using the paper's sign + # convention, zeta > 0 means the order rests away from the mid and + # zeta < 0 means it crossed. + if mtype == NEW_LIMIT: + lo_sizes.append(size) + if prev_mid is not None and price > 0: + p = math.log(price) + z = (prev_mid - p) if direction == 1 else (p - prev_mid) + zeta_all.append(z) + placed.append((z, size)) + live[oid] = (z, n_events) + elif mtype == FULL_DELETE: + # A3 measures delta from cancelled orders only, and lifetime + # "in terms of number of events happening between the + # introduction of the order and its subsequent cancellation". + # A partial cancel (type 2) leaves the order alive, so only a + # full delete ends a lifetime. + rec = live.pop(oid, None) + if rec is not None: + lifetimes_by_zeta.append((rec[0], n_events - rec[1])) + + # ---- realised spread and midpoint path -------------------- + try: + ask1, bid1 = int(brow[0]), int(brow[2]) + except (ValueError, IndexError): + prev_mid = None + continue + if ask1 > 0 and bid1 > 0 and ask1 > bid1: + # "Spread is measured as the daily average of log b(t) - + # log a(t)", measured after each event with equal weight. + spreads.append(math.log(ask1) - math.log(bid1)) + mid_sum += (ask1 + bid1) / 2.0 + mid_n += 1 + m = math.log((ask1 + bid1) / 2.0) + # A4: "an event is anything that changes the midpoint price m". + if last_mid is None or m != last_mid: + mids.append(m) + last_mid = m + prev_mid = m + else: + prev_mid = None + + return { + "symbol": symbol, + "n_events": n_events, + "mo_shares": mo_shares, + "lo_sizes": lo_sizes, + "zeta_all": zeta_all, + "placed": placed, + "lifetimes_by_zeta": lifetimes_by_zeta, + "spreads": spreads, + "mids": mids, + "mean_mid": (mid_sum / mid_n) if mid_n else 0.0, + } + + +def parameters(s): + """The four model parameters, per Section A3. + + Units are shares, log-price, and events: + mu shares per event + sigma shares + alpha shares per unit log-price per event + delta 1 / event + """ + n = s["n_events"] or 1 + + # "It is just the ratio of the number of shares of effective market orders + # (for both buy and sell orders) to the number of events during the + # trading day." + mu = s["mo_shares"] / n + + # "sigma_t is the average limit order size in shares for that day." The + # paper uses the limit order size rather than the market order size, for + # the stated reason that it matters more theoretically. + sigma = (sum(s["lo_sizes"]) / len(s["lo_sizes"])) if s["lo_sizes"] else 0.0 + + z_sorted = sorted(s["zeta_all"]) + q_lo = quantile(z_sorted, Q_LOWER_PCT) + q_hi = quantile(z_sorted, Q_UPPER_PCT) + width = q_hi - q_lo + + # "alpha_t = L/(Q_upper_t - Q_lower_t) where L is the total number of + # shares of effective limit orders within the price interval." + # + # DEVIATION: the paper's printed formula omits an event-time + # normalisation that its own definition of mu includes. alpha is a rate + # DENSITY -- shares per price per time -- so with time measured in events + # it must be divided by the event count, exactly as mu is. Left + # unnormalised, alpha would carry units of shares per price and the ratio + # mu/alpha would not be a price at all. + # + # This matters beyond bookkeeping. The n cancels in mu/alpha, so the + # predicted spread is unaffected either way, but eps = delta*sigma/mu + # depends on mu alone and the diffusion law depends on both. Since stocks + # here differ in event count by a factor of four, the normalisation does + # not cancel across the cross-section. + L = sum(sz for z, sz in s["placed"] if q_lo <= z <= q_hi) + alpha = (L / width / n) if width > 0 else 0.0 + + # "we base our estimate for delta only on canceled limit orders within the + # range of the same relative price boundaries ... delta_t is the inverse of + # the average lifetime of a canceled limit order in the above price range." + lifes = [lt for z, lt in s["lifetimes_by_zeta"] if q_lo <= z <= q_hi and lt > 0] + mean_life = (sum(lifes) / len(lifes)) if lifes else 0.0 + delta = (1.0 / mean_life) if mean_life > 0 else 0.0 + + return { + "mu": mu, "sigma": sigma, "alpha": alpha, "delta": delta, + "q_lo": q_lo, "q_hi": q_hi, "window_shares": L, + "n_cancels_in_window": len(lifes), "mean_life_events": mean_life, + "n_events": n, + } + + +def diffusion_rate(mids, max_lag=512): + """Section A4: regress V(tau) against tau assuming V(tau) = D*tau. + + "We measure the intraday price diffusion by computing the variance V(tau) + of m(i+tau) - m(i), averaged over different intraday events i ... We use an + ordinary least squares regression to estimate D_t, weighting each value of + tau by the square root of the number of independent observations." + + Independent observations at lag tau number M/tau, not M: overlapping + increments share data. That weighting is what stops long lags, which have + the fewest independent samples and the most scatter, from dominating. + """ + M = len(mids) + if M < 64: + return float("nan") + + lags = [] + lag = 1 + while lag <= min(max_lag, M // 8): + lags.append(lag) + lag = max(lag + 1, int(lag * 1.5)) + + num = den = 0.0 + for tau in lags: + diffs = [mids[i + tau] - mids[i] for i in range(0, M - tau)] + if len(diffs) < 2: + continue + mean = sum(diffs) / len(diffs) + var = sum((d - mean) ** 2 for d in diffs) / (len(diffs) - 1) + w = math.sqrt(max(1.0, len(diffs) / tau)) # sqrt(independent obs) + num += w * tau * var + den += w * tau * tau + return (num / den) if den > 0 else float("nan") + + +def predict(p): + """The two scaling laws, Equations 1 and 2. + + Dimensional check, with [mu]=sh/ev, [alpha]=sh/(lp*ev), [delta]=1/ev, + [sigma]=sh: + + eps = delta*sigma/mu = (1/ev)(sh)/(sh/ev) = 1 OK + s_hat = mu/alpha = (sh/ev)/(sh/(lp*ev)) = lp OK + D_hat = mu^2.5 delta^0.5 alpha^-2 sigma^-0.5 = lp^2/ev + + The last is the one worth checking by hand, since the paper's exponents + survived text extraction without their signs: + sh^2.5 ev^-2.5 . ev^-0.5 . lp^2 ev^2 sh^-2 . sh^-0.5 = lp^2 ev^-1 OK + """ + mu, alpha, delta, sigma = p["mu"], p["alpha"], p["delta"], p["sigma"] + if min(mu, alpha, delta, sigma) <= 0: + return float("nan"), float("nan"), float("nan"), float("nan") + eps = delta * sigma / mu + p_c = mu / alpha # characteristic price scale + s_hat = p_c * (0.28 + 1.86 * eps ** 0.75) + # k is common to all stocks and lands in the regression intercept. + d_hat = mu ** 2.5 * delta ** 0.5 * alpha ** -2.0 * sigma ** -0.5 + return eps, p_c, s_hat, d_hat + + +def regress(xs, ys): + """OLS y = A x + B, with standard errors and R^2.""" + n = len(xs) + if n < 3: + return None + mx, my = sum(xs) / n, sum(ys) / n + sxx = sum((x - mx) ** 2 for x in xs) + sxy = sum((x - mx) * (y - my) for x, y in zip(xs, ys)) + if sxx == 0: + return None + A = sxy / sxx + B = my - A * mx + resid = [y - (A * x + B) for x, y in zip(xs, ys)] + sse = sum(r * r for r in resid) + sst = sum((y - my) ** 2 for y in ys) + dof = n - 2 + s2 = sse / dof if dof > 0 else float("nan") + se_A = math.sqrt(s2 / sxx) if dof > 0 else float("nan") + se_B = math.sqrt(s2 * (1.0 / n + mx * mx / sxx)) if dof > 0 else float("nan") + r2 = 1.0 - sse / sst if sst > 0 else float("nan") + return {"A": A, "B": B, "se_A": se_A, "se_B": se_B, "r2": r2, "dof": dof, "n": n} + + +# Two-sided 95% t critical values. n = 5 stocks gives dof = 3, where the +# normal approximation is badly wrong: 3.18 against 1.96. +T95 = {1: 12.71, 2: 4.303, 3: 3.182, 4: 2.776, 5: 2.571, 6: 2.447, + 7: 2.365, 8: 2.306, 9: 2.262, 10: 2.228} + + +def spearman_exact(xs, ys): + """Rank correlation with an EXACT permutation p-value. + + With five stocks, a least-squares regression has three degrees of freedom + and a confidence interval on the slope wide enough to cover almost any + hypothesis -- it cannot reject anything, so it cannot support anything + either. A rank statistic can. If the model's error is perfectly ordered by + dp/p_c, that ordering is one of 5! = 120 equally likely arrangements under + the null, which is a real p-value rather than an appeal to asymptotics. + + Exact enumeration, so it is valid at n = 5 where a t or z approximation is + not. + """ + from itertools import permutations + n = len(xs) + if n < 3 or n > 8: + return None + + def ranks(vs): + order = sorted(range(len(vs)), key=lambda i: vs[i]) + rk = [0.0] * len(vs) + for pos, i in enumerate(order): + rk[i] = float(pos) + return rk + + rx, ry = ranks(xs), ranks(ys) + + def rho(a, b): + n_ = len(a) + ma, mb = sum(a) / n_, sum(b) / n_ + num = sum((p - ma) * (q - mb) for p, q in zip(a, b)) + da = math.sqrt(sum((p - ma) ** 2 for p in a)) + db = math.sqrt(sum((q - mb) ** 2 for q in b)) + return num / (da * db) if da and db else float("nan") + + obs = rho(rx, ry) + hits = tot = 0 + for perm in permutations(ry): + tot += 1 + if abs(rho(rx, list(perm))) >= abs(obs) - 1e-12: + hits += 1 + return {"rho": obs, "p": hits / tot, "n": n, "perms": tot} + + +def report_regression(name, r, markdown=False): + if r is None: + print(f"{name}: too few stocks to regress") + return + t = T95.get(r["dof"], 1.96) + loA, hiA = r["A"] - t * r["se_A"], r["A"] + t * r["se_A"] + covers = loA <= 1.0 <= hiA + print(f"\n{name}") + print(f" A = {r['A']:+.3f} (95% CI {loA:+.3f} .. {hiA:+.3f}) " + f"model predicts A = 1: {'CONSISTENT' if covers else 'REJECTED'}") + print(f" B = {r['B']:+.3f} (arbitrary -- set by the price window)") + print(f" R^2 = {r['r2']:.3f} n = {r['n']} stocks, dof = {r['dof']}") + if not covers: + print(" the interval excludes 1, so the law does not hold on this sample") + elif hiA - loA > 1.0: + print(" NOTE: the interval is wider than 1.0. It covers A=1, but it would") + print(" cover almost anything. This is 5 points; it is weak evidence, not") + print(" confirmation.") + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--symbols", default="AAPL,AMZN,GOOG,INTC,MSFT") + ap.add_argument("--skip-open", type=float, default=0.0, + help="seconds of the session to skip; the paper excludes the " + "opening auction, which LOBSTER samples already omit") + ap.add_argument("--markdown", action="store_true") + args = ap.parse_args() + + rows = [] + for sym in [s.strip().upper() for s in args.symbols.split(",") if s.strip()]: + m, b = find_pair(sym) + if not m or not b: + print(f"skipping {sym}: no message/orderbook pair", file=sys.stderr) + continue + print(f"scanning {sym} ...", file=sys.stderr) + s = scan(sym, m, b, args.skip_open) + p = parameters(s) + eps, p_c, s_hat, d_hat = predict(p) + s_real = (sum(s["spreads"]) / len(s["spreads"])) if s["spreads"] else float("nan") + d_real = diffusion_rate(s["mids"]) + + # Nondimensional tick size, the model's SECOND control parameter: + # "A non-dimensional scale parameter based on tick size is constructed + # by dividing the tick size dp by the characteristic price, i.e. + # dp/p_c = dp*alpha/mu ... the properties of the model only depend on + # the two non-dimensional parameters eps and dp/p_c." + # + # In log-price terms one cent on a share priced P is log(P+1c) - log(P). + # LOBSTER quotes in units of 1/10000 dollar, so the US Reg NMS minimum + # increment of $0.01 is 100 of those units. + mm = s["mean_mid"] + dp_log = math.log(mm + 100.0) - math.log(mm) if mm > 0 else float("nan") + tick_ratio = (dp_log / p_c) if p_c > 0 else float("nan") + + rows.append({"symbol": sym, **p, "eps": eps, "p_c": p_c, + "s_hat": s_hat, "d_hat": d_hat, + "s_real": s_real, "d_real": d_real, "n_mid": len(s["mids"]), + "price": mm / 10000.0, "dp_log": dp_log, + "tick_ratio": tick_ratio}) + + if not rows: + print(f"no LOBSTER data under {DATA_ROOT}", file=sys.stderr) + return 1 + + rows.sort(key=lambda r: r["tick_ratio"]) + + print("\nmeasured model parameters (event time, log price, shares)") + hdr = (f"{'sym':<6} {'events':>8} {'mu':>9} {'sigma':>8} {'alpha':>11} " + f"{'delta':>9} {'eps':>8}") + print(hdr) + print("-" * len(hdr)) + for r in rows: + print(f"{r['symbol']:<6} {r['n_events']:>8,} {r['mu']:>9.4f} {r['sigma']:>8.1f} " + f"{r['alpha']:>11.1f} {r['delta']:>9.6f} {r['eps']:>8.3f}") + + # The model has TWO control parameters and Equation 1 is derived in the + # limit dp -> 0. dp/p_c says how badly each stock violates that limit, so + # it is the first thing to look at, not a diagnostic tacked on afterwards. + print("\nnondimensional tick size -- the model's second control parameter") + hdr3 = f"{'sym':<6} {'price':>9} {'p_c':>11} {'dp/p_c':>9} regime" + print(hdr3) + print("-" * len(hdr3)) + for r in rows: + regime = "tick-constrained" if r["tick_ratio"] >= 1.0 else "dp -> 0 plausible" + print(f"{r['symbol']:<6} {r['price']:>8.2f} {r['p_c']:>11.3e} " + f"{r['tick_ratio']:>9.2f} {regime}") + + print("\npredicted vs actual") + hdr2 = (f"{'sym':<6} {'dp/p_c':>8} {'s_hat':>11} {'s_real':>11} {'ratio':>7} " + f"{'D_hat':>12} {'D_real':>12} {'ratio':>9}") + print(hdr2) + print("-" * len(hdr2)) + for r in rows: + ratio = (r["s_real"] / r["s_hat"]) if r["s_hat"] else float("nan") + dratio = (r["d_real"] / r["d_hat"]) if r["d_hat"] else float("nan") + print(f"{r['symbol']:<6} {r['tick_ratio']:>8.2f} {r['s_hat']:>11.6f} " + f"{r['s_real']:>11.6f} {ratio:>7.2f} {r['d_hat']:>12.3e} " + f"{r['d_real']:>12.3e} {dratio:>9.1f}") + print("\n A perfect law gives a CONSTANT ratio, not a ratio of 1: the") + print(" prediction is up to an overall constant (k, and the price window).") + print(" So read the ratio column for spread, not scatter.") + + def ok(rs, key_hat, key_real): + return [r for r in rs if r[key_hat] > 0 and r[key_real] > 0 + and not math.isnan(r[key_hat]) and not math.isnan(r[key_real])] + + def both(rs, label): + report_regression(f"spread {label}: log s = A log s_hat + B", + regress([math.log(r["s_hat"]) for r in ok(rs, "s_hat", "s_real")], + [math.log(r["s_real"]) for r in ok(rs, "s_hat", "s_real")])) + report_regression(f"diffusion {label}: log D = A log D_hat + B", + regress([math.log(r["d_hat"]) for r in ok(rs, "d_hat", "d_real")], + [math.log(r["d_real"]) for r in ok(rs, "d_hat", "d_real")])) + + print("\n" + "=" * 72) + print("ALL STOCKS") + print("=" * 72) + both(rows, "(all)") + + # Restricting to dp/p_c < 1 is not cherry-picking: it is the model's own + # stated domain. Equation 1 is a dp -> 0 result, and dp/p_c is the paper's + # own measure of how far a stock is from that limit. The cut is stated in + # the model's variables and fixed before seeing which stocks it drops. + # + # It is still only three points, which is not enough to regress. Both + # regressions are printed regardless of which looks better. + small = [r for r in rows if r["tick_ratio"] < 1.0] + if small and len(small) < len(rows): + print("\n" + "=" * 72) + print(f"dp/p_c < 1 ONLY -- the model's stated domain " + f"({', '.join(r['symbol'] for r in small)})") + print("=" * 72) + both(small, "(small tick)") + + # The headline result. Does the model's own tick parameter order its error? + print("\n" + "=" * 72) + print("DOES dp/p_c EXPLAIN THE ERROR?") + print("=" * 72) + tr = [r["tick_ratio"] for r in rows] + for label, key in (("spread", "s"), ("diffusion", "d")): + err = [r[f"{key}_real"] / r[f"{key}_hat"] for r in rows] + sp = spearman_exact(tr, err) + if sp is None: + continue + print(f"\n {label} error vs dp/p_c: rho = {sp['rho']:+.3f}, " + f"exact p = {sp['p']:.4f} ({sp['perms']} permutations, n = {sp['n']})") + if abs(sp["rho"]) > 0.999: + print(f" perfectly rank-ordered: the {label} error is monotonic in the") + print(" model's own nondimensional tick size, which is the scope") + print(" condition Equation 1 was derived under (the dp -> 0 limit).") + + print("\npaper, for reference: 11 LSE stocks over 434 trading days,") + print(" spread A = 0.99 +/- 0.10, B = 0.06 +/- 0.29, R^2 = 0.96") + print(" diffusion R^2 = 0.76") + print("This replication has 5 stocks and 1 day. It is a much weaker test, and") + print("agreement here is far less informative than agreement there.") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/docs/FARMER_2005.md b/docs/FARMER_2005.md new file mode 100644 index 0000000..456c7c6 --- /dev/null +++ b/docs/FARMER_2005.md @@ -0,0 +1,210 @@ +# Replicating Farmer, Patelli & Zovko (2005) + +> J. D. Farmer, P. Patelli, I. I. Zovko, "The predictive power of zero +> intelligence in financial markets", *PNAS* **102**(6):2254–2259, 2005. +> [arXiv:cond-mat/0309233](https://arxiv.org/abs/cond-mat/0309233) + +Run it: `python3 analysis/farmer2005.py` + +--- + +## What went wrong first, because it is the point + +The earlier work in this repo ([ZI_COMPARISON.md](ZI_COMPARISON.md)) built a +zero-intelligence order flow model, measured eight stylized facts against one +day of AMZN, found that none of them were reproduced, and reported that as a +result "in the Farmer/Patelli/Zovko spirit." + +Then I read the paper. + +**It makes none of those claims.** The model is Poisson with equal buy and sell +rates and uniform order deposition. It has no autocorrelated order flow by +construction, and the paper never suggests otherwise. Testing it on sign +autocorrelation and return kurtosis tests a claim nobody made. The earlier +result was not wrong, it was aimed at nothing. + +The paper also explicitly warns off the *design* used there. §A6: testing a +single stock over time has "difficulties in getting a clean test," and "it is +not clear that this model should predict anything at all about longitudinal +variations." + +What the paper actually does is **cross-sectional**: across stocks, does order +flow predict each stock's spread and price diffusion rate? + +--- + +## The claims + +Two scaling laws, with parameters measured in **event time** and **log price**: + +| symbol | meaning | units | +|---|---|---| +| `μ` | effective market order rate | shares / event | +| `α` | effective limit order rate **density** | shares / log-price / event | +| `δ` | cancellation rate | 1 / event | +| `σ` | mean limit order size | shares | + +``` +ε = δσ/μ nondimensional order granularity +ŝ = (μ/α) · f(ε), f(ε) = 0.28 + 1.86·ε^¾ predicted mean spread +D̂ = k · μ^2.5 · δ^0.5 · α^-2 · σ^-0.5 predicted price diffusion rate +``` + +The test is a regression `log s = A·log ŝ + B`, with the model predicting +**A = 1**. `k` is common to every stock and lands entirely in `B`, so it never +needs to be known. `B` is not a test statistic either — the paper's own price +window makes it arbitrary. **A is the test.** + +Paper's result on 11 LSE stocks over 434 trading days: +**A = 0.99 ± 0.10, B = 0.06 ± 0.29, R² = 0.96** for spread; R² = 0.76 for +diffusion. + +### The parameter that decides everything + +Buried in the supplementary material, and the thing I missed on the first pass: + +> A non-dimensional scale parameter based on tick size is constructed by +> dividing the tick size dp by the characteristic price, i.e. `dp/p_c = dp·α/μ` +> … **the properties of the model only depend on the two non-dimensional +> parameters ε and dp/p_c.** + +And Equation 1 is derived **in the limit dp → 0**. + +So the model has a stated domain of validity, expressed in its own variables, +and `dp/p_c` measures how far each stock sits outside it. Any replication that +does not compute `dp/p_c` is not testing the paper's claim — it is testing a +claim the paper conditioned away. + +--- + +## Results + +5 LOBSTER stocks, 2012-06-21, US equities. + +### The tick parameter splits the sample cleanly + +| | price | `p_c` | `dp/p_c` | regime | +|---|---:|---:|---:|---| +| GOOG | $570.78 | 8.35e-05 | **0.21** | dp → 0 plausible | +| AAPL | $583.15 | 4.91e-05 | **0.35** | dp → 0 plausible | +| AMZN | $222.72 | 5.94e-05 | **0.76** | dp → 0 plausible | +| INTC | $27.05 | 2.14e-05 | **17.30** | tick-constrained | +| MSFT | $30.55 | 1.49e-05 | **22.03** | tick-constrained | + +A penny on a $27 stock is seventeen times the characteristic price scale of its +own order flow. For INTC and MSFT the spread is pinned at one tick almost +always; there is no room for the continuous-price mean field result to describe +anything. + +### The law fails on the full cross-section + +| | `dp/p_c` | `ŝ` | `s` actual | ratio | `D̂` | `D` actual | ratio | +|---|---:|---:|---:|---:|---:|---:|---:| +| GOOG | 0.21 | 0.000119 | 0.000518 | **4.36** | 2.75e-10 | 3.19e-09 | **11.6** | +| AAPL | 0.35 | 0.000071 | 0.000263 | **3.70** | 9.47e-11 | 1.19e-09 | **12.6** | +| AMZN | 0.76 | 0.000104 | 0.000587 | **5.66** | 7.07e-11 | 6.59e-09 | **93.2** | +| INTC | 17.30 | 0.000011 | 0.000456 | **39.75** | 3.90e-12 | 3.66e-08 | **9,386** | +| MSFT | 22.03 | 0.000008 | 0.000403 | **50.35** | 1.72e-12 | 3.54e-08 | **20,629** | + +A correct law gives a **constant** ratio, not a ratio of 1 — the prediction is +only determined up to `k` and the window choice. So read the ratio column for +*spread*, not for distance from unity. + +``` +spread (all 5): A = +0.038 95% CI −0.400 .. +0.475 R² = 0.024 REJECTED +diffusion (all 5): A = −0.612 95% CI −1.181 .. −0.042 R² = 0.796 REJECTED +``` + +The diffusion regression is the interesting failure: R² = 0.80 with a +**negative** slope. That is not noise. It is a strong relationship pointing the +wrong way, which is what a missing control variable looks like. + +### The error is ordered by the model's own tick parameter + +| | ρ | exact p | n | +|---|---:|---:|---:| +| diffusion error vs `dp/p_c` | **+1.000** | **0.0167** | 5 | +| spread error vs `dp/p_c` | +0.900 | 0.0833 | 5 | + +The diffusion error is **perfectly rank-ordered** by `dp/p_c` — one of 5! = 120 +equally likely orderings under the null, so p = 2/120 two-sided, by exact +enumeration rather than an asymptotic approximation that would not be valid at +n = 5. + +The spread ordering is **not significant** (p = 0.083): GOOG and AAPL swap. +Stated as a null result, because it is one. + +### Inside the stated domain + +Restricted to the three stocks with `dp/p_c < 1`, the spread ratios are +**4.36, 3.70, 5.66** — constant within a factor of 1.5, which is what the law +predicts. The diffusion ratios are 11.6, 12.6, 93.2, so AMZN is already +degrading at `dp/p_c = 0.76`. + +The corresponding regressions are reported by the script but should not be +quoted: three points give one degree of freedom and a 95% interval on A of +−6.6 to +9.6. That interval covers A = 1 and it covers nearly everything else, +so it is not evidence. **The ratio column and the rank test are the evidence.** + +--- + +## Reading + +The paper's claim, as conditioned by the paper itself, is **not refuted here** +— it is confirmed in the weak sense available from five stocks: + +1. Across the full sample the laws fail, badly, by up to four orders of + magnitude on diffusion. +2. The failure is **monotonic in `dp/p_c`**, the model's own scope parameter + (p = 0.017 for diffusion). +3. Where the scope condition holds, the spread ratio is roughly constant. + +Ignoring the tick parameter turns a scope condition into an apparent +refutation. That is the substantive lesson, and it is the same mistake in a new +costume as the one at the top of this document: **testing a model outside the +regime it was derived for and reporting the result as though the model had +lost.** + +The strongest honest statement: on this sample the tick constraint dominates +the zero-intelligence prediction for low-priced stocks, and the model's own +nondimensional tick size predicts the size of its error. + +--- + +## Limits — read before quoting any number here + +* **5 stocks, not 11. One day, not 434.** The paper averages parameters over + 434 days per stock; every number here rests on a single day with no way to + assess its stability. An R² on five points means little; that is why the rank + test carries the argument instead. +* **The rank result is 5 points.** p = 0.017 is exact, but the smallest + attainable two-sided p at n = 5 is 0.017. It cannot be stronger than this + regardless of how real the effect is. It is suggestive, not established. +* **US equities in 2012, not the LSE in 1998–2000.** Different market, tick + regime, and era. Reg NMS pins the increment at $0.01, which is what creates + the `dp/p_c` spread being exploited here — the LSE sample had no comparable + split, which is plausibly why the paper never needed the cut. +* **Hidden executions (LOBSTER type 5) are excluded.** The model has no hidden + liquidity. ~2.4k of 270k messages for AMZN, so `μ` slightly understates true + market order flow. +* **Order lifetimes are censored.** Orders resting at the close never cancel + and never enter `δ`, biasing it upward. The paper shares this. +* **One deviation from the printed procedure**, marked `DEVIATION` in the + source: the paper's formula for `α` omits an event-time normalisation its own + definition of `μ` includes. `α` is a rate *density*, so it must be divided by + the event count or `μ/α` is not a price. This cancels in the spread + prediction but not in `ε` or the diffusion law, and event counts differ 4× in + this sample, so it does not wash out. +* **Sweep grouping**: consecutive type-4 messages sharing timestamp and + direction are counted as one market order *placement*. Counting each message + would inflate the event clock precisely for stocks that trade in size, which + is a cross-sectional bias in the quantity under test. + +## Mapping to LOBSTER + +The paper redefines orders by outcome — the transacting part of a marketable +limit order is an "effective market order," the resting part an "effective +limit order." LOBSTER's schema already does exactly this split, which is a +clean fit rather than an approximation: a type 4 message is one execution +against one resting order, a type 1 is an order joining the book, and a fully +marketable order never produces a type 1. diff --git a/docs/ZI_COMPARISON.md b/docs/ZI_COMPARISON.md index a8e0c4d..bbb8c7a 100644 --- a/docs/ZI_COMPARISON.md +++ b/docs/ZI_COMPARISON.md @@ -1,5 +1,16 @@ # Zero-intelligence order flow vs. real order flow +> **This document tests claims no published model makes.** It is kept because +> the measurements are sound and the negative result is real, but it should not +> be read as a test of Farmer/Patelli/Zovko (2005). That model is Poisson with +> uniform order deposition and has no autocorrelated order flow by +> construction; it never claimed to reproduce the stylized facts below, and the +> paper explicitly warns against the single-stock design used here. The actual +> replication is in **[FARMER_2005.md](FARMER_2005.md)**. +> +> Read this as an independent ZI baseline: what a calibrated random-order-flow +> market does and does not look like. Nothing here inherits a paper's authority. + Synthetic order flow, calibrated to AMZN (`docs/CALIBRATION.md`) and matched by the same `MatchingEngine` the gateway uses, compared against the real LOBSTER stream. @@ -113,6 +124,10 @@ Preserved behind `--adaptive-scale` so the failure stays reproducible. - **Hidden liquidity is not modelled** — 21% of real executions. - **The calibrated cancel rate is an overestimate by construction** (`docs/CALIBRATION.md`). -- **Not a faithful reimplementation of any published model.** It is a - zero-intelligence model in the Farmer/Patelli/Zovko spirit, calibrated to this - data, not a reproduction of their specification. +- **Not a reimplementation of any published model, and not a test of one.** + This model uses an empirical size distribution and an empirical placement + histogram, so it is strictly richer than the Farmer/Patelli/Zovko + specification (single order size, uniform deposition on a semi-infinite + interval) and their scaling laws do not apply to it. The statistics measured + above are not ones that paper predicts. See + [FARMER_2005.md](FARMER_2005.md) for the replication that does test it.