Skip to content
View ipezygj's full-sized avatar

Block or report ipezygj

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ipezygj/README.md

Ilpo Väätäinen

I measure whether reported numbers survive their own error.

Benchmark accuracies, leaderboard ranks, backtest Sharpes, deployment claims — they are published as point estimates, stripped of the uncertainty that produced them. My work is to recompute them from the raw data, attach the interval, and report what is left standing. Sometimes the answer is that the number holds. Sometimes the ranking was sampling noise, or the accuracy was the label's own algebra handed back.

Finland · remote · Finnish (native), English (professional)


Published

The Benchmark Number Is Not the Deployment Number Rank collapse and base-rate collapse, measured across four domains. 10.5281/zenodo.22109473

How the Choice of Sign Inventory Moves the Statistics of an Undeciphered Script Three published readings of one rongorongo corpus move the entropy statistic 4.1 times further than the structure it exists to detect — including the ligature artefact my own metric produced before I caught it. 10.5281/zenodo.22057706 (concept DOI, resolves to the current version) · code: rongorongo-catalogue-audit

How Much of a Published Statistic Belongs to the Choice Upstream of It The method behind the audits below, stated once: name the coding decision, assemble the codings other people published, recompute the statistic and its null under each, and report the movement against the effect the statistic exists to detect. Six worked cases across epigraphy, astronomy, archaeoastronomy, two leaderboards and clinical variant prediction — two of which return less than the audit set out to find — and a catalogue of the five ways the procedure misleads, each an error I made running it. 10.5281/zenodo.22162050

Audits — code, data, and the reproduction

dataco-late-delivery-audit Logistics ML's favourite number — ~97% accuracy at predicting late deliveries — is the label's own algebra. Leaked vs. clean reproduction (100% / 97.5% / 69%) and a census of 65 public works, 28 of them verifiably leaked.

swebench-rank-audit How much of a leaderboard ranking survives its own sampling error. Paired-resolution audit across four SWE-bench splits and MTEB.

preregistered-miss A preregistered hidden-test prediction, the miss that refuted it, and every number's receipt. Published because it failed.

ControlBattery — Lean 4 · mathlib The statistics a control battery rests on, proved rather than asserted: Šidák never exceeds Bonferroni, and Bonferroni's union bound needs no independence assumption. Kernel-checked, and CI rejects any proof resting on sorry.

Merged upstream

pmorissette/ffncalc_deflated_sharpe_ratio and calc_expected_max_sharpe (Bailey & López de Prado), plus the follow-up fix for a series with no dispersion.

uber/causalml — bootstrap confidence intervals for auuc_score() and qini_score().

hummingbot/hummingbot — Hyperliquid auth silently accepting wrong private keys; the KuCoin perpetual trade-stream topic and timestamp field; a Vertex testnet WebSocket URL.

Built end to end (source private — happy to walk through it live)

Ranger Sovereign Vault — delta-neutral funding-arbitrage vault deployed to Solana mainnet; honourable mention, 2026 hackathon. Babble Parenting — a published parenting app: sound recording, feeding timers, calendar, memory storage. Chimera — event-driven trading automation: real-time market data, risk controls, parallel strategy testing.


How I work

I direct AI tooling (Claude) through design, implementation and analysis, and I operate what I build. Every number on this page was recomputed from raw data before it was written down, and where the artefact is public the recomputation is public with it.

ipezygj2@gmail.com

Popular repositories Loading

  1. dataco-late-delivery-audit dataco-late-delivery-audit Public

    Audit: DataCo late-delivery 97% accuracy = the label's own algebra. Leaked-vs-clean reproduction (100%/97.5%/69%) + census of 65 public works (28 verifiably leaked).

    Python 5

  2. solana-sss-engine solana-sss-engine Public

    TypeScript 1

  3. rongorongo-catalogue-audit rongorongo-catalogue-audit Public

    How the choice of sign inventory moves the statistics of an undeciphered script (rongorongo): code, results, paper draft

    Python 1

  4. AI-Media-Processor-Pro AI-Media-Processor-Pro Public

    is a powerful desktop application that uses AI to transform your audio and video files. Process videos directly from YouTube or your local folders to create custom audio mixes, instrumental tracks,…

    Python

  5. hummingbot hummingbot Public

    Forked from hummingbot/hummingbot

    Open source software that helps you create and deploy high-frequency crypto trading bots

    Python

  6. ipezygj ipezygj Public

    Profile README — independent measurement audits: recomputing published numbers with their uncertainty attached.

    Python