Explicit Edit is an open benchmark for measuring how accurately coding agents and agent harnesses edit files across 226 deterministic, byte-exact tasks.
The tasks are small on purpose. None of them needs deep reasoning or domain knowledge: the agent finds the right text, changes it, and leaves every other byte as it was. That keeps the attention on what actually differs between setups, which is the model, the tools, and the harness around them.
There are 226 of them: replacements, insertions, deletions, copies, moves, large files, several file types, and Unicode edge cases. A verifier compares the result byte for byte.
What the tasks actually are: Benchmark tasks lists every task family, explains what each scale means, and shows how the generator increases context, matches, files, decoys, block size, and Unicode risk.
Where results live: the Benchmark Explorer shows current rankings, and the Hugging Face Dataset stores every accepted run.
The benchmark grows with the people who run it. Run it on your own harness and configuration, publish the result, and it joins the same database and counts towards the statistics.
There is room for more than this, too: longer and more involved edits are planned, closer to the work people do when they change software. Ideas are welcome as issues, pull requests are reviewed and merged when they help, and the author is open to discussing any of it.
Thank you to everyone who shares benchmark observations. Your work makes this public comparison possible.
| Contributor | Accepted runs | Configurations |
|---|---|---|
| @ashokkumards | 1 | 1 |
| @dirac-run | 4 | 4 |
| @TreyThomasCodes | 7 | 7 |
| @user2221 | 5 | 5 |
| @xuankunv1 | 6 | 6 |
| @Yugimob | 23 | 23 |
| Harness | Accepted runs | Configurations |
|---|---|---|
anchor-edit |
1 | 1 |
baseline-agent |
30 | 28 |
bb |
7 | 7 |
codex-cli-default |
2 | 6 |
d3ara1n-pi-hashline-edit |
1 | 1 |
dirac-default |
4 | 4 |
dsh-code |
2 | 6 |
dsh-standard |
2 | 6 |
github-copilot-cli-default |
2 | 6 |
jerryan-pi-hashline-edit |
1 | 1 |
oh-my-pi-default |
3 | 7 |
opencode-default |
2 | 6 |
personal-pi-extensions-opencode |
1 | 1 |
pi-agent-ide |
26 | 28 |
pi-apply-patch |
1 | 1 |
pi-better-edit |
7 | 7 |
pi-better-read-edit |
1 | 1 |
pi-codex-conversion |
1 | 1 |
pi-codex-edit |
1 | 1 |
pi-codex-minimal-tools |
1 | 1 |
pi-codex-tools |
1 | 1 |
pi-default |
27 | 21 |
pi-edit-safe |
1 | 1 |
pi-hash-anchored-edit |
1 | 1 |
pi-hash-edit |
1 | 1 |
pi-hashline-context-edit |
1 | 1 |
pi-hashline-edit |
1 | 1 |
pi-hashline-edit-pro |
24 | 24 |
pi-hledit |
1 | 1 |
pi-lean-edit |
1 | 1 |
pi-lector |
1 | 1 |
pi-mono-multi-edit |
1 | 1 |
pi-openai-codex-compat |
1 | 1 |
pi-semantic-edit |
1 | 1 |
pi-str-replace-editor |
1 | 1 |
pi-wayfinder |
1 | 1 |
- Linux or WSL2
- Node.js 24 or newer
- Python 3
- Bubblewrap (
bwrap) - the agent CLI you want to test, installed and logged in
- access to the model you want to test
- a free Hugging Face account to submit the result
Clone the repository, then install it and the Hugging Face CLI:
git clone https://github.com/alexshpunt/explicit-edit-benchmark.git
cd explicit-edit-benchmark
npm ci
python3 -m pip install --upgrade huggingface_hub
hf auth loginThe benchmark has ready adapters for these CLIs. Use the name in the --harness flag.
--harness |
Runs | Binary |
|---|---|---|
pi-default |
Pi | pi |
baseline-agent |
Pi with bash only and an empty system prompt | pi |
pi-agent-ide |
Pi Agent IDE | pi |
codex-cli-default |
Codex CLI | codex |
opencode-default |
OpenCode | opencode |
oh-my-pi-default |
Oh My Pi | omp |
github-copilot-cli-default |
GitHub Copilot CLI | copilot |
dsh-standard |
DeepSeek Harness, native tools | dsh |
dsh-code |
DeepSeek Harness, code mode | dsh |
The name in the flag is the published harness family, so a result lands in the leaderboard under the name you ran. The binary column is what actually runs; pass a different path with --command when yours is not on PATH. pi-default and pi-agent-ide both run pi, and differ only in the tools.
You install and sign in to that CLI yourself. The benchmark leaves your CLI alone, and your credentials stay yours: it copies only the file you point it at, for the length of a run.
For anything not in this list you write your own adapter. Benchmark automation and public data documents the config API, and examples/bb is a worked example for a harness that is more than one CLI call: it starts a server, runs another agent in a thread, and reads that thread's timeline. We ship the example and a smoke run, not a bb result.
Choose whether the observation is verified on GitHub or runs on your machine. Both paths use one command.
For a verified observation:
npm run benchmark -- run --official --harness pi-default \
--model openai-codex/gpt-5.6-luna --thinking lowThe first invocation checks gh and hf login, creates your public caller repository from the approved template when needed, copies your local Pi and Hugging Face credentials into GitHub Actions secrets through stdin, starts the workflow, and waits until the result is accepted. Credentials are never placed in command arguments. Use --task TASK_ID for another partial run, --caller-repository OWNER/REPO for an existing caller, or --no-wait to return after dispatch.
For an ordinary unverified observation:
npm run benchmark -- run --local --harness pi-default \
--model openai-codex/gpt-5.6-luna --thinking lowLocal mode keeps the existing smoke gate, runs the complete task set, validates the result, and opens the Dataset contribution. The Dataset marks approved GitHub workflow results as verified and ordinary local contributions as unverified. Both remain visible. See Official community runs for the trust boundary and recovery path.
The local command accepts the existing submission options. Its lower-level equivalent is:
npm run benchmark:submit -- --harness pi-default --model PROVIDER/MODEL --thinking low --concurrency 10You can run that yourself, or hand your coding agent this repository link and let it do the work: the skills in .agents/skills/ know how to route a harness to your account, run the observation, and open the pull request. Ask it to publish a result and it will use them.
| Flag | What it means |
|---|---|
--harness |
The harness from the table above. |
--model |
The exact model id, spelled the way that CLI expects it. |
--thinking |
The reasoning level, for example low, medium, or high. |
--concurrency |
How many tasks run at once. Default 10. It changes how long a run takes, and under heavy contention it can also change which trials fail. |
Then the command does the rest:
- builds a temporary configuration for the agent you picked, and checks the binary and its version;
- checks the Hugging Face login;
- runs one exact task for every selected configuration and compares the files byte for byte;
- runs all 226 tasks;
- exports the result, validates it, and opens a pull request on the Dataset.
If step 3 fails, the command stops there, so a misconfiguration costs you a minute instead of a full run. That check is part of the command rather than something you have to remember.
Every run sees the same 226 tasks. A run is grouped with others only when the rules match: how many recovery attempts it allows, how long one attempt may take, the task set, and the verifier. How many trials ran at once is recorded too, but it does not separate groups, because it is scheduling rather than a rule; it shows up in timings, and under contention sometimes in failures. If a harness needs longer, raise --timeout-seconds, and the result is compared with runs that used the same timeout.
Some agents need a provider route or a credential file. That is one more flag on the same command:
npm run benchmark:submit -- --harness codex-cli-default --model PROVIDER/MODEL --thinking high \
--provider-file provider.json \
--env-file private-env.jsonprovider.json holds routing metadata and no key. private-env.json holds the key itself. Keep both files private. The configuration guide explains each adapter's quirks, the runtime and mount rules, and the files each CLI expects.
If you keep your own benchmark.config.ts, submit it with --config benchmark.config.ts instead of --harness. Run npm run benchmark:submit -- --help for every option.
The pull request holds benchmark facts: task results, exact versions, timings, tool-call categories, and a safe copy of the configuration recipe. It holds no credentials, local paths, prompts, model prose, raw commands, command output, sessions, or workspaces, and your Hugging Face token only ever goes to the Hub API.
Keep the failures, timeouts, and recovery attempts in the result, because a published number is only useful when it is the real one. If a row is wrong, fix it at its source and rerun instead of editing the exported file.
Once your run is accepted, you can show its score in your own README:
[](https://huggingface.co/spaces/alexshpunt/benchmark-explorer?card=harness%3Api-agent-ide%40latest)The Dataset publishes one badge per harness family under badges/, named after the family. Clicking a badge opens the latest harness version’s card while keeping the full comparison visible. This is how they look right now:
| Family | Badge |
|---|---|
pi-default |
|
pi-agent-ide |
|
codex-cli-default |
|
opencode-default |
|
oh-my-pi-default |
|
github-copilot-cli-default |
|
dsh-standard |
|
dsh-code |
Each badge shows that family's score across its accepted configurations, on the same scale the Explorer uses. The color follows the score: 90% and up is bright green, then green from 75%, yellow from 50%, orange from 25%, and red below that.
The badge is a small JSON file served by the Dataset, so shields.io renders it and it refreshes whenever acceptance rebuilds the Dataset.
npm run checkIt runs formatting, linting, type checks, the unit and integration tests, and one deterministic sandbox trial. No paid model calls are involved, so it is a cheap way to check a clone or a change before running anything real.
| Document | What it adds |
|---|---|
| Official community runs | GitHub template, verified provenance, automatic acceptance, and submit-only recovery |
| Configuration guide | Per-adapter model and auth details, provider routes, mounts, lower-level commands |
| Share a result | The submission flow step by step, and what happens after you open the pull request |
| Benchmark automation and public data | The config API, the normalized data format, and the Dataset views |
| Benchmark tasks | Every task family, scale, language, Unicode variant, and generation rule |
| Methodology | What the scores mean, how exact verification works, and how recovery works |
| Architecture | Where each fact lives, and which layer owns what |
These are instructions that a coding agent loads on its own rather than documents to read. They live in .agents/skills/, and an agent that supports skills picks them up when a task matches, so it is enough to ask for the job.
| Skill | The job it covers |
|---|---|
configure-codex-account |
Run Codex on a ChatGPT subscription, an API-key route, or a Chinese provider such as Z.AI or DeepSeek |
configure-copilot-account |
Run Copilot on its own GitHub account, a provider key, or an OAuth-only plan through a local bridge |
add-benchmark-harness |
Add an adapter for another agent CLI |
publish-benchmark-observation |
Run a full observation, open the Dataset pull request, and report grades and usage |
review-benchmark-candidate |
Review, validate, and accept a contributed result |
