Skip to content

About

Benchmark coding agents and agent harnesses on 226 deterministic file-editing tasks with byte-exact verification and a public leaderboard.

Topics

Resources

Stars

37 stars

Watchers

0 watching

Forks

Repository files navigation

Explicit Edit Benchmark

Explicit Edit Benchmark

Explicit Edit is an open benchmark for measuring how accurately coding agents and agent harnesses edit files across 226 deterministic, byte-exact tasks.

CI/CD status MIT license Hugging Face Dataset Benchmark leaderboard

226 tasks Accepted benchmark runs Benchmark contributors Accepted benchmark observations Models in the benchmark dataset Configurations in the benchmark dataset Accepted harness families

The tasks are small on purpose. None of them needs deep reasoning or domain knowledge: the agent finds the right text, changes it, and leaves every other byte as it was. That keeps the attention on what actually differs between setups, which is the model, the tools, and the harness around them.

There are 226 of them: replacements, insertions, deletions, copies, moves, large files, several file types, and Unicode edge cases. A verifier compares the result byte for byte.

What the tasks actually are: Benchmark tasks lists every task family, explains what each scale means, and shows how the generator increases context, matches, files, decoys, block size, and Unicode risk.

Where results live: the Benchmark Explorer shows current rankings, and the Hugging Face Dataset stores every accepted run.

The benchmark grows with the people who run it. Run it on your own harness and configuration, publish the result, and it joins the same database and counts towards the statistics.

There is room for more than this, too: longer and more involved edits are planned, closer to the work people do when they change software. Ideas are welcome as issues, pull requests are reviewed and merged when they help, and the author is open to discussing any of it.

Data contributors

Thank you to everyone who shares benchmark observations. Your work makes this public comparison possible.

Contributor Accepted runs Configurations
@ashokkumards 1 1
@dirac-run 4 4
@TreyThomasCodes 7 7
@user2221 5 5
@xuankunv1 6 6
@Yugimob 23 23

Accepted harnesses

Harness Accepted runs Configurations
anchor-edit 1 1
baseline-agent 30 28
bb 7 7
codex-cli-default 2 6
d3ara1n-pi-hashline-edit 1 1
dirac-default 4 4
dsh-code 2 6
dsh-standard 2 6
github-copilot-cli-default 2 6
jerryan-pi-hashline-edit 1 1
oh-my-pi-default 3 7
opencode-default 2 6
personal-pi-extensions-opencode 1 1
pi-agent-ide 26 28
pi-apply-patch 1 1
pi-better-edit 7 7
pi-better-read-edit 1 1
pi-codex-conversion 1 1
pi-codex-edit 1 1
pi-codex-minimal-tools 1 1
pi-codex-tools 1 1
pi-default 27 21
pi-edit-safe 1 1
pi-hash-anchored-edit 1 1
pi-hash-edit 1 1
pi-hashline-context-edit 1 1
pi-hashline-edit 1 1
pi-hashline-edit-pro 24 24
pi-hledit 1 1
pi-lean-edit 1 1
pi-lector 1 1
pi-mono-multi-edit 1 1
pi-openai-codex-compat 1 1
pi-semantic-edit 1 1
pi-str-replace-editor 1 1
pi-wayfinder 1 1

What you need

  • Linux or WSL2
  • Node.js 24 or newer
  • Python 3
  • Bubblewrap (bwrap)
  • the agent CLI you want to test, installed and logged in
  • access to the model you want to test
  • a free Hugging Face account to submit the result

Clone the repository, then install it and the Hugging Face CLI:

git clone https://github.com/alexshpunt/explicit-edit-benchmark.git
cd explicit-edit-benchmark
npm ci
python3 -m pip install --upgrade huggingface_hub
hf auth login

Pick your agent

The benchmark has ready adapters for these CLIs. Use the name in the --harness flag.

--harness Runs Binary
pi-default Pi pi
baseline-agent Pi with bash only and an empty system prompt pi
pi-agent-ide Pi Agent IDE pi
codex-cli-default Codex CLI codex
opencode-default OpenCode opencode
oh-my-pi-default Oh My Pi omp
github-copilot-cli-default GitHub Copilot CLI copilot
dsh-standard DeepSeek Harness, native tools dsh
dsh-code DeepSeek Harness, code mode dsh

The name in the flag is the published harness family, so a result lands in the leaderboard under the name you ran. The binary column is what actually runs; pass a different path with --command when yours is not on PATH. pi-default and pi-agent-ide both run pi, and differ only in the tools.

You install and sign in to that CLI yourself. The benchmark leaves your CLI alone, and your credentials stay yours: it copies only the file you point it at, for the length of a run.

For anything not in this list you write your own adapter. Benchmark automation and public data documents the config API, and examples/bb is a worked example for a harness that is more than one CLI call: it starts a server, runs another agent in a thread, and reads that thread's timeline. We ship the example and a smoke run, not a bb result.

Run and publish

Choose whether the observation is verified on GitHub or runs on your machine. Both paths use one command.

For a verified observation:

npm run benchmark -- run --official --harness pi-default \
  --model openai-codex/gpt-5.6-luna --thinking low

The first invocation checks gh and hf login, creates your public caller repository from the approved template when needed, copies your local Pi and Hugging Face credentials into GitHub Actions secrets through stdin, starts the workflow, and waits until the result is accepted. Credentials are never placed in command arguments. Use --task TASK_ID for another partial run, --caller-repository OWNER/REPO for an existing caller, or --no-wait to return after dispatch.

For an ordinary unverified observation:

npm run benchmark -- run --local --harness pi-default \
  --model openai-codex/gpt-5.6-luna --thinking low

Local mode keeps the existing smoke gate, runs the complete task set, validates the result, and opens the Dataset contribution. The Dataset marks approved GitHub workflow results as verified and ordinary local contributions as unverified. Both remain visible. See Official community runs for the trust boundary and recovery path.

Local options

The local command accepts the existing submission options. Its lower-level equivalent is:

npm run benchmark:submit -- --harness pi-default --model PROVIDER/MODEL --thinking low --concurrency 10

You can run that yourself, or hand your coding agent this repository link and let it do the work: the skills in .agents/skills/ know how to route a harness to your account, run the observation, and open the pull request. Ask it to publish a result and it will use them.

Flag What it means
--harness The harness from the table above.
--model The exact model id, spelled the way that CLI expects it.
--thinking The reasoning level, for example low, medium, or high.
--concurrency How many tasks run at once. Default 10. It changes how long a run takes, and under heavy contention it can also change which trials fail.

Then the command does the rest:

  1. builds a temporary configuration for the agent you picked, and checks the binary and its version;
  2. checks the Hugging Face login;
  3. runs one exact task for every selected configuration and compares the files byte for byte;
  4. runs all 226 tasks;
  5. exports the result, validates it, and opens a pull request on the Dataset.

If step 3 fails, the command stops there, so a misconfiguration costs you a minute instead of a full run. That check is part of the command rather than something you have to remember.

Every run sees the same 226 tasks. A run is grouped with others only when the rules match: how many recovery attempts it allows, how long one attempt may take, the task set, and the verifier. How many trials ran at once is recorded too, but it does not separate groups, because it is scheduling rather than a rule; it shows up in timings, and under contention sometimes in failures. If a harness needs longer, raise --timeout-seconds, and the result is compared with runs that used the same timeout.

Some agents need a provider route or a credential file. That is one more flag on the same command:

npm run benchmark:submit -- --harness codex-cli-default --model PROVIDER/MODEL --thinking high \
  --provider-file provider.json \
  --env-file private-env.json

provider.json holds routing metadata and no key. private-env.json holds the key itself. Keep both files private. The configuration guide explains each adapter's quirks, the runtime and mount rules, and the files each CLI expects.

If you keep your own benchmark.config.ts, submit it with --config benchmark.config.ts instead of --harness. Run npm run benchmark:submit -- --help for every option.

What gets published

The pull request holds benchmark facts: task results, exact versions, timings, tool-call categories, and a safe copy of the configuration recipe. It holds no credentials, local paths, prompts, model prose, raw commands, command output, sessions, or workspaces, and your Hugging Face token only ever goes to the Hub API.

Keep the failures, timeouts, and recovery attempts in the result, because a published number is only useful when it is the real one. If a row is wrong, fix it at its source and rerun instead of editing the exported file.

Show your result

Once your run is accepted, you can show its score in your own README:

[![Explicit Edit Benchmark](https://img.shields.io/endpoint?url=https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark/resolve/main/badges/pi-agent-ide.json&style=flat-square)](https://huggingface.co/spaces/alexshpunt/benchmark-explorer?card=harness%3Api-agent-ide%40latest)

The Dataset publishes one badge per harness family under badges/, named after the family. Clicking a badge opens the latest harness version’s card while keeping the full comparison visible. This is how they look right now:

Family Badge
pi-default Explicit Edit Benchmark
pi-agent-ide Explicit Edit Benchmark
codex-cli-default Explicit Edit Benchmark
opencode-default Explicit Edit Benchmark
oh-my-pi-default Explicit Edit Benchmark
github-copilot-cli-default Explicit Edit Benchmark
dsh-standard Explicit Edit Benchmark
dsh-code Explicit Edit Benchmark

Each badge shows that family's score across its accepted configurations, on the same scale the Explorer uses. The color follows the score: 90% and up is bright green, then green from 75%, yellow from 50%, orange from 25%, and red below that.

The badge is a small JSON file served by the Dataset, so shields.io renders it and it refreshes whenever acceptance rebuilds the Dataset.

Check the code without spending money

npm run check

It runs formatting, linting, type checks, the unit and integration tests, and one deterministic sandbox trial. No paid model calls are involved, so it is a cheap way to check a clone or a change before running anything real.

Where to read more

Document What it adds
Official community runs GitHub template, verified provenance, automatic acceptance, and submit-only recovery
Configuration guide Per-adapter model and auth details, provider routes, mounts, lower-level commands
Share a result The submission flow step by step, and what happens after you open the pull request
Benchmark automation and public data The config API, the normalized data format, and the Dataset views
Benchmark tasks Every task family, scale, language, Unicode variant, and generation rule
Methodology What the scores mean, how exact verification works, and how recovery works
Architecture Where each fact lives, and which layer owns what

Skills for coding agents

These are instructions that a coding agent loads on its own rather than documents to read. They live in .agents/skills/, and an agent that supports skills picks them up when a task matches, so it is enough to ask for the job.

Skill The job it covers
configure-codex-account Run Codex on a ChatGPT subscription, an API-key route, or a Chinese provider such as Z.AI or DeepSeek
configure-copilot-account Run Copilot on its own GitHub account, a provider key, or an OAuth-only plan through a local bridge
add-benchmark-harness Add an adapter for another agent CLI
publish-benchmark-observation Run a full observation, open the Dataset pull request, and report grades and usage
review-benchmark-candidate Review, validate, and accept a contributed result

About

Benchmark coding agents and agent harnesses on 226 deterministic file-editing tasks with byte-exact verification and a public leaderboard.

Topics

Resources

Stars

37 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages