The Lab · Debugging
git bisect run on 1,000 Commits: 10 Runs and 3 Wrong Answers
I planted a bug in a 1,000-commit repo. git bisect run found it in 10 runs, after three bad test scripts sent it to the wrong commit first.
Each line jumps to its section
- git bisect run found a planted bug in 1,000 commits with 10 test runs, a linear walk needed 389
- A test file that only existed in newer commits made bisect blame commit 2 instead of 613
- Without exit code 125, 41 broken builds made it blame commit 560; with it, 17 runs found 613
- Timing a slowdown once gave 26 different wrong culprits in 30 bisects, a median of 9 runs never missed
I planted a bug in a 1,000-commit repo and let git bisect run hunt it down. It found the right commit in 10 test runs and under half a second, where walking back one commit at a time took 389 runs.
It also blamed the wrong commit three separate times, with total confidence. My test script caused all of them, and none took more than two lines of shell to fix.
One Planted Bug in 1,000 Commits, Found in 10 Runs
The repo comes from a shell loop. Every commit appends a line to a changelog and rewrites a small build file, so no two commits are identical, and a few commits plant specific changes. The one that matters first is commit 613, which "simplifies" a price formatter and drops the zero padding on cents:
// commit 612
return euros + "," + pad(cents % 100) + " EUR";
// commit 613
return euros + "," + String(cents % 100) + " EUR";
formatPrice(1905) goes from 19,05 EUR to 19,5 EUR. Easy to miss in review, obvious on a receipt.
The test is four lines of Node with assert, and a tiny shell script runs it. Then:
git bisect start c1000 c1 # bad, then good
git bisect run ~/bisect/check.sh
git bisect reset # back to where you were
I tagged every commit c1 to c1000 so I could name them. Small side lesson: 53 of those 1,000 names were also valid short object hashes in that repo, and Git printed warning: refname 'c997' is ambiguous every time I used one. Don't name tags like hex.
Bisect ran the script 10 times, took 0.45 seconds over three repeats and printed commit 613 as the first bad commit. Ten is what binary search on 1,000 items should take, because 2 to the power of 10 is 1,024.
The naive alternative is checking out commits from the top until the test passes. That took 389 runs and 17.8 seconds, about 46 ms per step for the checkout plus Node starting up.
At 46 ms a step the gap barely registers, so scale it to a real project. If each step needs a 90-second build, 389 steps is close to 10 hours and 10 steps is 15 minutes. That's plain arithmetic on my numbers, not a benchmark, but it's the reason I reach for bisect before I reach for reading diffs.
When the bug only shows up in production and a local test can't reproduce it, this doesn't apply. There I bisect by deploy log instead, the way Debugging a Production Bug I Cannot Reproduce Locally describes.
The Test File That Didn't Exist Yet
You usually write the failing test after you notice the bug. In my repo that test, test.js, lands in commit 700, well after the bug in 613.
So the obvious script, node test.js, runs inside a working tree that changes with every checkout. On every commit older than 700 the file isn't there. Node exits with code 1 and Cannot find module, and bisect reads exit code 1 as "bad". After 10 runs it announced that commit 2 was the first bad commit, with exactly the same confident output as the correct run.
The fix is to keep the check outside the repo, so every checkout runs the same test:
#!/bin/sh
# ~/bisect/check.sh lives outside the repo
node ~/bisect/test-price.js
// ~/bisect/test-price.js
const assert = require("assert");
const { formatPrice } = require(process.cwd() + "/price");
assert.strictEqual(formatPrice(1905), "19,05 EUR");
assert.strictEqual(formatPrice(3500), "35,00 EUR");
The test reads the code through process.cwd() because bisect runs the script from the top of the working tree, so the current directory always holds whichever commit is checked out. That part worked on the first try.
What matters is that your script only fails for the bug. Anything else that exits with 1 gets counted as evidence. These are the codes, from the git-bisect docs, plus what I saw on Git 2.54:
| Exit code | What bisect does |
|---|---|
| 0 | marks the commit good |
| 1 to 127, except 125 | marks it bad |
| 125 | skips it, the commit can't be tested |
| 128 and up | aborts the whole run |
I tried two of the edges. A script that exits 130 (what Ctrl-C gives you) stopped the run with exit code 130 ... is < 0 or >= 128. A typo in the script path, which comes back as exit code 127, didn't mark a thousand commits bad either. Git 2.54 stopped with error: bogus exit code 127 for good revision. That's the safe outcome, though you still end up with nothing found.
Broken Builds Need Exit Code 125
Real histories have commits that don't build. My second repo has 41 of them, commits 560 to 600, where a helper file is missing a closing brace. Any test run on those commits crashes on a syntax error with exit code 1.
With the outside-the-repo script from above, bisect landed in that range, read the crash as the bug and blamed commit 560. Wrong again, and the output looked exactly like a real answer.
Exit code 125 tells bisect "this commit can't be tested, pick another one". So the script checks that the code loads before it tests anything:
#!/bin/sh
node -e 'require("./price")' 2>/dev/null || exit 125
node ~/bisect/test-price.js
That run took 17 test runs and 1.27 seconds and found commit 613. The 41 broken commits cost 7 extra runs.
If you already know where the history is broken, you can tell bisect before it starts. git bisect skip c559..c600 marked all 41 as untestable up front, and the plain script without the 125 line then found 613 in 9 runs. I still keep the 125 line, because the broken commits you know about are rarely all of them.
There's a limit worth knowing. If a skipped commit sits right next to the culprit, Git can't tell which of them came first. I forced that by skipping 612 and 613, and instead of one answer it printed three hashes under The first bad commit could be any of:. You still get a short list, just not a single commit.
I now write the 125 line first, before the test itself, even when I think every commit builds. For a compiled project the check is the build command. For a web app it might be npm run build or a type check, anything that separates "can't test" from "tested and failed".
Fewer broken commits also means fewer skips. A pre-commit hook that runs the type check keeps most of them out of history in the first place, and that's one of the hooks in The 6 Git Hooks I Copy Into Every New Repo.
Bisecting a Slowdown Instead of a Bug
Nothing is broken in commit 777. It just replaces a Set with indexOf:
// commit 776
return [...new Set(list)];
// commit 777
return list.filter((t, i) => list.indexOf(t) === i);
On 20,000 tags with 50 distinct values, the median of five calls went from about 0.21 ms to about 1.06 ms. That's five times slower, from a one-line change that reads like a cleanup.
"Good" and "bad" read oddly for speed, so bisect lets you rename them:
git bisect start --term-old=fast --term-new=slow c1000 c1
git bisect run ~/bisect/check-speed.sh
The output then says is the first slow commit. The check times the function and exits 1 above 0.5 ms, a line between the two warm numbers:
const { uniqueTags } = require(process.cwd() + "/tags");
const list = Array.from({ length: 20000 }, (_, i) => "tag-" + ((i * 7919) % 50));
const runs = Number(process.env.RUNS || 5);
const t = [];
for (let r = 0; r < runs; r++) {
const s = process.hrtime.bigint();
uniqueTags(list);
t.push(Number(process.hrtime.bigint() - s) / 1e6);
}
t.sort((a, b) => a - b);
process.exit(t[Math.floor(runs / 2)] > 0.5 ? 1 : 0);
I ran the whole bisect 30 times per setting. Timed once, 0 of 30 runs found commit 777, and they came back with 26 different wrong answers. With the median of 5 it was 29 of 30 (the miss said 776). The median of 9 found 777 all 30 times.
The single run failed because the first call in a fresh process is slow. Across 12 cold runs the fast version took 0.47 to 0.56 ms, right on top of my 0.5 line, while its warm median sat near 0.21. My guess is JIT warm-up, but the fix doesn't depend on the cause: time it several times and use the median.
And run the check on both ends before you bisect. Those 12 single runs on commit 776 would have shown the problem in seconds, since half of them landed above the line.
Bottom Line
git bisect run is the fastest debugging tool I know whenever a script can tell good from bad. 10 runs for 1,000 commits is the easy part. The script is where it goes wrong, and it goes wrong silently, with the same "first bad commit" message it prints when it's right.
So I write the script before I type git bisect start. It lives outside the repo and exits 125 whenever the code doesn't build. If it measures time, it takes a median and never trusts the first run. Before handing it to bisect I run it once by hand on the known-good commit and once on the known-bad one, and check that the exit codes are 0 and 1.
If a session goes sideways halfway through, git bisect log prints everything marked so far. Save that to a file and delete the wrong line. After git bisect reset, git bisect replay with the file puts you back where you were, minus the mistake.
After that, bisect does the searching and you only have to read one diff.
Which bug in your backlog could you describe in a five-line check script tonight?