I’m interested in ways AI can make processes more effective, rather than merely “movement along the curve” of effort-vs-quality.
Many popular uses of AI boil down to, “You can put in way less effort if you’re willing to accept an inferior-quality result!” That occasionally has its place, but is generally uninteresting to me, because people tend to accept this tradeoff even when it’s not a good one, simply because they’re lazy. So it’s delightful for me when I find something that AI actually improves– i.e., a rightward shift of the effort-vs-quality curve– rather than just enabling laziness.
I was first alerted to this perhaps a year ago, when someone pointed out that coding agents are very good at fixing merge conflicts, and can usually do it unsupervised. (I cannot find the original source for this, although I suspect it was Mitchell Hashimoto.)
This was notable, as fixing merge conflicts is one of the most unpleasant and error-prone tasks there is, and agents generally were extremely bad at most stuff back then, compared to what they are today.
I tried it for a while, first letting the agent move one step at a time through each conflict, making it explain the conflict context and what it did to fix it, then gradually I started stepping back further and further. The error rate was close to zero– certainly lower than my error rate at doing the same thing. The LLM was indisputably better at rebasing than me. Now I just rebase stuff, say “fix conflicts”, and don’t even look at the result because I’m pretty sure it’ll be right.
Git history autism
As someone who has Git history autism, this was fascinating to me.
Git history autism is fairly common among good developers: obsession with having a “clean” commit history, with meticulous and well-organized commit messages, perfectly atomic commits, and a commit sequence that tells a clear story about how a change came to be. Git history autism is a classic “nerd trap”: 99% of the time, it has zero value. Your perfect Git history has no consumers; no one will ever read it, so making it “easy to read” is a pointless activity– but it triggers an autistic compulsion anyway.
This is an interesting point, because the idea that someone may at some point want to “check the git history” from years ago is widespread, but in practice, no one ever does it. This is largely due to the fact that beyond a certain level of analysis, Git history becomes impenetrable for humans.
Everyone feels like Git history is an important artifact that ought to be carefully managed, but on some level there’s always a little bit of embarrassment that you’re spending so much effort on something that’s ultimately pointless. Keeping a hyper-meticulous changelog of everything is a waste of time, because no one is ever going to read it again, including you. It just gets lost in the scrollback and disappears.
Humans can’t read Git history
Theoretically, everything you’d want to know about how a piece of code came to be is contained in the Git history. In practice, piecing it together is usually prohibitively time-consuming and tedious.
For example, suppose you want to answer the question, “Why was XYZ done this way?” This is a very useful question, especially in a larger code base. (Sometimes it resolves to “There’s no reason, it’s outdated, you can get rid of it”, which is awesome when you’re refactoring.)
git blame will tell you when a line was last edited; but then you have to check what that commit did, and what the commits around it did, whether related parts of the code were changed since then, what PR or ticket it was associated with, whether any alternatives were tried previously, and so on. Even a simple investigation quickly turns into a graph of commits and dependent changes, which becomes infeasible for humans to mentally track after very few steps.
Often the last change to a line was non-semantic, like renaming a variable or something. So you have to check what the previous commit that touched that line was, which hardly anybody knows how to do. Most human investigations into the history of a change would be aborted by something as simple as that.
Git history and LLMs
That’s right, no one will ever sift through that huge pile of text to extract a few crucial details that end up mattering later– wait a minute. You know who’s really good at that? Yes.
Unlike humans, LLMs can scan through huge amounts of text extremely fast. LLMs get value out of Git in a way that humans never could.
The question from above, “Why was XYZ done this way?”, becomes easy to answer. The complicated git commands, multiple commit hashes and their contents, and all the other things that humans are extremely bad at, LLMs can do effortlessly in a minute or two.
This has the effect of making Git history autism somewhat meaningful again. Old commits are now likely to be read again, and making commits clear and logical makes it easier for LLMs to understand them as well as humans.
Git commit messages
A commit has contents (actual code changes) and a description (commit message). Like any morally upstanding person, I was initially disgusted by LLM commit messages. Unbeknownst to me at the time, this initial reaction was a reflection of my rule of LLM output, which states that your LLM output should never be shown to another human. This was based on the assumption that the primary consoomer of Git commit messages is other developers.
However, if we apply the reframing that commit messages are mainly for the benefit of LLMs, then the baroque and floccinaucinihilipilificandum-dense writing style of LLMs becomes fine, and is arguably ideal. LLMs are unfazed by this writing style, and occasionally the wall of jargon embeds a nugget which is actually useful.
Commit messages have the useful property that they’re collapsed by default (showing only the “title”, i.e., 1st line), so this doesn’t even contribute unduly to context-bloating.
Conclusion
In my view, traversing version control history should be seen as a job for LLMs.
- LLMs should be encouraged to examine the historical context frequently (as far as I can tell, most popular agent frameworks already appear to include instructions to do that).
- Asking LLMs questions about commit history is often more efficient than checking it yourself, as they can explore far more context than you can, and far faster, too.
- LLM-written commit messages are often fine, and not necessarily a violation of the “LLM output rule”, because they’re not really meant for human consumption.