Back to all posts

AI Code Review Should Sound Like a Careful Engineer

Sep 22, 2026
8 min read
AI Code Review Should Sound Like a Careful Engineer

There is a strange cost to a queue of Pull Requests: it rarely looks like a problem.

One PR can be reviewed in the evening. The second between meetings. The third looks small, just dependencies. The fourth already passed CI, so it probably feels fine. Very quickly, review stops being focused engineering work and becomes a memory task: where exactly am I expected to pay attention?

The routine itself was not what bothered me. Review is part of the job. What bothered me was that the most repeatable risks require the most boring kind of attention. Did an auth guard disappear? Did unsafe innerHTML show up? Did a dependency update add an install lifecycle script? Was an export removed? Did an API, schema, or config contract change in a way the author may not have noticed?

These are not the places where you need a genius reviewer. They are the places where it is unpleasantly easy to miss something.

So I built a personal AI reviewer for GitHub Pull Requests. Not because I wanted to replace human review. And definitely not because I wanted another bot leaving comments like "Potential security issue." I wanted a tool that keeps a baseline level of risk attention running consistently, while still speaking to the author like a reasonable engineer.

I did not want a rule bot

The simplest version of this tool is obvious: scan the diff for patterns, leave comments.

That becomes bad review very quickly.

The machine sees a fragment. It does not know the whole product history, the migration path, the team's local agreement, the temporary rollout plan, or why the author made that trade-off. If it writes with too much certainty, it starts sounding smarter than it has earned the right to be.

At the same time, I did not want a harmless toy that is afraid to say something real. Good review does not have to be soft. If a PR removes a permission guard or adds process execution from user-controlled input, the reviewer should stop and ask for changes. But even then, the job of the comment is to move the author toward verification and a decision, not to drop a red flag and leave.

The principle became:

Automated review should be specific, checkable, and modest about what it knows.

That modesty is not politeness theater. It reflects the reality of review. The reviewer often sees the diff, but not the whole context.

The local reviewer

The tool lives locally, outside cloned repositories. That was the first practical decision. I did not want a normal repository cleanup, reinstall, or update to accidentally delete the system that watches my review queue.

It runs through macOS launchd: it starts after login and then runs a review cycle roughly once an hour. Each cycle finds GitHub Pull Requests where my account is listed as a requested reviewer.

This is not a global GitHub scanner. It is my personal review inbox. If someone explicitly asks me to review a PR, the tool can put it into its queue.

Before reviewing, it waits until the checks for the current head commit are green. That constraint matters. If CI is already failing, an automated reviewer tends to comment on noise: formatting, types, broken tests, obvious things the author needs to fix before a real review. I want the tool to review something close to a human-reviewable state, not compete with CI.

It also stores local state: which PRs it has already reviewed and which head SHA it saw at the time. If new commits are pushed to a previously reviewed PR, the agent reviews it again, even if the requested-reviewer event did not happen again.

That small detail changes the behavior. A PR is not a static document. The author may respond to feedback, rewrite part of the logic, update dependencies, or fix tests. A reviewer that only remembers "I have seen this PR" becomes blind to the most important part of the process: what changed after the first round.

What it does, and what it does not do

The main path is simple: the agent reads the diff and publishes comments on the relevant diff lines. If GitHub does not accept an inline comment — for example, because the diff position is no longer valid or the API cannot attach the note to that line — it falls back to a body-only review and preserves the meaning of the finding.

That matters more than it sounds. Bad automation often breaks on a small API detail and silently drops the useful signal. I would rather leave a body review than pretend there was nothing to say.

The checks are intentionally simple and explainable: dynamic code execution, unsafe innerHTML, process execution, install lifecycle scripts, removed exports, risky API/schema/config changes, and disappearing auth or permission guards.

This is not a complete model of code quality. It is not a senior engineer replacement. It is a first, repeatable pass over known classes of risk.

There is also a boundary around approval. The agent does not approve PRs with unresolved review threads; it leaves a comment instead. It can auto-approve safe dependency-bot PRs, but normal PRs are limited by the number of reviews per run.

I want automation to help, not turn into an uncontrolled conveyor belt of approvals.

The most important iteration was the comment style

Technically, building the loop was not the most interesting part. Finding PRs, checking requested reviewers, waiting for checks, reading diffs, and leaving reviews are all annoying in the details, but conceptually straightforward.

The most important iteration was the language of the comments.

Early automated review text tends to collapse into dry warnings:

"Potential security issue."

"Validate compatibility."

"This may break consumers."

Those comments can be technically correct and still be bad review. They do not tell the author what to check. They do not connect the concern to the specific change. They sound like a confident alarm with missing context.

So I rewrote the rules around how comments should be phrased. A comment should name the specific risk, attach it to the change, explain why the diff raised concern, and ask a question the author can answer.

For example, instead of "Validate compatibility," a better comment is:

Can you confirm whether this exported type is still used by downstream packages? I could be missing the migration path, but removing it here looks like it may break consumers outside this repo.

The value is not that the sentence became softer. The value is that it gives the author a concrete check: downstream usage and migration path.

Another example:

Can we make sure this HTML is sanitized before it reaches dangerouslySetInnerHTML? The diff makes the rendered value look closer to user-controlled content than before.

That does not scream "security bug." But the risk is direct: sanitization and user-controlled content.

One more:

Can you confirm this route still requires the same permission check? I see the guard moved out of this branch, and I want to make sure the access rule did not become implicit.

This is the tone I want. The reviewer does not say, "you removed auth." It says: I see a change, it looks like a risk, help confirm the invariant.

That is the difference between humanizing text and pretending to be human. The goal is not to disguise AI as a teammate. The goal is to automate the best parts of good engineering review: specificity, respect for context, an argument behind the concern, and the willingness to be wrong.

Why the control panel matters

Automation that runs somewhere in the background can become a source of anxiety very quickly.

Is it checking anything? Did it crash? Is it writing a comment right now? Did it skip a PR because checks were red, or because it did not see the requested reviewer? Has it already reviewed this SHA? Did a new commit appear after the last run?

If those questions are hard to answer, the tool saves time in one place and spends trust in another.

So the reviewer has a local browser control panel, available only on localhost. It shows whether the service is running, the launchd state, and the current process state. It has Start, Stop, Restart, Refresh, and logs.

The panel shows the timer until the next review cycle, the last run time, and whether a cycle is currently active. It has an event feed: PRs found, skips, errors, published reviews, and links to the PR or review.

It also shows counters and a tracked PR table: total tracked PRs, how many are visible, how many need rereview, which SHA was reviewed, and which SHA appeared later.

The logs are not hidden. I can open the service output and see what actually happened.

That is not decoration. Observability is part of the product. If automation leaves comments in your name, you should not need a terminal investigation to understand what it did.

The boundaries I want to keep

I do not want a reviewer that pretends to understand everything.

It does not know the product better than the team. It should not approve normal PRs just because it did not find a risk pattern. It should not argue with the author instead of a person. It should not turn a simple dependency update into a philosophical review if the known safe conditions are satisfied.

Its value is narrower: it looks consistently at classes of risk that are easy to miss, waits for a reasonable check state, remembers reviewed SHAs, returns after new commits, and makes its comments visible inside the conversation humans are already having.

Most importantly, it does not have to be perfect to be useful. But it does have to be inspectable. I need to see what it did, why it skipped a PR, which SHA it reviewed, where it posted the review, and when the inline-comment fallback became a body-only review.

Mature automation should not feel like magic. Magic is hard to debug.

What I built, really

I did not start with the idea of replacing review. I started with the feeling that some attention could become continuous.

A Pull Request is still a conversation between engineers. The author knows the context of the change. The reviewer sees risk from the outside. CI checks formal properties. An agent can add another layer: a repeatable pass over known dangerous areas, with comments that help the author respond in a useful way.

That is the shape I like. Not a cold judge. Not autopilot approvals. Not a bot that writes "Potential issue" and disappears.

More like a careful junior reviewer: arrives on time, remembers what it saw, asks specific questions, and leaves the decision to people.

The value is not that AI replaced a reviewer. The value is that attention to risk became regular, observable, and part of the process.

That is the kind of automation I am interested in: limited, checkable, useful to humans, and modest enough not to get in the way of a normal engineering conversation.

If you're building an AI reviewer of your own, I'd be curious to hear where you draw the line between useful vigilance and automated noise. Reach out and tell me what earned your trust, and what lost it.

Telegram

More than a blog post

I share frontend news and the reasoning behind it throughout the day. Pick the language that feels natural to you.

Need to discuss your project? Get in touch.