Skip to content

The Work Around the Agent Is Engineering Work

By Maksym Kuzmitskyi
9 min read
The Work Around the Agent Is Engineering Work

A coding agent returns a plausible diff. The local tests pass. It has followed the requested format and written a sensible explanation.

There is still a fairly ordinary question left: what does that evidence actually establish?

Perhaps the tests cover only one service. Perhaps a hook added files after the diff was inspected. Perhaps another change already solved the problem, so the proposed update is now unnecessary. The output can be quite competent while the decision to ship it remains wrong.

That is the part of working with agents that increasingly interests me. Not another way to phrase the request. The conditions around the request that make a result possible to check, continue, and sometimes reject.

I have been designing and iteratively refining that environment as part of my engineering work. Repository instructions, connected tools, validation steps, review handling, and session handoffs became things to maintain rather than instructions to repeat in every conversation.

I would describe this as agentic workflow engineering. It is still software engineering. The responsibilities have moved around a little; they have not disappeared.

The diff is one artifact in a longer operation

Here is a deliberately generalized composite, not a reconstruction of one identifiable project.

Two services share a contract. An agent updates a dependency in one repository. The local build accepts the new types, and the tests confirm the behavior that repository knows about. A separate integration check finds that the other service still expects a different representation.

The first diagnosis is incomplete. Before making another patch, the engineer checks the actual package contents, current contracts, base branch, and related changes. A compatible solution has already arrived through another change. Continuing the original update would now introduce a regression rather than finish the task.

The useful result is to stop and close the obsolete proposal.

There is no new feature to celebrate. There is a clearer account of the system and one unsuitable change that did not enter it. I think that is worth recognizing as an engineering outcome, without pretending it measures a productivity gain.

The agent helped investigate. A person still had to decide what the evidence meant, correct the initial assumption, and change the workflow so freshness checks happened earlier next time.

My contribution is the operating arrangement

There is an unhelpful choice in some descriptions of AI-assisted work: either the engineer wrote everything personally, or they merely asked a chatbot to do it.

Neither explains what I am doing.

My contribution includes defining the task, organizing access to relevant context, choosing constraints, examining decisions, interpreting failures, and refining the next attempt. The agent can inspect sources, propose changes, run checks, and prepare reviewable artifacts inside that arrangement.

I do not need to claim manual authorship of every generated line to own those engineering decisions. Equally, a generated implementation does not make the environment around it automatically good.

I have integrated existing tools rather than built every tool involved. MCP connections help bring requirements and documentation into reach. Git, CLI/API access, tests, type checks, builds, and browser verification contribute different kinds of evidence. Each needs a clear reason to participate in the task.

The distinction matters when describing the work: integration, configuration, development, evaluation, and adoption are separate claims. I would rather use the smaller accurate verb than let a convenient headline imply the larger one.

Context needs authority and a lifetime

Giving an agent more text is not necessarily giving it better context.

Personal working preferences, shared engineering standards, repository rules, and task requirements answer different questions. A general suggestion should not override the API contract in the current code. A remembered convention should not defeat a test result observed now.

I use saved context to reduce repeated explanation and support continuation. I still expect the agent to verify current sources. Memory can become stale, and instructions can preserve a decision after the reason for it has changed.

For an active task, a compact continuation record might look like this:

Goal: accept the new response shape without changing existing callers Current observation: one caller still consumes the previous field Decision: retain the compatibility adapter for this release Verified: caller inventory and focused adapter tests Unverified: downstream deployment sequence Next step: confirm the rollout contract before removal

This is an illustrative record, not an internal artifact. It separates observations, decisions, checks, and uncertainty. A new session can challenge it against the repository instead of treating the summary as a complete memory of reality.

I do not promise identical behavior across models or sessions. The more defensible promise is a documented place to resume and a way to notice when it no longer matches the work.

Autonomy is partly a permissions problem

Written boundaries are useful: which operations are allowed, when to ask, and what must not be disclosed. They are not enforcement by themselves.

Tool credentials, available commands, environment isolation, branch protections, and executable gates also determine what can happen. A rule that says "do not release" is weaker than an environment where the agent cannot release without an accountable person's action.

The design depends on the task. Low-risk investigation and focused checks can usually proceed without a sequence of tiny approvals. Business requirements, architectural commitments, sensitive actions, and release decisions need the appropriate owner.

I want the agent to have enough room to work, not enough authority to settle every question it encounters.

This also makes a stop meaningful. If acceptance criteria are unclear or current evidence contradicts the request, continuing to produce code is not necessarily progress. The next useful artifact may be a question or a bounded proposal.

A small change deserves a discriminating check

Before the first edit, I ask the agent to locate the code that controls the behavior, state a local hypothesis, and choose a check that could disprove it.

That is less impressive than a long repository survey, but often more useful. The goal is not to read every nearby file before acting. It is to know which assumption the next small edit will test.

After that edit, run the focused check before expanding the change. If it fails, decide whether it exposed a local defect or showed that the hypothesis was wrong. Broader tests, lint, build, CI, and acceptance checks follow according to the risk and repository requirements.

The focused check does not replace the broader gates. It keeps us from spending a large patch on an assumption we could have disproved cheaply.

For a small, clear task, I do not need a specification document that repeats the ticket. For a risky or multi-session change, writing down assumptions and a plan before implementation can be valuable. The timing is part of the evidence: a plan written after the work cannot demonstrate that the work was specification-driven.

Tools have outputs that need their own review

A command can succeed and still change more than the task authorized.

Generators, package installation, hooks, formatting, and deployment tooling can all transform the artifact after the initial edit. The working diff we reviewed may not be the commit tree that a hook produced.

I have used that kind of failure to refine the workflow: inspect the result after automation, separate unrelated generated drift, and verify what is actually being proposed.

It is a fairly undramatic control. That is one of its useful qualities. It does not require the agent to recognize every possible reason a generator changed a file. It requires the scope of the result to be checked before it leaves the local operation.

The next two articles, Green Checks Only Cover What You Checked and The Command Succeeded. The Scope Changed., examine those boundaries in more detail.

Review is evidence attached to a changing object

I have also developed personal PR-review tooling with AI assistance. The work includes tracking revisions, avoiding repeated unchanged findings, dry-run checks, and guarding actions against stale state.

That is a statement about mechanisms I developed, not a claim that the tool is continuously operating today, deployed across a company, or safely replacing human approval.

The engineering problem is interesting on its own. A finding can remain relevant while its line number changes. A new commit can remove the risk but leave an old comment. A review decision can become stale when the object it described changes.

I want the automation to provide useful evidence without silently taking over accountability. A Review Belongs to a Commit follows that boundary, including the limitations of duplicate suppression and head checks.

Improvements need careful attribution

I participate in evaluating reusable agent skills against real engineering work. That does not make me the owner of an organization-wide transformation, and it does not establish that a particular framework caused a measured improvement.

A successful task can depend on repository instructions, tool access, a well-chosen test, current domain knowledge, or a person's correction. A skill may contribute, add ceremony, or simply describe a practice already present.

I try to keep those explanations apart. Record the outcome, corrections, residual risks, and what the available evidence can attribute. Unknown is a reasonable answer when the task does not support a stronger one.

What a Skill Changed, and What It Didn't will make that distinction concrete without inventing comparative results.

The useful specialization is not a new claim about the model

This work does not mean I train models, conduct ML research, or operate production MLOps. It also does not mean role-based agent configurations have become reliable multi-agent execution simply because the definitions exist.

It means I design the conditions in which coding agents participate in software delivery: context, tools, boundaries, checks, human decisions, and a way to learn from ordinary failures.

My frontend and full-stack background is part of that, not something to replace with a new label. Interfaces, data ownership, service contracts, tests, and release behavior provide the questions against which generated output gets evaluated.

The diff at the beginning can remain plausible. I still need to know what changed, what was checked, and who owns the unresolved decision. I am comfortable letting an agent do a substantial part of the work when those answers are available. I am less comfortable calling the work complete because the explanation sounds finished.

If your agent workflow is reliable until one ordinary exception appears, tell me which assumption the exception exposed. That is usually a more useful starting point than another prompt template.

About the author

Maksym Kuzmitskyi is a Senior Software Developer focused on React and TypeScript. His work includes enterprise applications, shared frontend components, Node.js integration, testing and accessibility. Alongside frontend work, he designs AI-assisted delivery workflows and has developed personal review tooling.

Telegram

More than a blog post

I share frontend news and the reasoning behind it throughout the day. Pick the language that feels natural to you.

Need to discuss your project? Get in touch.