Four levels of programming

New levels of programming.

Four levels of programming

"Vibe coding" is a horrible expression because it took the most consequential phase change in how software is developed and reduced it to a "slur" that makes good engineers hide their tools and careless ones hide their output. I understand it wasn't the initial intention of the expression but it became a thought terminating label that is applied without scrutiny to everything that was done with (and times without) AI. This gave rise to other, worse expressions, like "meat proxy" or "slop", as if the times before LLM tooling weren't plagued with the problems that these words claim to describe.

Every engineering org now uses AI at some level, and I think there is a missing qualifier for how much AI is being used, and that's why every AI touched contribution gets filed under the same dismissive labels. Reviewers cannot tell the difference between code that the author understands from code they never actually read, so they start treating both with the same levels of suspicion. This quietly punishes the people doing careful and well reviewed agentic work the most, because they're the ones who care. They then hesitate to make any contribution outside of a directly assigned scope, and the team loses the contributions it should want more of.

The real problem was never that an LLM produced the code or wrote it, they get better every day. The issue is that no human can vouch for the output and these are two different things that these single "words" confound.

I'm proposing a small classification for code, built around accountability:

  • A0: human authored, with AI at most as a search engine or sounding board
  • A1: AI assisted, autocomplete, snippets, boilerplate etc but the human is the author
  • A2: Agent built and human-reviewed (ie: the dev can explain every line and is accountable as if they'd written it)
  • A3: Agent built and unreviewed (which is what "vibe coding" initially described), good for prototypes and throwaway tools. This contains agent built & agent reviewed.

Put this way, what we can actually agree to avoid is mountains of A3 code accumulating in places that demand A2 or better. Asking people to label their contributions themselves tells reviewers or downstream users how much scrutiny to apply and whether they want to take the risk or not.

The standard runs both ways though, if the author has to vouch for every line, the reviewer must vouch for every objection. Shouting "slop" is not a finding but a vibe because throwing a diff into an agent or "slop detector" and forwarding its verdict is A3 reviewing and worth exactly as much as the A3 code it's shouting at. So reviewing should carry the same burden as the code, and perhaps even more: name the specific defect or the label itself is unreviewed output.

Perhaps there are better labels and boundaries will blur with time but the underlying question won't: can the person proposing this code explain it if asked? Can someone else understand it? Is it doing what it is supposed to do correctly? Therefore these questions, and not the tool, should be what the team argues about.

This post is A1.