Judgment is codeable. Conviction is not.
I run most of my work through agents now. The research, the drafting, the first pass at nearly everything I used to do by hand. When I automate something these days I'm rarely asking whether it can be done. I'm asking what's left after.
The answer I keep hearing, from people who do roughly the work I do, is judgment. Taste, or instinct, depending on who's saying it. Automate the rest, but that part is yours, and it's the part that keeps you employable.
I think that's mostly wrong. I thought it was wrong enough to bet a company on the opposite, so I started Contextful on the premise that judgment is a thing you can codify, hand to a machine, and get back at scale.
Judgment is codeable. Conviction is not.
Judgment is pattern recognition, the read on which option is best, and that is what we are automating. Conviction is the harder thing underneath it: committing to a problem before it can be proven, and paying for it when you are wrong. It is easy to mistake having an opinion for conviction. A model can hand you a hundred good options and not care which one is right. It has no skin in the game. You do.
What judgment actually is
Most of the comfort in "judgment is safe" comes from leaving the word vague.
When I say someone has good judgment, I mean something fairly specific. Put a set of options in front of them and they pick well. They look at three designs and know which one will confuse people. Given a plan, they see where it breaks before anyone has built it. And on a piece of work they can usually say what's wrong before they can fully explain why.
That's evaluation. Given candidates, rank them.
I'm aware that's narrower than the whole of what people mean. Framing the problem in the first place, or generating the candidates nobody had thought of, is real work and it isn't ranking. I'll come back to it, because it turns out not to sit where you'd expect. But ranking is most of what the word covers in practice.
And evaluation is pattern recognition. The reason the senior person picks well is that they've seen several hundred versions of this before and the current one resembles the ones that went badly. That's the whole mechanism. The felt experience is instant and wordless, which is why we call it instinct.
Pattern recognition over a large body of prior cases is the single thing we've gotten best at automating. That's a description of what these models already are, not a forecast about where they're heading.
That's the definition I'm using. Accept the rungs and the conclusion follows mechanically.
It's how I work. When I run an agent over a long job now, I have another model grade the output against criteria I wrote down. It reads the work, applies the standard, and tells me where it failed. That's a judgment call, made without me, a few hundred times a day. The first time I set one up I expected to be disappointed. It's now load-bearing in how I build, and I trust it more than I trust my own attention at the end of a long day.
The obvious objection is that I still wrote the criteria, so what I automated was the checking and not the judgment. I thought that for a while. Then I looked at what the criteria actually were: a distillation of a few hundred earlier cases where I'd seen the same mistakes, written down in an afternoon. Setting the standard turned out to be the same evaluation as before, done once at a longer interval.
The uncomfortable part isn't that a machine can evaluate. It's how little of what I was proud of turned out to be more than evaluation.
The bet
The shift I was watching was that teams stopped being all human.
It happened one task at a time. An agent takes over something a junior used to do, then something a senior used to do, and after a while a meaningful share of the actual work in a company is executed by something that isn't a person.
Every company I've worked inside carried a body of practice it built over years. How we handle this kind of customer. What we never promise. Which corners we cut when the deadline is real and which ones we don't. Very little of it is written down anywhere useful, and the parts that are written down live in a document nobody reads. It exists in the people who've been there long enough.
If agents are going to do the work, they have to follow that practice. The model's own code, its instructions, the skills you give it, none of that carries it. Those cover how to perform a task, which is a different thing from how this particular company decides.
The matching is automated; the case material it matches against is where the work now sits. A model arrives carrying the internet's precedent and none of yours, so it decides like the average of everyone. Call it the decision layer: the part that has to be codified before an agent decides the way this company does. I hold the name loosely, the way I hold any name I coin.
The mental model I kept coming back to was legal practice. A lawyer facing a question doesn't reason it out from first principles. They look for comparable prior cases. If enough of them line up, the answer is close to settled and the work is applying it. If the situation is genuinely novel, precedent runs out.
That mapped cleanly onto what agents were getting wrong. Where there's a deep body of prior cases, an agent can decide well and I'd rather it did, because it's faster and more consistent than a tired person on a Friday. Where precedent is thin or absent, letting it decide is reckless, and the danger is that it doesn't sound reckless. A model with no comparable cases behind it produces an answer with exactly the same confidence as one with a thousand. The confidence is free.
So the product became a guardrail. Agents act where the precedent supports it, and escalate where it doesn't.
Some of why I cared about this is not strategic at all. For years I was the person every decision routed through, and the queue behind me never got shorter. I was the bottleneck, and I was also the reason the bottleneck existed. What I couldn't see at the time is that two different things were stuck in that queue. Some of it was my evaluation, which the organisation genuinely couldn't get elsewhere yet. The rest was my accountability, which it could not get elsewhere at all. Nobody separated them, so everything waited on me.
That's the bet: the genuinely codifiable part of deciding can be lifted out of individual people, so it stops being trapped in whoever happens to have been there longest.
The line the product made me draw
To make any of it work, I had to draw that line explicitly. There's no version of the product that avoids it.
It took me longer than it should have to notice that the line was the answer to the question I'd been arguing about publicly. I had built the boundary into a product before I understood I was building the argument.
On the near side of the line, decisions are evaluation. That side is codeable, and I want it codeable, because that's where the bottleneck was.
The far side works differently. When precedent runs out, nobody can tell you the answer is right, because the thing that would establish it doesn't exist yet. Someone has to commit anyway and then carry what happens.
That's not judgment at all.
The other half
So what is it.
Start with commitment while the evidence is still thin, and stake: a wrong call lands on you. Then add the part that does the work, which is that the cost is actually borne rather than risked in theory.
That's conviction, and it's the thing I mean when I say it isn't codeable. Not because the reasoning is too subtle for a machine. Because there is no mechanism by which a model bears a consequence.
The counterexamples come quickly. Models do get things imposed on them. They're retrained on outcomes and deprecated when they underperform. And you can construct stakes contractually, by putting capital behind an agent and insuring its decisions.
All of that is a cost applied from outside, which is a different thing from bearing one. The question is whether there's anyone there to lose something. A deprecated model doesn't experience the deprecation, so there is no continuous party that was better off before and worse off after. The contractual version has the same problem one step back: the capital behind the agent belongs to someone, and that someone is who actually pays. Follow any of these arrangements to the end and a person is standing there.
The word gets used loosely, so it's worth being blunt about what it excludes. Believing something before it was popular doesn't count, if believing it cost you nothing. I've held plenty of those, and they cost nothing to produce, which is also why a model is good at it on demand.
Conviction is not a capability you have. It's a price you pay.
There's a harder objection. I've written before that layers built to cover a model's temporary weakness get eaten by the next release. Precedent accumulates. Models keep getting better at generalising from thinner case material. So the line I described should be sliding: what needed a human last year is settled this year, and conviction is just the next layer waiting to go under.
The first half of that is right, and I'm building on the assumption that it is. The line has moved twice in the time I've been working on this, both times in the same direction.
But two boundaries are being confused. One is what can be evaluated, which is a capability question, and it moves fast. The other is who carries the cost when the evaluation is wrong. Getting better at generalising from thin precedent changes nothing about who is standing there when the answer turns out to be wrong.
And the movement runs the other way from what you'd hope. The commitments that get absorbed first are the cheap ones, the calls where enough people have already done the thing that precedent exists. What's left over is the expensive end. The territory doesn't shrink so much as it concentrates, and the average weight of what's still in it goes up.
The consolation on offer, when people say judgment is what keeps you safe, is that you already have the thing that saves you and can carry on. What survives instead is an exposure you keep accepting. It feels like being on the hook.
A test you can run
Take any skill or instinct you believe protects you from being automated. Ask what it costs you to be wrong when you use it. If the answer is nothing, you just evaluate again, then it is judgment, and the model is coming for it. If being wrong costs you something you can't get back, it is conviction, and it is the part that's yours.
One correction, because on its own the question gives a flattering answer. Cost alone isn't enough. Think of any role where a single call carries a heavy consequence and the call itself rests on thousands of comparable prior ones. Expensive to get wrong, and still not conviction: the consequence was assigned by the org chart, while the decision was settled by precedent. Both conditions have to hold. Nothing could have settled it in advance, and the bill has to be yours.
Run it honestly on your own calls and it's less flattering than it sounds. Mine came back as judgment. The read on a design, the sense for which feature would land, the instinct for where a plan breaks. Good work, all of it, and all of it evaluation. When I'm wrong about which of three designs is better, I look at the fourth. Nothing is spent.
The architecture calls are the ones that taught me the distinction, because they split down the middle. Choosing between two known patterns is ranking, and it's cheap to be wrong about at the whiteboard. Committing the company to one of them, in a market nobody had built for yet, was a different act with the same words attached, and I lived inside that one for a year.
The things that came back as conviction were fewer and less comfortable. Choosing a market before anyone could show me it was there. Staying with a problem well past the point where staying was defensible. Being wrong about those costs time I don't get back, money, and the credibility of having said it out loud.
Those are the ones nobody can take, and they're also the ones I have the least appetite for on any given morning.
That leaves the part I set aside earlier. Framing a problem, or generating the option nobody had, is not ranking, and it isn't a third capability sitting off to one side either. It's what the far side of the line is made of. Reasoning your way to an answer nobody can confirm puts you at the moment where someone has to commit without proof.
The practical use is where you put your next year. Getting sharper at evaluating is the obvious plan, and it was mine for a long time, but it puts the year into the half that's automating fastest. The other half gets better only by taking positions where being wrong actually costs you something.
That's harder to arrange than it sounds, and it's why I'd stop short of calling this career advice. Conviction goes with the bill, and the bill isn't always with the person doing the work. I spent 15 years in roles where being wrong cost my employer more than it cost me, and I'm not sure I understood the difference until I was the one paying.
What's left
I'm still building the codeable half, and I think that's the right thing to build. The bottleneck was real. It was me, for years, and there was nothing noble about it.
What changed is that I stopped thinking of it as taking something from people. Evaluation was never the part that was ours in any deep sense. It was the part we happened to be the only ones able to do, which is a different claim, and one that had a clock on it.
The line I had to draw in the product turned out to be a line about work in general. What's left is the calls where nothing settles it and someone has to commit and carry it.
$ ./essays --more
$ ./subscribe
Stay in the loop
Occasional updates when I ship something new or write a note worth reading.