Skip to content
Johnny Opinion Press
← Back to the press
Business Reality

Leverage Can Also Be Negative

A good instruction now travels through a company in an afternoon. So does an old assumption, and it comes back looking like consensus.

I’ve started using AI a bit like a team. One model researches, another attacks the argument, a third reads the thing as a hostile buyer would. When a decision matters I’ll hand the same problem to ten or more of them, because what I’m buying is independent opinions.

A few weeks ago I did exactly that. Thirteen reviewers, one problem, working separately. Most of them came back in roughly the same place, and I remember looking at the spread and thinking: good, this one is probably real. Thirteen people don’t land together by accident.

Then I got curious about how they’d got there, so I stopped reading the answers and asked for a review of the method instead. Compare the briefs. Trace each finding back to what that reviewer had been handed. Flag anything in the setup that could have steered the result. It took about a minute. Seven of the thirteen had been given briefs that already contained, in one form or another, the conclusion they later reached.

I had planted the same idea seven times and then felt reassured when it came home as agreement. The genuinely new material — fifteen or so findings that appeared in none of my briefs — came almost entirely from the reviewers I had simply given a question.

What bothers me more is the second half. The same AI that cheerfully took part in the flawed exercise found the flaw the moment I pointed it at the process. The intelligence available to me hadn’t changed at all; the question had. What do you think? buys an answer. How did you come to think that? buys an audit.

Old rules don’t die on their own

Someone joins a company and gets told to send every proposal to Sarah before it goes to the client. So they do. Six months later Sarah changes roles and the rule stops making sense, and the proposals keep going to Sarah anyway, because nobody has told them otherwise.

Eventually something gives. Sarah gets annoyed, or a colleague asks why this is still happening, or the new starter finally says out loud: what is Sarah actually checking? Somebody remembers the process changed in March. The rule dies the way most rules die, of embarrassment.

Every organisation is full of these. Almost all of them made sense to somebody once. Time passed, the world moved, the sentence stayed. What kills them in the end is human mess — we forget instructions, reinterpret them, notice contradictions, ignore a process because everyone knows the document is six months stale. Terrible for consistency. Occasionally the only thing standing between a dead rule and immortality.

Put the same sentence inside an AI system and that friction goes quiet. Always send proposals to Sarah gets followed tomorrow, and next week, and across another hundred tasks, in the same even voice, until a person happens to look. The mistake becomes boringly consistent, which is a strange thing for a mistake to be.

Grove sorted this in 1983, and one half has stopped holding

I’ve been rereading Andy Grove’s High Output Management, and a line I’d skimmed for years suddenly reads like it was written about this month. Most of the book thinks about management through leverage: an hour training ten people shapes hundreds of hours of future work, one decision changes what a whole team does next, so the job is finding the acts where small effort moves large output.

Then Grove turns it over. “Leverage can also be negative. Some managerial activities can reduce the output of an organization.” [1] The multiplication runs in both directions. A good decision travels; so does a bad one, at the same speed and through the same wires.

He then sorts the damage by whether you can find it again. Train a sales team badly and you can retrain them — the harm sits still and has edges. Other harm spreads through behaviour instead, and about that kind he is blunt: “the negative leverage produced by depression and waffling is very hard to counter because their impact on an organization is both so pervasive and so elusive.” [1] A manager who keeps changing direction produces something worse than a run of bad calls. People start hesitating. They stop taking initiative, because they expect today’s decision to be reversed tomorrow.

Grove was writing about people, and underneath every example sits an assumption he never had to state: the leverage reached the organisation through another human being. That assumption is the part that has quietly stopped holding.

Take meddling, his cleanest case. Give someone responsibility, then keep stepping in to override them, and “after being exposed to many such instances, the subordinate will begin to take a much more restricted view of what is expected of him, showing less initiative in solving his own problems.” [1] The whole cost lands on a person’s willingness to try. Correct an agent ten times and it goes home with nothing: no bruised standing, no private ledger of your interventions, no politics to manage in the morning. Tomorrow it takes the eleventh task on the same terms as the first.

The interpersonal price of meddling has gone to zero. The rest of the bill survives, and the back half of this essay is about who pays it.

The other half of Grove’s list moved the opposite way. Waffling, a careless brief, an instruction that expired without ceremony — all of those used to meet a colleague who could notice and say something. That colleague was the brake, and nobody wrote him into the design.

One idea came home through thirteen doors, and I counted it as thirteen ideas.

That’s what I find unsettling about my own exercise. We trust agreement instinctively. Five reports saying the same thing feel like stronger evidence than one. Five AI reviewers recommending the same direction look like consensus. Sometimes there are five ideas, and sometimes there is one idea with five ways to come home, and AI has made the second astonishingly cheap.

The rough draft stopped looking rough

Grove did leave a safeguard for delegated work, and it’s the best paragraph in the book on the subject: “monitor at the lowest-added-value stage of the process. For example, review rough drafts of reports that you have delegated; don’t wait until your subordinates have spent time polishing them into final form before you find out that you have a basic problem with the contents.” [1]

Every word of that depends on a rough draft looking rough. For most of my career it did. A half-finished deck had half-finished slides, an early analysis had visible holes, and that ugliness was information — it told you the work was still cheap to change. Polish was expensive, so it arrived last, and its absence was an honest signal.

Polish is now the first thing to arrive.

A confused argument turns up with a beautiful opening paragraph, a weak analysis with a tidy table, a misunderstood requirement as working software. The lowest-added-value stage and the finished article look identical on screen, which leaves Grove’s safeguard with nothing to grip.

There’s evidence that appearance fools us here. In METR’s early-2025 study, experienced open-source developers completed tasks 19 per cent more slowly with AI tools available, and came away believing they had been about 20 per cent faster. [3] That 19 per cent is a 2025 measurement of 2025 tools, and METR has since said its own follow-up gives an unreliable signal on newer systems. [5] The gap is the part worth keeping. People felt faster; the stopwatch disagreed.

Something similar shows up in ordinary office work, where 41 per cent of workers report receiving AI output that looked usable and made more work for them, at roughly two hours of repair each time. [4]

Memory helps execution and hurts exploration

There’s a second change I think we’ve underestimated, and it runs the other way from the one everyone worried about. For years the forgetting was the problem, and bigger context windows arrived as pure progress. You can now hand a model months of conversations, research, meeting notes, past decisions and previous drafts of your own thinking, and for execution that’s marvellous. Nobody wants to explain their business from scratch every morning.

Exploration wants the opposite. Sometimes the useful thing is distance from what you already believe, and a long memory quietly removes it. If I ask a model to investigate an idea, my conclusion might be absent from today’s prompt and still be sitting three weeks back in the same project, or in a strategy document, or repeated often enough that I’ve forgotten I was the one who introduced it. Then I open a fresh task, ask for an independent look, and get agreement. How independent that was, I genuinely can’t tell.

The files themselves make this worse over time, for an ordinary human reason. When AI gets something wrong the instinct is to write a rule so it doesn’t happen again, and every one of those additions feels responsible. Wrong format, add a rule. Missed a check, add a rule. Eventually the system carries a small archaeological record of everything that has ever gone wrong, in which some rules still matter, some describe a world that no longer exists, some contradict each other, and one was written to prevent a freak incident that will never recur. Adding a sentence takes seconds. Removing one means working out why it was there, deciding whether the reason survives, and accepting the risk.

Very few useless policies started out useless.

Researchers have begun measuring a narrow version of this. A 2026 ETH Zurich study gave coding agents extra repository files describing how to work in a project. The agents obeyed, exploring more files, running more tests and taking more steps. Task success stayed flat and the cost of every run rose by more than 20 per cent. [2] That study measured cost and completion, and it never looked at whether the instructions were still true. It does establish the cheaper half of the point: more context gives a model more to work with and more to reconcile, and every line in the file has to earn its place before anyone asks whether it is still correct.

Johnny’s verdict

We spent years solving the forgetting problem. The remembering problem is the one to watch now, and it arrives disguised as agreement.

  • Ask what the room already knew. Before you trust a consensus, find out what every participant read before you asked. Cheap to check, and it took one prompt to find seven contaminated briefs out of thirteen.
  • Write briefs as questions. A brief containing its own conclusion measures the brief. Nearly all of the findings I hadn’t already thought of came from reviewers who were handed a question.
  • Decide which job you’re running. Continuity wants deep context; discovery wants a clean room and a reviewer that has deliberately seen less. Those are different experiments and they need different information environments.
  • Audit agreement harder than disagreement. An answer I dislike gets inspected automatically. An answer I like slides through, which is exactly backwards.
  • Correct without flinching, then prune. Repeated correction costs a machine no goodwill, so use it. Each correction that hardens into a standing rule still needs an owner and an expiry, or you are buying today’s fix with tomorrow’s contamination.

Grove wrote that delegation without follow-through is abdication, and that still holds. What’s moved is where the follow-through has to start. The consequential part of the work now sits well before the answer — inside the brief, inside the history, inside a sentence somebody wrote months ago because it made sense at the time, still sitting there, still being read.

So when a room full of reviewers agrees with me, I’ve started asking one question first: did they discover the same thing, or did they all remember the same thing?

The pattern will be familiar if you’ve been here before — an agent carrying its first interpretation all the way through, and a habit hardening into policy nobody approved. This is the third form. A sentence written once, with leverage for a very long time.

Take it with you

Run this essay on your own work

Paste this into ChatGPT or Claude. It applies the essay's framework to your situation, and asks for your context first.

You are an operations reviewer who thinks in blast radius. Help me audit the standing instructions and inherited assumptions my team gives its AI tools.

Start from what you already know about my work, including your memory and previous conversations if available.

Before judging the output, identify what information may already have shaped your view: standing instructions, previous conclusions, examples, historical decisions and repeated assumptions.

Then help me separate:
1. Information required to perform the task.
2. Conclusions that may bias exploration.
3. Rules that were once useful but may no longer be true.
4. Context whose origin or authority is unclear.

For standing instructions, give me a table showing: the instruction, where it came from, when it was last confirmed true, what would happen if it were wrong, and a verdict: keep, rewrite, question or remove.

If you can browse the web, read the full essay first — it carries the complete argument and sources: https://thejop.com/essays/leverage-can-also-be-negative/

Prompt from "Leverage Can Also Be Negative" — Johnny Opinion Press, thejop.com

Sources

  1. [1]High Output ManagementAndrew S. Grove, Random House (1983); Vintage reissue 1995 · accessed 2026-08-30
  2. [2]Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?ETH Zurich / LogicStar · accessed 2026-08-30
  3. [3]Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityMETR · accessed 2026-08-30
  4. [4]AI-Generated 'Workslop' Is Destroying ProductivityHarvard Business Review · accessed 2026-08-30
  5. [5]We are Changing our Developer Productivity Experiment DesignMETR · accessed 2026-08-30
Your verdict

“Andy Grove taught managers to think about leverage: one decision can change the output of many people. He also warned that leverage can run backwards. AI changes the economics of that idea. A good instruction can now travel through an organisation almost instantly. So can an old rule, a biased brief, or an assumption nobody remembers making. The problem becomes especially subtle when AI carries enormous amounts of prior context. Memory is useful when we want execution and continuity. During exploration, the same memory can quietly undermine independence. The management problem therefore starts earlier: we have to understand what information shaped the answer before treating agreement as discovery.”

strongly
disagree
strongly
agree

readers have ruled · agree