A rule for every failure made our agents obey less. The fix was scope, not wording
Adding an instruction for every agent failure makes the whole prompt less reliable. Our fix was not a better sentence. Each human decision is now stored with its scope, and each rule gets a lifetime. Our first fix for a related check was worse than the bug.
The short answer. The standard advice for a failing agent is to add a rule for the failure. People who run agents in production report the opposite effect: the more rules a prompt carries, the worse the model follows any one of them. We ran into the same thing in our own system. Rewording the rule doesn't fix it. What fixes it is storing each human decision with the scope it was made for, and giving each rule a lifetime. We also got the first repair wrong, and that part is below too.
The advice, at its strongest
The case for adding rules is a fair one. A failure you saw once will come back. Writing it down is cheap. Examples and constraints narrow what the model can do, and narrowing is usually the goal. A widely read 2026 guide to prompt engineering for agents still recommends this as good practice: more diverse examples and more constraints. That is the advice a buyer's own team is likely to quote.
When the prompt is small and the rules don't overlap, we agree with it.
What production reports
Practitioners who run agents in production call the pattern the curse of instructions. Every rule you add makes the model less likely to follow any single rule. Their write-ups describe agent manuals of 2,000 words and 4,000 tokens, built up one edge case at a time. In those accounts each addition left the agent less reliable than before.
That is someone else's report. Here is ours.
What it looked like in our system
The setup was an internal multi-role agent system. A handler put the last 40 messages of one chat into every request. Four separate role prompts each told the model to treat every reason for a rejection as a permanent rule with no expiry.
every request
context: last 40 messages of one chat
└─ includes one remark, 3 September, about one campaign
role prompts (4 of them):
"treat every rejection reason as a permanent rule"
result
remark about one campaign
-> rule with no scope and no expiry
-> applied to everything produced afterwards
-> overrides an authoritative source on the same questionThat one sentence, written on 3 September about a single campaign, ended up governing everything the system produced after that date.
From the outside, nothing looked broken. The output stayed plausible. It was just narrower than anyone had asked for. Where an authoritative source and the old remark covered the same question, the remark won, and nothing reported the override.
Our view: this is the real cost of piling on instructions. There's no error message. The output gets narrower, and no test catches "narrower than asked". You only see it if you read the output against what was actually requested.
The fix: a decision is stored with its scope
The fix wasn't better wording on the four prompts. It was a change to what a decision is. A decision is now a record, and one of its fields says where it applies.
decision
verdict: no
reason: "..."
scope: post | campaign | project | topic | allAt read time the rule is mechanical:
for each decision:
scope == all -> apply everywhere
scope matches this deliverable -> apply, reason is binding here
otherwise -> ignore; it was about something elseall is the only global scope. A reason given for one campaign binds that campaign and nothing else. It doesn't turn into a global rule over time. If the owner wants a lesson applied everywhere, they have to say so, and that is a separate decision with its own scope.
The second half of the fix is lifetime. Our view is that a rule nobody has been held to for a long time is a candidate for deletion, not for more emphasis. A rule has to keep showing it is still needed. Otherwise it's just the pile described above, getting bigger.
Scope also keeps the context small. When each decision knows where it applies, the context holds only the rules relevant to the current job. You don't have to curate that set by hand. It follows from the data.
Our first fix was worse than the bug
The same week we repaired a check nearby and got it wrong. We think the mistake is common enough to be worth writing up.
The check had three possible outcomes, not two: pass, fail, and could not run. The old version reported "could not run" as a green pass. That green pass hid a gap of fourteen migrations between the repository and the database.
Our patch went the other way and treated "could not run" as a failure. Within twenty minutes it was clearly worse. The check failed on every commit, for a reason no commit could fix. Once a signal is red for everyone all the time, it stops meaning anything, just like the one that was always green. We rolled it back.
outcome before our patch now
pass green green green
fail red red red
could not run green (hid red on every commit routed to whoever can
14-migration (nobody can act) act on it, when they
gap) can act on itFolding the third state into either of the other two destroys it. The fix was to route "could not run" based on who can answer it and when, not on how severe it is. The author of a commit can't fix a database that is missing migrations, so the check shouldn't block them. It should reach the person who can fix it.
Our view is that both bugs are the same mistake. The first threw away a decision's scope. The second threw away an outcome. In each case a distinction that looked like overhead was carrying the meaning.
What would change our mind
We would change our minds if someone showed a long, flat list of rules that keeps its adherence as rules are added, on a real workload over weeks. We'd also reconsider if scoped decisions started drifting in the same way unscoped ones did. So far we have seen neither.
What to check this week
1. count the instructions in your system prompt
2. for each one: who said it, about what, and when
3. search for any line telling the model to generalize from feedback
4. list every check's outcomes; find where "could not run" goesFAQ
Isn't this just "write shorter prompts"? No. In our experience, shorter prompts treat the symptom. The cause is decisions stored without the context they were made in. Record the scope and the prompt stays short without anyone having to cut it by hand.
Won't scoping make the agent forget real lessons?
Only the ones nobody chose to keep everywhere. A lesson that should apply everywhere can be raised to all by the person who owns it, on purpose. What we stopped doing is letting the system make that promotion on its own.
How do I tell that an old rule is running my agent? Look for output that is plausible but narrower than the request. Or look for a place where a documented source says one thing and the agent does another. In our case that disagreement was the only visible symptom.
Should "could not run" count as a failure? Not by default. We tried it and rolled it back within twenty minutes. Send it to whoever can act on it, and don't block the people who can't.
A related case
Kindling, one of our products, faces the same question from the other side: what should a system learn from a person's edits? In Kindling every edit teaches, but a learned change shows up as a visible, optional observation instead of a silent rewrite. The person decides whether it becomes a rule. Read the Kindling case.