Skip to content

A refusal with no scope becomes a law: three ways feedback breaks an AI product

Three kinds of feedback reach an AI product, and most systems misread all three. A no without a scope becomes a standing rule. A yes from someone who never checked becomes truth. Silence — the draft started and never sent — is never counted. Store the scope, split the ratings, count the unsent.

evals · trust · process

Three kinds of feedback reach an AI product: a refusal, an approval, and silence. In our experience most systems misread all three, and the three misreadings have one root. The signal is stored as a value when it should be stored as a value plus the context it arrived in.

A refusal comes with a reason, and the reason is kept without the boundary it was said inside. An approval tends to come from the user least likely to have checked the answer. Silence never comes at all, so nothing counts it.

Three fixes. All of them are storage decisions, not model decisions: store the scope with the refusal, split the ratings by what the user actually did, count the work that was started and never sent.

1. A refusal with no scope becomes a law

This one is ours.

In an internal tool of ours, a person reviews an agent's work and either approves it or turns it down with a reason. The handler passed the agent the last 40 messages of that review chat. Four role prompts told it to treat every refusal reason as a standing rule. So one sentence said about one job on 3 September governed everything the system produced afterwards.

Nothing looked broken. Our own note on it reads: “the output stayed plausible, it was simply narrower than anyone had asked for.” That is the expensive part. A refusal that leaks turns into a quality problem you cannot see, because the output it produces is still fine — it is just answering a smaller question than the one you asked.

The fix was one field.

# before — a reason with no boundary
decision:
  verdict: no
  reason:  "too long, cut it"

# after — the same reason, with the boundary it was said inside
decision:
  verdict: no
  reason:  "too long, cut it"
  scope:
    kind: post        # post | campaign | all
    ref:  "8f3c..."   # what exactly it applies to

The reading rule matters as much as the field. A stored decision is loaded into the context of a new job only if that job falls inside the decision's scope. all is the only global scope, and it has to be picked on purpose. A person saying “make it shorter” about one piece is not writing a house style, and the system should not be able to promote them to it by accident.

Our view: if your agent reads a rolling window of chat history, you do not have a memory of decisions. You have a memory of sentences, and sentences do not carry their own limits.

2. An approval from someone who did not check is not a verified label

A survey of 544 generative-AI users found that trust pulls in two directions at once. Higher trust makes a user more willing to give the system feedback, and it also makes them less likely to verify the output before giving it. The authors call the risk feedback without verification.

Read that as a founder with a thumbs-up button on an AI feature. The people most willing to rate are the ones least likely to have checked what they rated. A learning loop trained on that stream is not learning quality. It is learning what confident users wave through, which includes your errors.

The fix is not a better button. It is refusing to store a rating as one bit.

rating(answer_id, value)                  # one bit, unlabeled
rating(answer_id, value, verified_by)     # what the user did before rating

verified_by ∈ { edited, copied, opened_source, none }

All four are things a product usually logs already. You are not asking the user for anything new — you are joining the rating to the behavior that surrounded it. Then split the report:

select verified_by,
       count(*)   as ratings,
       avg(value) as avg_score
from ratings
where created_at > now() - interval '30 days'
group by verified_by
order by ratings desc;

If none carries most of your volume and the highest average score, your quality metric is measuring trust, not quality. Our view: an edit is the strongest label on that list, because the user paid for it with work. A copy is next. A bare rating is the weakest one you have, and it is the one a dashboard usually reports.

3. Silence is the cheapest signal you have, and nobody counts it

The third kind never arrives in the feedback table at all.

A user who does not like an AI feature usually does not complain. They open it, start, and do not finish. The rejection is an unfinished draft — the feature was used, and the output was not good enough to send. We take this as a hypothesis worth measuring rather than a finding: it comes from one short talk with a very small audience, the speaker is not ours to name, so we paraphrase it and test it instead of citing it.

Testing it is cheap, and a founder can pull the number tonight.

-- of the drafts started this month, how many were sent?
select count(*)                                    as started,
       count(*) filter (where sent_at is not null) as sent,
       round(100.0 * count(*) filter (where sent_at is not null)
             / nullif(count(*), 0), 1)             as pct_sent
from drafts
where created_at >= date_trunc('month', now());

Swap drafts for whatever your feature produces. Generated and never exported. Opened and never saved. Suggested and dismissed without a click. The shape is the same: a unit of work that started and did not reach its exit. We expect that ratio to move before your ratings do, because abandoning something costs the user nothing and complaining costs them a message.

This is also why one of our products is open for free while it is being tested. We expect defects, and the thing that reliably finds them is a live person deciding not to finish.

What we cut, and what would change our mind

We cut a weighted rating — a model-estimated confidence multiplier on each thumbs-up. A weight is a guess about the user. verified_by is an observation of the user, it costs one column, and anyone can audit it.

What would change our mind on the split: if verified and unverified ratings track each other closely in our own data over a few months, the split buys nothing and we drop it and say so here.

What would not change our mind is the scope field. That one is already paid for.

FAQ

Was the problem that the agent had too little history? No. It had 40 messages, which was plenty. The missing thing was not volume, it was a boundary on each stored decision. More history would have made the leak larger.

Which scopes are worth storing? We store a decision against the piece it was said about, the campaign it belongs to, or everything. The exact list matters less than having all be a separate, deliberate choice rather than the default reading of every sentence.

Is a rating from someone who did not verify useless? It is not useless, it is a different measurement. Filed on its own, it is a good proxy for how much your users trust the feature. Mixed into a quality score, it inflates it.

Our product has no drafts. What counts as silence? Anything with a start and an exit. Generated then discarded, opened then abandoned, suggested then ignored. Count the starts, count the exits, watch the ratio weekly.

Related

The short version of all three, as a checklist you can save and run against your own tables, is here: store the scope, split the ratings, count the unsent.

Cases and process at iloblique.com.