Skip to the content.

SF-009 · Exposure counted as usage

Status: stable | Family: C · Wrong Evidence | Confidence: high | as of 2026-09-19

One line. Content is present in every request, so a search for it returns a large number — which is read as usage.

Symptom

You measure adoption by counting occurrences of an identifier in a corpus (logs, transcripts, prompts, rendered pages). The count is large. The conclusion is “heavily used”.

Then you look at one instance and find that the corpus contains the entire inventory on every entry, because a roster, a catalogue, a capability list or a system prompt is emitted with each request.

occurrences of item X in logs: 268
actual invocations of item X:    1

The factor is not noise — it is the fan-out of whatever ambient container repeats the inventory. 268 ≈ (number of requests) + (number of genuine uses).

Why it is silent

A signal that can prove the claim is missing, and the inflation is structural rather than statistical: it comes from the shape of the data, so no amount of averaging removes it. The metric is precise, reproducible and wrong.

This is the same family as SF-008, with a different mechanism. SF-008 uses an artifact from the wrong lifecycle stage; SF-009 uses text from the wrong role — an inventory listing looks identical to a use in a substring search, because both are just the identifier appearing.

A further trap: the measurement is usually taken with a tool that cannot distinguish roles, and the correction requires knowing which site emitted the text. Counting is easy; attribution is not — and the count without attribution is the false metric.

Minimal reproduction

# every request emits the full roster (a system prompt, a catalogue, a menu)
def build_prompt(user_msg, roster):
    return f"Available items: {', '.join(roster)}\n\nUser: {user_msg}"

# audit, later: how often was item X used?
log = read_logs()
hits = sum(1 for entry in log if "item-X" in entry.text)   # ← counts listings
print(f"item-X occurrences: {hits}")

Observed: occurrences: 268 — Expected after correction: invocations: 1 (268 occurrences were roster listings)

Self-check

  1. Inspect one instance of a hit, not the aggregate. Read the surrounding text and identify which site produced it.
  2. Classify the sites into listing (the item is named because it is available) and firing (the item is named because it was chosen).
  3. Count only the firing sites. A pattern that marks the firing site (a tool-call marker, an event id, a distinct log template) is the only usable counter.

Rule of thumb: if a metric’s value is close to your request count, it is measuring exposure, not usage.

Fix

Count events at the firing site only, and make the exclusion explicit in the query.

FIRING   = re.compile(r"\[Tool\] tool '(?P<name>[^']+)'")   # fires
LISTING  = re.compile(r"Available items:")                  # ambient inventory

invocations = [m["name"] for line in log
               if not LISTING.search(line)
               for m in [FIRING.search(line)] if m]

usage = collections.Counter(invocations)
print(f"invocations: {sum(usage.values())} ({len(usage)} distinct)")

Operational notes that make this stick:

Negative control Expected result
Add a request that does not invoke the item usage count unchanged (exposure rises, usage does not)
Invoke the item once usage count +1
Remove the ambient roster both counts converge — proving the gap was the inflation

中文要点