Skip to the content.

SF-004 · Sentinel value on failure

Status: stable | Family: A · Vacuous Verification | Confidence: high | as of 2026-09-19

One line. A function returns 0.0 / "" / [] when it fails, so a crash becomes indistinguishable from a bad result.

Symptom

def f1_score(preds, labels) -> float:
    try:
        ...
    except Exception:
        return 0.0            # ← the bug

The evaluation script now reports F1 = 0.0. That number is ambiguous: it is what you would see if the model performed terribly, and also what you would see if the code threw at line 47. The two situations require completely different responses and are indistinguishable in the report.

Common sentinels and their collisions:

Sentinel Collides with
0.0 a genuinely zero metric
"" an empty-field case (see SF-007)
[] “no results found” — a legitimate, meaningful outcome
-1 any legitimate negative count that is not supposed to occur, until it does
None “not computed yet”, “not applicable”, “failed” — three different states

Why it is silent

A signal that can prove the claim is missing: the return channel carries a plausible value instead of a distinguishable state.

The failure is not that the error was lost — the error was laundered into data. Downstream consumers see a valid value, aggregate it, plot it, and reason about it. The error has become a finding.

This is the most expensive of the Family A entries, because it does not merely hide a problem — it manufactures a false observation that other people then act on.

Minimal reproduction

def score(preds, labels):
    try:
        assert len(preds) == len(labels)
        return sum(p == l for p, l in zip(preds, labels)) / len(preds)
    except Exception:
        return 0.0

print(score([1, 0], [1, 0, 1]))    # length mismatch — a bug
# → 0.0   indistinguishable from "the model got everything wrong"

Observed: 0.0 — Expected after fix: an exception, or a structured {ok: false, reason: ...}

Self-check

For each function that returns a value in a try/except:

  1. Enumerate every sentinel it can return, and every legitimate value it can return.
  2. If the two sets intersect, you have this bug.
  3. Grep for except blocks containing return — each is a candidate.

A useful sharper question: “if this function fails, what will the report say, and is that distinguishable from a real result?”

Fix

Raise. Let the failure be a failure.

def f1_score(preds, labels) -> float:
    if len(preds) != len(labels):
        raise ValueError(f"length mismatch: {len(preds)} vs {len(labels)}")
    if not preds:
        raise ValueError("empty input — cannot compute a metric")
    ...

Where the caller genuinely needs a value back, return an explicitly typed outcome rather than an in-band sentinel:

@dataclass
class Outcome:
    ok: bool
    value: float | None
    reason: str | None

and make the aggregation fail on any ok=False.

Note for metric code specifically: a metric function that returns 0.0 on error means every evaluation table in the project silently mixes “the model is bad” with “this code is broken”. Refuse both: raise on bad input, and assert non-empty input.

Negative control Expected result
Pass inputs of mismatched length raises ValueError, run aborts; no 0.0 is ever produced
Pass an empty input list raises; never reports a metric for zero records

中文要点