How much uncertainty sits behind the next decision?

Decision Risk Score makes the weak points in a measurement-led decision visible. We assess five factors, document the evidence for each rating, and connect the findings to a bounded fix.

This is Calyxra’s working diagnostic rubric. Version 1.0 has not been validated as a predictor of commercial loss. A score of 60 means more measurement concern than 20 under this rubric; it does not mean a 60% probability of failure or 60% wasted spend.

Unit of assessmentOne decision.
One evidence window.
One accountable owner.
01 / 30% weight

Consistency

Do equivalent records tell the same story?

Matched eligible order and event records, with population, currency, timezone and refund treatment aligned. Report failed checks and their commercial materiality.

  1. 0Checks pass within the agreed tolerance.
  2. 1Isolated exceptions; no decision changes.
  3. 2Recurring exceptions affect a decision input.
  4. 3Material unexplained mismatches change the recommendation.
  5. 4A reproduced collection or transformation failure invalidates the input.

Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.

02 / 25% weight

Attribution gap

What remains unexplained after the definitions are aligned?

A reconciliation bridge with eligibility, attribution windows, consent and modelled credit explicitly separated. Use the unexplained residual on a comparable population, never the sum of overlapping platform claims.

  1. 0Residual is explained or inside the agreed tolerance.
  2. 1Small residual; sensitivity check leaves the decision unchanged.
  3. 2Residual is material but the decision holds within tested bounds.
  4. 3Plausible allocation of the residual reverses the recommendation.
  5. 4The decision relies on overlapping or irreconcilable credit as company revenue.

Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.

03 / 15% weight

Data freshness

Is the evidence current enough for this decision?

Source event time, last completed refresh, late-arriving records and required refresh cadence. Divide the age of the last complete data by the pre-agreed maximum acceptable age.

  1. 0Age is at most 1× the allowed age.
  2. 1More than 1× and at most 2×.
  3. 2More than 2× and at most 3×.
  4. 3More than 3× and at most 5×.
  5. 4More than 5×, or a confirmed refresh failure.

Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.

04 / 15% weight

Channel volatility

Does the decision survive the movement in its inputs?

Daily spend, eligible conversion value and campaign or release changes over an agreed window, typically 28 complete days. Test the recommendation under documented channel ranges and separate trading changes from collection changes.

  1. 0The recommendation holds across the tested range.
  2. 1Short-lived variation; recommendation unchanged.
  3. 2Repeated variation needs narrower operating guardrails.
  4. 3A plausible channel range reverses the recommendation.
  5. 4An observed regime change makes the comparison window unusable.

Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.

05 / 15% weight

Confidence interval

Could estimation uncertainty change the action?

A named estimator, sample unit, method, observation window, assumptions and interval. For a complete ledger reconciliation there may be no sampling interval; report this as not applicable and use the other factors separately.

  1. 0A supported interval is within the pre-agreed precision target and clear of the decision threshold.
  2. 1The interval misses the precision target but stays clear of the threshold.
  3. 2The interval touches or crosses the action threshold.
  4. 3The interval includes both a material gain and a material loss.
  5. 4A computed interval is unusable because its documented assumptions fail.

Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.

Every point has a source.

Rate each factor from 0 to 4 using the anchors above. Multiply the rating by its weight, divide by 4, then add the five contributions and round once at the end.

Score = round(Σ weight × rating ÷ 4)

Weights total 100. They are declared judgement weights, not fitted statistical coefficients. We freeze the rubric, thresholds, materiality and data window before scoring so a later result can be reproduced.

If a factor is unknown or not applicable, we publish the factor findings and coverage without an overall score. We do not silently remove a factor or change its weight.

Constructed example / all five factors assessed
53/ 100
Worked calculation
FactorRatingPoints
Consistency2 / 415
Attribution gap3 / 418.75
Data freshness1 / 43.75
Channel volatility2 / 47.5
Confidence interval2 / 47.5
Unrounded total52.5

Priority: investigate before committing to the affected decision.

0–24

Proceed with checks

Document the remaining exceptions and the monitoring owner. The score is not an approval to spend.

25–49

Narrow the decision

Set operating limits, resolve the highest-contributing factor and collect the missing evidence.

50–100

Investigate first

Prioritise the incident before a material commitment that relies on the affected numbers.

A reproduced defect that invalidates the controlling metric overrides the band. A low total never clears a broken input for use.

A score starts the work. Verification closes it.

Each review produces the decision record, source inventory, reconciliation bridge, factor ratings with evidence, and an agreed correction or escalation path.

Discuss your decision

After a correction

Re-run the affected checks on comparable data. Record the code or configuration changed, the approver, rollback path, exceptions, and observation window. Keep the original score and definition alongside the re-assessment.

About statistical confidence

A confidence interval belongs to a defined estimator under stated assumptions. A budget range or a dashboard screenshot cannot produce one. When the required source data is absent, the interval is “not estimated”.

Our statistical terminology follows NIST’s explanation of confidence intervals. NIST does not validate or endorse Calyxra’s scoring rubric.

Keep the claim proportionate

Tracking integrity, agreement between reports, and incrementality answer different questions. A successful reconciliation cannot establish the causal revenue effect of advertising.