How much uncertainty sits behind the next decision?
Decision Risk Score makes the weak points in a measurement-led decision visible. We assess five factors, document the evidence for each rating, and connect the findings to a bounded fix.
This is Calyxra’s working diagnostic rubric. Version 1.0 has not been validated as a predictor of commercial loss. A score of 60 means more measurement concern than 20 under this rubric; it does not mean a 60% probability of failure or 60% wasted spend.
Unit of assessmentOne decision. One evidence window. One accountable owner.
01 / 30% weight
Consistency
Do equivalent records tell the same story?
Matched eligible order and event records, with population, currency, timezone and refund treatment aligned. Report failed checks and their commercial materiality.
Risk rating / 0–4
0Checks pass within the agreed tolerance.
1Isolated exceptions; no decision changes.
2Recurring exceptions affect a decision input.
3Material unexplained mismatches change the recommendation.
4A reproduced collection or transformation failure invalidates the input.
Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.
02 / 25% weight
Attribution gap
What remains unexplained after the definitions are aligned?
A reconciliation bridge with eligibility, attribution windows, consent and modelled credit explicitly separated. Use the unexplained residual on a comparable population, never the sum of overlapping platform claims.
Risk rating / 0–4
0Residual is explained or inside the agreed tolerance.
1Small residual; sensitivity check leaves the decision unchanged.
2Residual is material but the decision holds within tested bounds.
3Plausible allocation of the residual reverses the recommendation.
4The decision relies on overlapping or irreconcilable credit as company revenue.
Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.
03 / 15% weight
Data freshness
Is the evidence current enough for this decision?
Source event time, last completed refresh, late-arriving records and required refresh cadence. Divide the age of the last complete data by the pre-agreed maximum acceptable age.
Risk rating / 0–4
0Age is at most 1× the allowed age.
1More than 1× and at most 2×.
2More than 2× and at most 3×.
3More than 3× and at most 5×.
4More than 5×, or a confirmed refresh failure.
Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.
04 / 15% weight
Channel volatility
Does the decision survive the movement in its inputs?
Daily spend, eligible conversion value and campaign or release changes over an agreed window, typically 28 complete days. Test the recommendation under documented channel ranges and separate trading changes from collection changes.
Risk rating / 0–4
0The recommendation holds across the tested range.
3A plausible channel range reverses the recommendation.
4An observed regime change makes the comparison window unusable.
Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.
05 / 15% weight
Confidence interval
Could estimation uncertainty change the action?
A named estimator, sample unit, method, observation window, assumptions and interval. For a complete ledger reconciliation there may be no sampling interval; report this as not applicable and use the other factors separately.
Risk rating / 0–4
0A supported interval is within the pre-agreed precision target and clear of the decision threshold.
1The interval misses the precision target but stays clear of the threshold.
2The interval touches or crosses the action threshold.
3The interval includes both a material gain and a material loss.
4A computed interval is unusable because its documented assumptions fail.
Evidence missing? Mark “unknown”. Absence of evidence is never rated 0.
The calculation
Every point has a source.
Rate each factor from 0 to 4 using the anchors above. Multiply the rating by its weight, divide by 4, then add the five contributions and round once at the end.
Score = round(Σ weight × rating ÷ 4)
Weights total 100. They are declared judgement weights, not fitted statistical coefficients. We freeze the rubric, thresholds, materiality and data window before scoring so a later result can be reproduced.
If a factor is unknown or not applicable, we publish the factor findings and coverage without an overall score. We do not silently remove a factor or change its weight.
Constructed example / all five factors assessed
53/ 100
Worked calculation
Factor
Rating
Points
Consistency
2 / 4
15
Attribution gap
3 / 4
18.75
Data freshness
1 / 4
3.75
Channel volatility
2 / 4
7.5
Confidence interval
2 / 4
7.5
Unrounded total
52.5
Priority: investigate before committing to the affected decision.
From score to action
0–24
Proceed with checks
Document the remaining exceptions and the monitoring owner. The score is not an approval to spend.
25–49
Narrow the decision
Set operating limits, resolve the highest-contributing factor and collect the missing evidence.
50–100
Investigate first
Prioritise the incident before a material commitment that relies on the affected numbers.
A reproduced defect that invalidates the controlling metric overrides the band. A low total never clears a broken input for use.
The handoff
A score starts the work. Verification closes it.
Each review produces the decision record, source inventory, reconciliation bridge, factor ratings with evidence, and an agreed correction or escalation path.
Re-run the affected checks on comparable data. Record the code or configuration changed, the approver, rollback path, exceptions, and observation window. Keep the original score and definition alongside the re-assessment.
About statistical confidence
A confidence interval belongs to a defined estimator under stated assumptions. A budget range or a dashboard screenshot cannot produce one. When the required source data is absent, the interval is “not estimated”.
Tracking integrity, agreement between reports, and incrementality answer different questions. A successful reconciliation cannot establish the causal revenue effect of advertising.