DeePicks
Insights / Data quality

How DeePicks handles missing data

Missing is a state of knowledge, not a football result. DeePicks keeps absent inputs, unavailable predictions, and observed zeros separate so the interface does not manufacture certainty.

Missing does not mean zero

Zero is an observed or modelled numerical value. Missing means the required value was not supplied, could not be validated, or is not available for that market. Replacing missing data with zero would assert something that the source did not establish.

Hypothetical teaching example: a recorded first-half corner total of zero means the data says no first-half corners occurred. A blank first-half corner field means the total is unknown. Grading an Over 0.5 selection as lost from the blank field would be a false conclusion, even though zero would produce that settlement.

Explicit availability in the prediction contract

DeePicks validates each supported market as AVAILABLE or UNAVAILABLE. An unavailable market carries one of a limited set of reasons: model not available, insufficient history, missing input data, source data unavailable, or model disabled. The unavailable object cannot also contain probability data.

Available markets must contain the fields required by their contract. For example, first-half goals requires both supported total lines, while cards requires all four supported match-total lines. Other markets allow only their declared structures. This prevents an incomplete object from quietly appearing as a complete prediction.

How the public interface responds

Public free tips are selected only from upcoming, READY fixtures and explicitly available supported lines. Started fixtures, final fixtures, future-dated snapshots, missing selections, and probabilities outside the documented public-tip band are excluded. If nothing qualifies, the honest result is an empty state—not a sample pick or a copied percentage from another market.

Paid Today and Parlay views can also include clearly labelled added-league research predictions, but only selections whose market status is AVAILABLE and whose model value is a finite probability. Production fixtures win when the same fixture identifier appears in both sources. Source identity and generation time remain attached to research selections.

Settlement data follows the same principle

Bet Tracker's manual grading requests only the statistics needed by the saved selections. A goals market needs the relevant score, first-half goals needs both halftime scores, corners needs the corner total, cards needs the cards total, and shots on target needs its own total.

When a required statistic is null, the selection becomes NEEDS_REVIEW. It is not silently won or lost. A leg containing any selection still needing review cannot be marked won; a confirmed loss still makes the leg lose. This protects the distinction between “the event did not occur” and “the application does not yet have the evidence to decide.”

AVAILABLE is not a quality certificate

Availability answers a narrow implementation question: is there a structurally valid estimate for this market and line? It does not establish that the estimate is calibrated, independently validated, or profitable. Calibration would require checking many comparable forecasts against later outcomes using a declared method. Profitability would additionally depend on real prices, timing, settlement, and costs.

A model can produce an AVAILABLE probability and still be wrong on the next match. It can also pass schema validation while remaining unproven over a meaningful sample. DeePicks therefore avoids inventing performance figures, combined probabilities, bookmaker odds, or returns from availability alone.

Practical checks for readers

Keep the market, selection, exact line, fixture, kickoff, source label, and snapshot time together. Treat an unavailable value as no published estimate. Treat a zero only as zero when the source actually supplies it. If a bookmaker wager is involved, consult that operator's current settlement rules because its data provider, correction window, void policy, and definition of an incident may differ from the application's display.

This approach cannot remove uncertainty or guarantee a correct prediction. It can make uncertainty more visible and prevent a missing value from being converted into a confident but unsupported answer.