Today's correct score breakdown — VIP Gold live
Last updated 5 minutes ago
→
Football Analysis Guide

How to Build a Football Scoreline Probability Matrix

Turn expected home and away goals into a coherent joint distribution for plausible football scorelines, then test whether the resulting probabilities are calibrated and decision-useful.

How to Build a Football Scoreline Probability Matrix

A football scoreline probability matrix assigns a probability to every home-and-away goal combination: 0-0, 1-0, 1-1, 2-1 and so on. It is more than a correct-score tool. Once the matrix is internally coherent, win-draw-win, both teams to score, totals and combination-market probabilities can all be derived from the same joint distribution.

The arithmetic is relatively simple. The difficult part is estimating each team's expected goals, deciding whether the two goal counts can be treated as conditionally independent, and quantifying the uncertainty in those assumptions. This guide builds the standard independent-Poisson matrix first, then examines tail handling, market aggregation, dependence, parameter uncertainty and validation.

Define the probability object before choosing a model

Let H represent home goals and A represent away goals. A scoreline matrix contains the joint probabilities P(H = h, A = a) for non-negative integers h and a. Rows can represent home goals and columns away goals, or vice versa, but the orientation must be documented because reversing it silently reverses home and away wins.

Every cell is a mutually exclusive outcome: a match cannot finish both 1-0 and 1-1. Across an infinite matrix, the cells are collectively exhaustive, so their probabilities sum to 1. A displayed matrix is normally finite, meaning that its visible cells sum to less than 1 unless high scores are grouped into an overflow category.

This distinction matters. The matrix is not a list of separate predictions generated independently. It is one probability distribution. Raising the 1-1 probability while leaving every other cell unchanged makes the total exceed 100%. Probability mass must be removed from other outcomes, or the whole distribution must be re-estimated or renormalised under a stated rule.

Separate model output from interpretation. A cell probability of 11% means that the model assigns that scoreline to roughly 11 of 100 repeated matches played under comparable modelled conditions. It does not mean that the score is close to certain, even if it is the largest cell. Football probability is usually spread across many plausible outcomes, so the modal scoreline often remains unlikely in absolute terms.

Estimate one scoring rate for each team

The basic matrix requires two inputs: an expected home-goal rate, λH, and an expected away-goal rate, λA. These are conditional means, not predictions that either team will score exactly that many goals. A value of λH = 1.60 says that the home side's average goal count would be 1.60 across repeated matches played under the modelled conditions.

A simple strength model can estimate the rates on a logarithmic scale. Conceptually, log(λH) can be built from a competition baseline, home advantage, the home team's attacking strength, the away team's defensive strength and contextual variables. The away equation uses the corresponding away-attack and home-defence terms. A log link keeps rates positive and allows team effects to combine multiplicatively.

The rates should be estimated rather than chosen because a team recently scored two or three goals. Possible inputs include goals, expected goals, shot quality, venue, opposition strength, rest, lineup information and red-card-adjusted match states. The exact variables depend on data quality and the forecasting time. A prematch model must not use information that became available after kickoff.

Team-strength estimates also need shrinkage. A side with a small sample of unusually strong attacking results should normally be pulled toward the competition average. Without regularisation, temporary finishing variance can become an exaggerated λ estimate and produce a matrix that is too confident about high scores.

Recency weighting is another modelling choice. Heavy weighting adapts quickly to tactical or personnel changes but reacts strongly to noise. Slow weighting is more stable but can lag behind genuine changes. The decay rate should be selected through time-ordered validation rather than narrative judgement.

Convert expected goals into marginal goal distributions

The standard starting point is a Poisson distribution. For a team with scoring rate λ, the probability of scoring k goals is P(K = k) = exp(-λ) × λ^k / k!. The model therefore converts one mean into probabilities for zero, one, two and every higher goal count.

For implementation, calculate the zero-goal probability as exp(-λ), then use the recurrence P(K = k + 1) = P(K = k) × λ / (k + 1). The recurrence avoids repeated factorial calculations and makes it straightforward to extend the goal range until the remaining tail is below a chosen tolerance.

The Poisson specification imposes a particular shape: conditional mean equals conditional variance. That can be an acceptable baseline, but it is an assumption rather than a law of football. Tactical uncertainty, uncertain lineups and changing game states can create more variation than a fixed-rate Poisson permits. Other matchups may suppress the tails.

If the marginal vectors are expanded far enough, each should sum approximately to 1. If a finite display is used instead, its missing marginal mass must be retained as a tail or overflow category. A short vector summing below 1 is not necessarily an error; silently treating it as complete is.

Form the joint matrix and retain the omitted tail

Under conditional independence, each cell is the product of the corresponding marginal probabilities: P(H = h, A = a) = P(H = h) × P(A = a). In matrix terms, this is the outer product of the home-goal vector and the away-goal vector.

Independence has a precise meaning here. Once λH, λA and all included covariates are known, learning the realised home-goal count provides no additional information about the away-goal count. This is stronger than saying the teams were estimated separately. Match state can violate the assumption because a goal changes incentives, pressing, risk and substitution decisions.

The illustrative matrix below uses fictional rates of λH = 1.60 and λA = 1.10. It displays only scores from zero to four goals for each team. The home marginal probability of zero through four goals is 97.63%, and the away equivalent is 99.46%. Under independence, the displayed 25-cell probability is their product: approximately 97.10%. The remaining 2.90% covers outcomes in which at least one team scores five or more.

Do not divide the displayed cells by 97.10% merely to force them to sum to 100%. That would treat high-score outcomes as impossible and inflate every visible scoreline. Retain an explicit tail, add a 5+ row and column, or expand the matrix until omitted probability is below a tolerance relevant to the event being priced.

The appropriate tolerance depends on use. A small omitted tail may have little effect on a three-way result probability but can matter materially to high total-goal lines or correct-score groups. Control truncation with reference to the event being priced, not only a universal display convention.

Aggregate cells into market-level probabilities

Once the joint distribution is available, common event probabilities are sums over defined regions of the matrix. A home win is the sum of cells where h > a. Draw probability is the main diagonal where h = a. An away win is the region where h < a. These three regions sum to 1 only when the complete tail is included.

A total-goals event is organised by diagonals running in the opposite direction. Over 2.5 goals includes every cell where h + a is at least 3. Under 2.5 contains cells where the sum is 0, 1 or 2. Under independent Poisson assumptions, total goals follow a Poisson distribution with mean λH + λA. Direct matrix summation remains useful because it also works under non-independent models.

Both teams to score is the region with h > 0 and a > 0. It can be calculated as 1 - P(H = 0) - P(A = 0) + P(0-0). Under independence, this simplifies to [1 - exp(-λH)] × [1 - exp(-λA)]. The sensitivity chart shows the effect of changing the away scoring rate while keeping the home rate fixed.

Combination events must be summed from their intersections. A home win with both teams scoring consists of cells where h > a and a > 0. Multiplying separate home-win and BTTS probabilities imposes an additional independence assumption that generally does not hold. The matrix preserves the required overlap structure.

A full-time matrix does not describe the route taken to the final score. Half-time/full-time probabilities require a temporal model or separate first-half and second-half scoring components. The same limitation applies to first-team-to-score and in-play events: a terminal score distribution does not identify event order.

Illustrative scoreline matrix: λH = 1.60 and λA = 1.10
Home goals ↓ / Away goals →01234
06.72%7.39%4.07%1.49%0.41%
110.75%11.83%6.51%2.39%0.66%
28.60%9.46%5.20%1.91%0.52%
34.59%5.05%2.78%1.02%0.28%
41.84%2.02%1.11%0.41%0.11%

Decide whether independent Poisson is adequate

Independent Poisson is a baseline, not the finished answer in every competition. Its main advantages are transparency, coherent aggregation and low parameter requirements. Its main weaknesses concern dependence, low-score behaviour, dispersion and uncertainty in the input rates.

The Dixon-Coles adjustment modifies probabilities around 0-0, 1-0, 0-1 and 1-1 because those cells can be systematically misrepresented by naive independence. Its correction parameter must be estimated from relevant historical data. Arbitrarily boosting draws because a fixture appears tight is not a Dixon-Coles implementation.

A bivariate Poisson model adds a shared scoring component, creating positive covariance between the teams' goal counts. It can represent some common match-level influences, but its dependence structure is restrictive. Real game-state effects need not produce uniformly positive covariance: an early lead may make one team safer while forcing the other to take more risk.

Mixed-Poisson, negative-binomial or simulation models can represent overdispersion. Instead of treating λH and λA as known constants, a model can assign distributions to them and average the resulting scoreline probabilities. Propagating rate uncertainty generally produces heavier tails than a fixed-rate model using only the mean estimates.

Simulation is useful when the match model includes nonlinear events such as red cards, tactical-state transitions or lineup scenarios. Each simulation produces one scoreline; repeated simulations estimate the matrix. The assumptions should remain auditable, and the number of simulations should be sufficient to make Monte Carlo error immaterial for the probabilities being reported.

BTTS sensitivity to the away scoring rate

The home scoring rate is fixed at 1.60 while the away rate varies. Values use the independent-Poisson calculation [1 - exp(-1.60)] × [1 - exp(-λA)].

6246.53115.50λA 0.70λA 0.90λA 1.10λA 1.30λA 1.50

Illustrative scenario only.

Validate the entire distribution, not only the favourite score

A scoreline model should be evaluated on matches strictly later than the data used to estimate it. Randomly mixing past and future fixtures can leak team-strength, lineup or competition information across the split. Rolling or expanding-window evaluation better reproduces how the matrix would have been generated at the time.

Full-distribution log loss evaluates the probability assigned to the observed scoreline: lower average negative log probability is better. It heavily penalises outcomes to which the model assigned tiny probabilities, making it useful for detecting overconfident matrices. Because exact scorelines are sparse, market-level Brier scores can also be monitored for home win, draw, away win, BTTS and total-goal events.

Market-level accuracy alone is insufficient. Two matrices can generate similar win-draw-win probabilities while disagreeing materially on 0-0, 1-1, 2-1 and the high-score tail. Validation should therefore include exact-score likelihood, marginal goal distributions, the total-goal distribution and several aggregated events.

Calibration asks whether events assigned a given probability occur at roughly that rate. Exact-score cells may need to be grouped because individual outcomes are rare. Useful groupings include predicted-probability bands, goal totals and score families such as home clean-sheet wins. Assess calibration out of sample and attach uncertainty intervals; a small number of matches cannot reliably distinguish 8% from 10% events.

Implementation checks are separate from forecasting validation. Each cell must be non-negative, the complete matrix must sum to 1, row sums must recover the home marginal and column sums the away marginal. Home-win, draw and away-win probabilities must partition the matrix without overlap. These identities do not prove the football assumptions are correct, but failing them proves the implementation is wrong.

Scoreline model choices and the assumptions they change
MethodWhat it capturesMain limitationAppropriate role
Independent PoissonSeparate home and away scoring rates with a transparent joint matrixConditional independence and equal mean-variance structureBaseline model and benchmark for more complex methods
Dixon-Coles adjustmentEmpirical correction for selected low-scoring cellsDoes not provide a general model of all dependenceWhen validated low-score bias is stable in the target competition
Bivariate PoissonPositive covariance through a shared goal componentRestrictive dependence pattern and added estimation burdenWhen data support common match-level scoring variation
Mixed-Poisson or simulationRate uncertainty, overdispersion and nonlinear match scenariosMore assumptions, computation and calibration riskWhen lineups, game states or tail behaviour materially affect the use case

Treat the matrix as a probability accounting system

A useful scoreline matrix is not created by guessing several plausible scores and assigning percentages afterward. It begins with explicit scoring-rate assumptions, produces coherent marginal distributions and combines them through a stated dependence model. Every derived event is then an aggregation of the same underlying probability mass.

The independent-Poisson matrix is valuable because each step is inspectable. Its limitations are equally inspectable: rate uncertainty, overdispersion, game-state dependence and low-score bias. Retain more complex models only when they reduce out-of-sample error or improve calibration. Complexity without measurable gain makes the matrix harder to audit.

Questions about the model

How large should a football scoreline matrix be?

Expand it until the omitted tail is immaterial for the intended decision. A 0-5 or 0-6 display may be adequate for many ordinary fixtures, but high expected-goal rates require more rows and columns. Calculate omitted mass rather than assuming that a fixed grid is sufficient.

Should the visible scoreline cells add to 100%?

Only if the matrix includes every possible score or uses explicit overflow categories such as 5+. A finite grid without overflow should sum to less than 100%. The difference is the probability of scorelines outside the displayed range and should not be discarded silently.

Can bookmaker prices be used to estimate the two goal rates?

They can provide constraints after prices are adjusted for margin, especially when win-draw-win, totals and BTTS markets are used together. However, two Poisson rates may not reproduce every market simultaneously, and different rate pairs or dependence assumptions can produce similar aggregate prices. Report the fitting method and residual errors.

Is the highest-probability cell the predicted final score?

It is the modal scoreline under the model, but describing it as a single prediction can hide substantial uncertainty. If the largest cell has a probability of 12%, the model still assigns 88% to all other outcomes combined. The full distribution is usually more informative than the modal cell.

Lehua Kahale

Lehua Kahale

BTTS & Win Markets · 9 years experience