5. The Maher Model
The Maher Model
声明:本文为本人毕业研究报告《The Exploration of Pairwise Comparison in Football Application》中的部分内容摘录与整理,仅用于学习与交流。
Introduction
Firstly, according to Maher’s research [2], we provide a brief introduction to the historical background of football prediction studies, along with a more detailed explanation.
Maher’s research highlights that the Negative Binomial distribution provides a better fit for football match data than the Poisson distribution. The Poisson distribution, typically used to model rare events within fixed time or space intervals, assumes that the mean and variance of data are equal. However, in football scoring data, the variance often exceeds the mean, a phenomenon known as overdispersion [1]. This suggests that the Poisson model may not accurately reflect the variability in goal numbers. The Negative Binomial distribution, with an additional parameter allowing the variance to exceed the mean, is better suited for modeling count data that shows significant overdispersion, such as football goals.
Over time, the understanding of football dynamics has evolved. Initially, unpredictability was seen as a key factor in football, with outcomes believed to be dominated by chance (opportunity, luck). Later studies, however, have indicated that skill plays a more decisive role than chance in determining match outcomes. Despite these findings, Maher critiques the notion that football is primarily a game of chance. He references earlier studies suggesting that chance outweighs skill in influencing football results. Nevertheless, Maher also explores the Poisson distribution’s applicability to football scoring, considering the sport’s frequent ball control and scoring opportunities. Each offensive action, though it carries a small but consistent chance of scoring, can lead to a goal distribution that approximates a Poisson distribution when aggregated over many attacks.
Model Formulation 5
This model employs the Poisson distribution to simulate the number of goals each football team scores in a match. It presumes that each team’s goal tally is independent, influenced by the team’s offensive and defensive strengths.
The model sets specific constraints that the sum of the home teams’ offensive capabilities (
For a match where team
Goals by away team
Where:
is the attacking strength of the home team , is the defensive weakness of the away team , is the defensive weakness of the home team , is the attacking strength of the away team .
With parameter constraints:
The parameters can be estimated using maximum likelihood estimation (MLE), where the log-likelihood for home teams’ scores is:
The MLEs satisfy:
(using the same approach for
Model Derivation 5
For a Poisson random variable
where
In the Maher model, the number of goals scored by the home team
The likelihood function
Log-likelihood:
Setting partial derivatives to zero for MLEs:
Hence,
These require iterative numerical methods. Initial estimates use the normalisation factor
Repeat similarly for
Conclusion
To sum up, the Maher model is more complex than previous models because it uses the Poisson distribution. This aligns more closely with real-world scenarios, allowing direct calculation of each team’s score while considering offensive and defensive strengths, improving prediction accuracy from win probabilities to score predictions.
However, real scenarios also need to account for home/away advantages and low-scoring situations. The Dixon–Coles model addresses these issues.
References
- Penn State University. Lesson 7: Multinomial response models and logit models — stat 504, 2024. [Online; accessed 9-April-2024]
- M. J. Maher. Modelling association football scores. Statistica Neerlandica, 36:109–118, 1982.
“觉得不错的话,给点打赏吧 ୧(๑•̀⌄•́๑)૭”
微信支付
支付宝支付