Probability
Probability is a single number squeezed between 0 and 1, and that number carries the entire weight of how likely something is to happen. Toss a fair coin and the answer is plain. Heads and tails are equally probable, no other outcome is possible, so each lands at one half. You can write that as 0.5, or as 50 percent, and the larger the number, the more likely the event. From that humble coin grows a discipline that reaches into statistics, finance, gambling, machine learning, game theory, and philosophy. How did a word once tied to a witness's nobility become a tool for measuring uncertainty itself? Who first dared to put numbers on chance, and why did it take so long? And how can the same mathematics describe both a casino's guaranteed profit and the trembling of particles too small to see? Those questions wait ahead.
The Latin probabilitas could mean probity, a measure of how much authority a witness carried in a European legal case. That authority often correlated with the witness's nobility. This is a strange ancestor for the modern idea, which instead measures the weight of empirical evidence and is reached through inductive reasoning and statistical inference. Richard Jeffrey traced the shift, noting that before the middle of the seventeenth century the term probable, from the Latin probabilis, meant approvable. A probable action or opinion was one that sensible people would undertake or hold in the circumstances. In legal contexts especially, the word probable could also apply to propositions backed by good evidence. That older sense lingered while the numerical sense was still being born.
Two competing camps disagree about what a probability fundamentally is. Objectivists assign numbers to describe an objective or physical state of affairs. Their most popular version is frequentist probability, which says the probability of a random event is the relative frequency of that outcome when the experiment is repeated indefinitely. It is the relative frequency in the long run. A modification called propensity probability reads probability as the tendency of an experiment to yield a certain outcome, even when it is performed only once. Subjectivists take a different path entirely, treating probability as a degree of belief. That belief has been described as the price at which you would buy or sell a bet that pays one unit of utility if an event happens and nothing if it does not, though not everyone agrees with that framing. The leading subjective version is Bayesian probability, which folds in expert knowledge alongside experimental data. The expert knowledge enters as a prior probability distribution, the data enter through a likelihood function, and their normalized product becomes a posterior distribution holding everything known to date. By Aumann's agreement theorem, Bayesian agents with similar priors end up with similar posteriors. Yet sufficiently different priors can lead to different conclusions, no matter how much information the agents share.
Gambling proves that people wanted to quantify chance throughout history, yet exact mathematical descriptions arrived much later. Games of chance gave the impetus, but fundamental issues stayed obscured by superstition. The sixteenth-century Italian polymath Gerolamo Cardano showed the value of defining odds as the ratio of favourable to unfavourable outcomes. Past Cardano's elementary work, the doctrine of probabilities dates to the 1654 correspondence of Pierre de Fermat and Blaise Pascal. Christiaan Huygens gave the earliest known scientific treatment in 1657. Jakob Bernoulli's Ars Conjectandi, published after his death in 1713, and Abraham de Moivre's Doctrine of Chances in 1718 treated the subject as a branch of mathematics. Ian Hacking's The Emergence of Probability and James Franklin's The Science of Conjecture chronicle how the very concept took shape.
Roger Cotes's Opera Miscellanea, published after his death in 1722, marks an early thread in the theory of errors. A memoir by Thomas Simpson, prepared in 1755 and printed the following year, first applied the theory to errors of observation. Its 1757 reprint laid down axioms that positive and negative errors are equally probable and that assignable limits bound all errors. Simpson also discussed continuous errors and described a probability curve. Pierre-Simon Laplace produced the first two laws of error. The first, published in 1774, expressed the frequency of an error as an exponential function of its magnitude, ignoring sign. The second, proposed in 1778, made the frequency an exponential function of the square of the error, and this is the normal distribution, also called the Gauss law. One wry remark noted that it is difficult to credit the law to Gauss, who despite his precocity had probably not made the discovery before he was two years old. Daniel Bernoulli, in 1778, introduced the principle of the maximum product of the probabilities of a system of concurrent errors. Adrien-Marie Legendre developed the method of least squares in 1805, presenting it in his Nouvelles methodes pour la determination des orbites des cometes. Unaware of Legendre, the Irish-American writer Robert Adrain, editor of The Analyst in 1808, deduced the law of facility of error and offered two proofs. Gauss gave the first proof known in Europe in 1809, and further proofs followed from Laplace, James Ivory, Friedrich Bessel, Morgan Crofton, and others. Andrey Markov introduced Markov chains in 1906, and Andrey Kolmogorov built the modern measure-theoretic theory in 1931.
Like other theories, probability is a representation of its concepts in formal terms, terms held apart from their meaning. Those terms are manipulated by the rules of mathematics and logic, and the results are translated back into the problem at hand. There have been at least two successful attempts to formalize it. In Kolmogorov's formulation, sets are interpreted as events and probability as a measure on a class of sets. In Cox's theorem, probability is taken as a primitive that is not further analyzed, and the emphasis falls on assigning consistent probability values to propositions. In both, the laws of probability come out the same except for technical details. Other ways to quantify uncertainty exist, such as Dempster-Shafer theory and possibility theory, but those are essentially different and not compatible with the usual laws of probability.
Roll a die and six results are possible. The collection of all possible results is the sample space, and the power set of that space gathers every collection of results. The subset containing 1, 3, and 5 is one such collection, the event that the die lands on an odd number. A probability assigns every event a value between zero and one, with the event made of all results assigned a value of one. For mutually exclusive events, those sharing no common results, the probability that at least one occurs is the sum of their individual probabilities. The complement of an event A is the event of A not occurring. If two events both happen on a single performance, that is their intersection or joint probability. When two events are independent, the joint probability is the product, so two flipped coins both landing heads follows that rule. Events that can never happen together are mutually exclusive, like rolling a 1 or a 2 on a six-sided die. When events overlap, the counting must avoid double work. Drawing from a 52-card deck, the chance of a heart or a face card must reckon with the 13 hearts, the 12 face cards, and the 3 cards that are both, counting the overlap once. Conditional probability, written as the probability of A given B, captures how one event shifts the odds of another. In a bag of 2 red balls and 2 blue balls, drawing a red ball first changes what remains, leaving 1 red and 2 blue for the next draw. Bayes' rule connects the odds of an event before and after conditioning, summarized as posterior proportional to prior times likelihood, a form going back to Laplace in 1774 and to Cournot in 1843.
Laplace's demon imagines a deterministic universe built on Newtonian concepts, where probability would vanish if every condition were known. A roulette wheel teases this idea. Know the force of the hand and the period of that force, and the resting number becomes a certainty, at least on a wheel not perfectly levelled, as Thomas A. Bass's Newtonian Casino revealed. That certainty would also demand knowledge of the wheel's inertia and friction, the ball's weight, smoothness, and roundness, and every variation in hand speed. In practice a probabilistic description proves more useful than tracking all of that. Physicists meet the same wall in the kinetic theory of gases, where the molecule count runs to the order of the Avogadro constant, 6.02, so only a statistical description is feasible. At the smallest scales the obstacle becomes fundamental. A discovery of early twentieth-century physics was the random character of all physical processes at sub-atomic scales, governed by quantum mechanics. The wave function evolves deterministically, but under the Copenhagen interpretation it deals in probabilities of observing, resolved by a wave function collapse when an observation is made. This loss of determinism did not win everyone over. Albert Einstein wrote to Max Born, "I am convinced that God does not play dice". Erwin Schrodinger, who discovered the wave function, also believed quantum mechanics was a statistical approximation of a deeper deterministic reality, an unease that still echoes through modern interpretations invoking quantum decoherence.
Common questions
What is probability in mathematics?
Probability is a number between 0 and 1 that describes how likely an event is to occur, where a larger number means the event is more likely. It is often expressed as a percentage from 0 percent to 100 percent, so a fair coin gives each side a probability of one half, written as 0.5 or 50 percent.
Where does the word probability come from?
The word probability derives from the Latin probabilitas, which can also mean probity, a measure of the authority of a witness in a European legal case that was often correlated with the witness's nobility. Before the middle of the seventeenth century, the term probable meant approvable.
Who developed the mathematical theory of probability?
The doctrine of probabilities dates to the 1654 correspondence of Pierre de Fermat and Blaise Pascal, with earlier elementary work by Gerolamo Cardano. Christiaan Huygens gave the earliest known scientific treatment in 1657, and Andrey Kolmogorov developed the modern measure-theoretic theory in 1931.
What is the difference between objective and subjective probability?
Objectivists assign probability to describe an objective or physical state of affairs, with frequentist probability defining it as the relative frequency of an outcome when an experiment is repeated indefinitely. Subjectivists treat probability as a degree of belief, and the most popular subjective version is Bayesian probability, which combines expert knowledge with experimental data.
How is probability used in everyday life?
Probability theory is applied in risk assessment and modeling, with the insurance industry and markets using actuarial science to set pricing and trading decisions. It is also used to design games of chance so casinos make a guaranteed profit, to analyze trends in biology and ecology, and in reliability theory for consumer products like automobiles and electronics.
Why is probability important in quantum mechanics?
Probability theory is required to describe quantum phenomena because physical processes at sub-atomic scales have a random character governed by quantum mechanics. Under the Copenhagen interpretation, the wave function deals with probabilities of observing, resolved by a wave function collapse, though Albert Einstein objected with the remark that God does not play dice.
All sources
25 references cited across the entry
- 2BookThe Logic of Statistical InferenceIan Hacking — Cambridge University Press — 1965
- 3JournalLogical foundations and measurement of subjective probabilityBruno de Finetti — 1970
- 4JournalInterpretations of ProbabilityAlan Hájek — 2002-10-21
- 5BookProbability Theory: The Logic of ScienceE.T. Jaynes — Cambridge University Press — 2003
- 6BookIntroduction to Mathematical StatisticsRobert V. Hogg et al. — Pearson — 2004
- 8A Brief History of ProbabilityWilliam Abrams — Second Moment
- 9BookQuantum leap : from Dirac and Feynman, across the universe, to human body and mindVladimir G. Ivancevic et al. — World Scientific — 2008
- 10BookThe Science of Conjecture: Evidence and Probability Before PascalJames Franklin — Johns Hopkins University Press — 2001
- 11JournalThomas Simpson and the arithmetic meanEddie Shoesmith — November 1985
- 12"Adrien-Marie Legendre" (version 9)Eugene William Seneta
- 13Markov ChainsRichard Weber — University of Cambridge
- 14JournalAndrei Nikolaevich KolmogorovPaul M.B. Vitanyi — 1988
- 15BookUnderstanding and applying basic statistical methods using RWilcox, Rand R. — 2016
- 16JournalReginald Crundall Punnett: First Arthur Balfour Professor of Genetics, Cambridge, 1912Genetics Society of America — September 2012
- 17JournalMathematical analyses of casino rebate systems for VIP gamblingJ.Z. Gao et al. — April 2011
- 18JournalManagement InsightsMichael F. Gorman — 2010
- 19BookA First course in ProbabilitySheldon M. Ross — Pearson Prentice Hall — 2010
- 20ProbabilityEric W. Weisstein
- 23Interpretations of Negative ProbabilitiesMark Burgin — 2010
- 25BookSchrödinger: Life and ThoughtW.J. Moore — Cambridge University Press — 1992