About Me

My photo
Science communication is important in today's technologically advanced society. A good part of the adult community is not science savvy and lacks the background to make sense of rapidly changing technology. My blog attempts to help by publishing articles of general interest in an easy to read and understand format without using mathematics. You can contact me at ektalks@yahoo.co.uk

Saturday, 29 August 2026

What Has Probability to do With Scrabble, Dice, and Extreme Climate Events? An Introduction to Probabilities for the Non-specialist.

 The blog is written with the express aim to encourage everybody, particularly students, to carry out fun experiments with dice to learn how probabilities work.

************************************************************************

My wife and I regularly play scrabble and over the years jotter books were piling up in the filing system.  It occurred to me that  scrabble scores from our games - 359 over four years - might be a good set of data (collected unintentionally) that one can analyse to see if any identifiable patterns are present.  I thought that since in every move in scrabble chance is an overriding factor (sitting with six vowels can be so frustrating), the final scores must reflect this (the role of chance) in some way. 

I added up both scores and with a bin size of 20 points, plotted the distribution and fitted a Gaussian Probability Function to obtain the following result (slide 1)

The standard error in the mean (defined as standard deviation divided by the square root of the number of points) is 10.8, and if we repeat this exercise many times, then 2 out of 3 times, the mean score will be in the range 559 to 581  (more on this later).

This was all good fun  - I was not expecting such a good fit to a random distribution that the Gaussian represents.  Obviously repeating this exercise many times is not very practical,  and it seemed sensible to consider throwing dice as a much better option to understand outcomes in random events - so that is what I did next. 

Why a good understanding of probability is so important?    Probability measures the  likelihood of an event happening on a scale of 0 to 1. Weather forecast, risk assessment in finance and business, health test results are some examples of how we depend on probability outcome of events.  Probabilities govern every aspect of our lives - after all the most successful scientific theory (quantum mechanics) tells us that what happens in nature is governed by interacting microscopic particles that do not follow exact predictable paths but follow those defined by chance and probability. It is built into our lives and there is nothing you or I can do about it. The only option we have is to educate ourselves.

Humans are hopelessly poor in their understanding of probabilities: Daniel Kahneman's book Thinking, Fast and Slow is my favourite read.  It tells us about the various cognitive biases and mental shortcuts that human intuition depends on, making us fundamentally unequipped for statistical (probabilistic) thinking. In making decisions, we give much weight to cognitive biases like the confirmation bias (stick to what we already believe), the availability heuristic (give much importance to what readily comes to mind), gambler's fallacy (past random events influence future outcomes) etc.  The result is poor, flawed decision making.  

Collectively, we are unappreciative of the role that mathematics plays in our daily lives - many surveys find that people do not understand how mathematics governs our lives and there is little incentive to educate ourselves that way.  

In the following, I shall follow the example of my experiments with throwing dice - an excellent way to understand probabilities and outcome of random events.  Then I discuss how probabilities are related to disorder in a system, and why extreme climate events appear to be happening more frequently.

My Experiments with Six Dice:  I found this to be an ideal way to learn about probability outcome in random events.  A die is a cube with each of the six sides marked with a different digit from 1 to 6.  Throwing a die results in its top face displaying the number printed on it.  The number is totally random, and the chance of the number being any one from 1 to 6 is equally likely.

I took 6 such dice, put them in a cylindrical box, shook them well and emptied the box on to a plane surface. The exposed faces had numbers which were noted and the throw repeated a total of 100 times.  

We note that each throw is completely independent of the previous throws, and the 100 events provide an excellent statistically random dataset.  Obviously, one would expect that each of the numbers 1 to 6 will appear 100 times in the final count as there is equal chance of each number appearing in a throw.  But, remember that in random events the standard deviation is square root of N ( = ✓N) - that is (explained later) 2 out of 3 times (68% chance), the result will be within the range 士 ✓N. I found the following for the 100 throws:

        Number appearing      Frequency 

                   1                          107 

                   2                          102 

                   3                            98

                   4                          104

                   5                            93 

                   6                            96  

The results are between 90 and 110 and show that the sample is statistically sound.

In the next step, I added the six numbers displayed in each throw.  One expects the sum to lie in the range from 6 (all dice display number 1) to 36 (all dice display number 6).  In order to increase the frequency of occurrence, I binned them by adding two adjacent bins and obtained 14 data points.  These are plotted in the next slide along with a Gaussian Distribution fit to the data. Slide 2:

The slide shows a reasonable fit to the data but it is not as good as in the case of the scrabble data (Slide 1).  This represents the smaller sample size of 100 dice throws as compared to 359 for the scrabble data.  The bigger the sample size, the better approximation to a gaussian distribution would be expected.

Interestingly, the dice experiment also predicts several other probabilities that we can check in the data.  

For example, one could check the probability of all numbers being different - how many times in the 100 throws, did the six dice display all six numbers. For a totally random set, the prediction is 1.54% or 1.54 times in 100 throws.  In the data, I had 2 occasions when all six numbers were different - a really good result.

In fact, there are several ways that we can look at the data to work out probabilities of different combinations happening.  I show some of the combinations with their predicted probabilities and the actual observed events. For ease of understanding, I shall label the six die faces as A, B, C, D, E and F.  The results are shown in Slide 3.


Making Sense of the Results:  Now, we have sufficient information to analyse and understand the behaviour of systems where we observe an outcome determined by the internal ordering of its components.  The outcome is something we can observe/measure and is called a macrostate. The macrostate is formed in the different ways the internal components of system combine - we can not observe these internal states (they are called the microstates of the system).

A very good example is that of an assembly of gas molecules in a container.  The large number of gas molecules move randomly and collide with each other and with the container walls to produce the observable properties like the pressure and temperature of the gas.  Pressure and temperature are the macrostates of the gas - we can observe them.  The positions and velocities of the molecules in the container are the microstates of the system. Gas molecules change their position and velocities millions of times every second and the microstates of the system change accordingly.  Many different microstates can produce the same observable macrostate.

 Back to the Experiment with Dice:  We are in a good position to understand how the number of microstates in a system determine the likelihood of different macrostates manifesting, and provide a definition of the probability of realising a particular outcome. 

Remember that microstates are specific arrangement of components of the system and every individual microstate is equally likely.  The system fluctuates between microstates constantly.  Macrostates are general observable states.

One Die:  In one die, the six faces are numbers 1 to 6.  The throw of the die results in one of the numbers facing upwards.  This may be any number from 1 to 6.  

The macrostate of the die is the number showing on the top face and there are six possible macrostates of a die.  The microstates of the die are the six numbers on its sides - any one of them has the same likelihood of being on top.  Hence the probability of throwing any number is 1/6.  

Two Dice:  In two dice, the six faces are numbers 1 to 6. When two 6-sided dice are rolled, there are 6X6 = 36 possible microstates -> 11,12,13,...; 21,22,23,...; 31,32,33,...;  41,42,43...; 51,52,53... and 61,62,63.... 

Any of the 36 combinations has equal probability to appear on the top faces and hence its probability is 1/36.

If we wish to calculate the likelihood of a double appearing then there are six different ways (11, 22, 33, 44, 55 and 66) this may happen.  Hence there is 6 in 36 or 1 in 6 (16.7%) chance of a double appearing (probabilities are often quotes as %). If we want a double six (a specific double) then the chance to obtain a double six is 1/36 or 2.78%.

A larger number of microstates for a combination increases the likelihood of that combination appearing.  For example, if we wish to obtain a sum equal to 8 for the two up-facing side, then one needs to work out the different numbers or microstates that can add to 8.  In the 2 dice case, the possibilities are   2,6;  3,5; 4,4; 5,3 and 6,2  -  five microstates in all.  The probability of sum 8 appearing is 5/36 = 13.9%.  

Can you work out the probabilities of the sum of other sums appearing? Remember the possible range is from 2 to 12. combination 7 is the most likely at 16.67% and 4 or 10 have a probability of 8.33%.  

It is easy to verify these predictions by throwing two dice and repeating the process at least 100 times to accumulate enough statistics - the more trials the closer the results will be to the predicted probabilities.

As a final point, if we throw six normal dice, then the possible combinations (number of microstates) is 6^6 = 46,656   --> a very large number and each of them is equally probably in a throw of 6 such dice.


The Gaussian Distribution Curve: Also know as the Normal Distribution or the Bell Curve.  We are now ready to discuss as to why the frequency, with which a particular combination (macrostate) appears, seems to follow the bell-shaped curve depicted in Slides 1 and 2. In random events (the sum in a scrabble game or throwing dice), the final result depends on how the microstates have organised themselves - and that is indeed a random process independent of what had appeared in the past.  Each microstates is equally likely and what happens in a single throw is a random selection of one of the possibilities.  When the number of throws is increased substantially, the combinations with greater likelihood (greater number of microstates) happen more often and the frequency distribution starts to show a bell-shaped curve (known as the Gaussian Distribution in mathematics).  I present Slide 2 again in the following to make this point.

In the throw of 6 dice, the sum of over 30 is highly improbable as several of the dice will have to throw multiple sixes and fives; and there are only a relatively small number of combinations - for example to get a score of 34, one need to have combinations one four+5 sixes or two fives+four sixes (a total of 6 + 15 = 21 microstates) and an overall probability of 21/46656 or 0.045%.  

However, a throw for the sum to be around 20 has 32 combinations with 4221 different ways (microstates) to roll a 20 with an overall probability of 4221/46656 or 9.05%, and hence shows up more often.  

(please go back to the section on the throw of 2 dice to see how the sum of 7 was more probable than the sum of 4 or 10).

Slide 3 presents a breakdown for the probabilities of several combinations when 6 dice are thrown repeatedly.  The observed probabilities do not always agree exactly with the predictions but the agreement overall is very good indeed.  If we can extend the number of throws to 500 from 100 used for slides 2 and 3, then the agreement is expected to get much better.

The six dice measurements take a couple of hours and the analysis is another two hours of counting.  I highly recommend this exercise as it gives a great feel for probabilities happening in real life. 

Returning back to the Gaussian Distribution Curve, the question one might ask is what a bell-shaped curve to do with measurement of random events.  Most random events concentrate around a mean that is the most probable outcome - it has the highest number of microstates.  To obtain a result near either extreme will require a few microstates to concentrate at the extremes at the same time (see above example of 6 sided dice rolling).

In fact, A fundamental rule in statistics (The Central Limit Theorem) states that when you add many independent variables together, their total sum forms a bell-shaped Gaussian curve.  I now discuss the main features of the Gaussian Curve.

The Gaussian Distribution Curve: 



Gaussian distribution is an excellent way of describing and understanding random events.  If enough measurements are made the distribution of the values of a random variable follows the bell-shaped Gaussian curve.  We can quantitatively understand the process under study.  Importantly, it tells us what the most probable value of the mean is and how the probabilities of other values diminish as we look for measured values away from the mean. In particulae 99.6% of probable outcomes are with 3 standard deviation of the mean value. However, there is a finite probability of a measurement falling at 4 sigma or even at 5 or 6 sigma away from the mean albeit with vanishingly small probability - note that this value is not zero and unexpected events do happen in life.  
For example, the mean height of young boys in OECD countries follow a Gaussian distribution of mean 176 cm with a standard deviation sigma of 7.1 cm. 

 Brandon Marshal from Suffolk measured 224 cm that is more than 6 standard deviation away from the mean!  

I give another example of extreme weather events happening more frequently due to global warming - a consequence of climate change.

An Example of Extreme Weather Events:  The weather is a great example of random events and we all talk about, particularly in the UK, of how unpredictable the weather is.  What is noteworthy is that in recent years, there has been many extreme weather events relating to hot days, flooding, very heavy rain falls, droughts, wild fires etc.  Once in a decade events are happening every year and it is easy to understand why such a shift in frequency of extreme events might be expected from global warming.  Let us talk about daily maximum temperatures that have been observed with much higher frequency. The following is a qualitative description of what happens in a warming world.

In a stable climate, the distribution of  maximum temperatures at a place on a particular day of the year will follow a Gaussian distribution curve with a defined mean and a standard deviation.  Temperatures that are say 3 sigma away from the mean (on either side of the maximum) will be much less likely and could happen once a century.

However, if the distribution is shifted towards higher temperatures by half the standard deviation then temperatures that were 3 sigma away from the mean are only 2.5 sigma away on the hotter side and 3.5 sigma away on the colder side.  The probability of extremely hot days increases significantly and even mildly hotter days become much more probable - see slide below.  On the other side extremely cold days become even less probable - exactly what we have been experiencing in the UK and globally.

https://ektalks.blogspot.com/2021/12/slides-part-1-relating-to-climate.html 

End Note:






   https://math.stackexchange.com/questions/3076058/calculating-the-probability-of-obtaining-exactly-four-distinct-values-when-a-die   A pair and 4 distinct numbers               









 



No comments: