As can be seen, the orange variation is clearly sampled much more than the blue variation. And how do we acknowledge this? The variance! We saw earlier that a posterior probability gets skinnier with more sample data insights, so given that the blue variation is still chubbier, we can conclude it is not sampled enough.
At this point, if we decide to randomly sample two points, one from each variation, and compare them both, what are the chances the orange variation would be higher?
If the sample from the blue variation comes from the right half of the plot, then it would have better probability to be higher
If the sample from the blue variation comes from the left half of the plot, then it would likely be lower than the orange variation
What happens if we decide on the variation to show next based on which has the higher value in this random sampling?
→ If the blue variation wins, it would then be shown next to the audience, furthering its sampling while also narrowing around a fixed probability for its true mean value. Here, we see two additional possibilities:
→ If the blue variation loses, the orange variation is shown
Therefore, sampling takes care of the explore-exploit dilemma for us, always making the best decision on our behalf.