Statistics

Fundamentals of Statistics for Data Analyst

By Andrew Sam

Students and professionals who are just getting started in Data Science learning journey and have basic knowledge of statistics can better leverage data insights. Statistics is very important for all the sciences especially computer and information sciences, biological science and physical sciences. Statistical techniques help you to find structures in data and improve on your insights. Karl Pearson a British mathematician once stated “Statistics is the grammar of science.” Some of the fundamentals of statistics are mentioned below:

Random Variables

The outcome of random processes can be stored in random variables. For example, a random process can be flipping a coin and its value can be stored in a random variable X which can be 1 in case of head and 0 in case of tail or vice versa. Here in case of flipping a coin there are two possible outcomes – {0,1} this is called sample space. When a particular random process is repeated it can be called as an event. In event when you want to address the chances of particular outcome, that is called probability. 

Mean, Variance and Standard-deviation

Generally, it is not possible to run calculations on the entire vast population of observations. So, a small sample is taken into account which represents almost similar structure to the entire population of observations.  

If a finite set of numbers is taken and central value is tracked then this central value is called Mean. It is also called as the average value of set of numbers. Let’s assume that the set has following values x1, x2, x3,…xn then there mean can be denoted by 

The variance measures how far individual data points are compared to their average value. Variance can also be understood as sum of square of difference of data values and average (mean). 

Standard deviation has same unit as data points so many a times is considered more than variance. Standard deviation can be found by taking square root of variance. A major drawback in standard deviation is that it can be impacted by outliers. 

Covariance

The covariance is a measure of the joint variability of two random variables and describes the relationship between these two variables. It is defined as the expected value of the product of the two random variables’ deviations from their means. 

Correlation

The correlation is also a measure for relationship and it measures both the strength and the direction of the linear relationship between two variables. If a correlation is detected then it means that there is a relationship or a pattern between the values of two target variables. 

Probability Distribution Functions

A function that describes all the values that any sample space or random variable or probability can take is called probability distribution functions (pdfs). Pdfs are also known as probability density. These functions should satisfy two conditions first being, the sum of all probabilities should be 1. The Second condition being the value of every probability must be between 0 and 1. Probability functions are classified into two categories: continuous and discrete. Discrete probability functions are the ones in which the output sets are limited and away. For example, in flipping coin example mentioned above there were only two outcome possibilities. Whereas in case of continuous probability distribution the sample space is continuous and may also be in a sequence. 

Binominal Distributions

Binominal distribution is the discrete probability distribution of the number of successes in a sequence of an independent experiments, each with a Boolean valued outcome: p for success and q for failure. Assume that random variable x follows binomial distribution. Then the probability of observing k successes in an independent trial can be expressed as  

Poisson Distribution

Poisson’s distribution is discrete probability distribution function of number of events occurring in a specified time period, given the average number of times that event occurs in a particular time period. 

Normal Distribution

Unlike Poisson’s distribution, Normal distribution is a continuous probability distribution function. Normal distribution is one of the most popular probability distribution functions and is also called as Gaussian distribution.  

Bayes Theorem

Named after famous statistician Thomas bayes, Bayes theorem is arguably the most powerful theorem of statistics and probability. This theorem is also known at Bayes law. Bayes theorem defines the probability of event, where information about conditions that might affect the event is available. For example, there are higher chances of older people getting infected to virus x. So, Bayes theorem could help us to include person’s age as a factor in determining probability of him getting infected. The concept of conditional probability is at the base of bayes theorem. Bayes theorem can be described as following: 

Pr(X|Y): The probability of event X occurring given that event Y has already occurred. 

Pr(Y|X): The probability of event Y occurring given that event X has already occurred. 

Pr(X) and Pr(Y) stands for their individual probability of occurrence. 

Here we come to an end of fundamentals of statistics. Do consider subscribing to our newsletter to be up to date with latest happenings in Advanced tech industry and also to get such timely tutorial tech guides.