Posts

Showing posts with the label Decision tree

Concept of Random Forest | Mathematics | Machine Learning | ML Algorithm | Data Science

Image
Concept of Random Forest | Mathematics | Machine Learning | ML Algorithm | Data Science Photo by David Kovalenko on Unsplash Trees don't have the same level of accuracy as the other prediction algorithm, so the random forest came up in the limelight, it uses trees as a building block to form a more powerful algorithm. In the random forest, the process of finding the root node and leaf node runs randomly and it is made of more than one decision trees. So, it is called Random Forest. Ensemble Technique:  Basically, sometimes we use more than one model together to increase the efficiency of model and accuracy of predictions. So, it is called Ensemble Technique . It has further two types i.e. Bagging and Boosting . The bagging is also known as the bootstrap aggregation. In bagging the different base models feed with the different sample of data from the main dataset for the purpose of training of the models. After training of all models, a test dataset is fed to all the trained models...

Concept of Decision Tree Classification | Machine Learning | Data Science | Mathematics

Image
Concept of Decision Tree Algorithm | Machine Learning | Data Science | Mathematics Decision Tree Algorithm for Classification Decision Tree Algorithm is one of the most popular algorithms and widely used in machine learning. It is a type of supervised learning-based algorithm, can be used for both classification and regression. Photo by Fabrice Villard on Unsplash Let's see first how it works? A simple decision tree example So, now we are enough aware of the decision tree, so let's get deeper. Impurity It is a measurement that how much our data is impure, means how much homogeneity is present in your data. Image Source: Research Gate For measuring impurity we have several measures from which we will learn these two:  1. Entropy: Entropy is nothing but the randomness in your dataset. Which increase predictability. It is directly proportional to the non-homogeneity in your dataset. It measures the purity of the split. Use:  We analyse the entropy on every node in the decision tr...