Get in Touch
 Duration 21 hours

Course Outline

Introduction

This section offers a foundational overview of when to apply 'machine learning,' key considerations, and the broader context, including advantages and disadvantages. It explores data types (structured, unstructured, static, and streamed), data validity and volume, the distinction between data-driven and user-driven analytics, and the differences between statistical and machine learning models. Additionally, it addresses challenges in unsupervised learning, the bias-variance trade-off, iterative evaluation, cross-validation strategies, and the paradigms of supervised, unsupervised, and reinforcement learning.

MAJOR TOPICS

1. Understanding Naive Bayes

  • Foundational concepts of Bayesian methods
  • Probability theory
  • Joint probability
  • Conditional probability and Bayes' theorem
  • The Naive Bayes algorithm
  • Classification using Naive Bayes
  • The Laplace estimator
  • Incorporating numeric features into Naive Bayes

2. Understanding Decision Trees

  • The divide and conquer approach
  • The C5.0 decision tree algorithm
  • Selecting optimal splits
  • Pruning decision trees

3. Understanding Neural Networks

  • Evolution from biological to artificial neurons
  • Activation functions
  • Network topology
  • Determining the number of layers
  • Direction of information flow
  • Sizing the number of nodes per layer
  • Training networks via backpropagation
  • Deep Learning

4. Understanding Support Vector Machines

  • Classification using hyperplanes
  • Identifying the maximum margin
  • Handling linearly separable data
  • Handling non-linearly separable data
  • Utilizing kernels for non-linear spaces

5. Understanding Clustering

  • Clustering as a machine learning objective
  • The k-means clustering algorithm
  • Using distance metrics for cluster assignment and updates
  • Determining the optimal number of clusters

6. Measuring classification performance

  • Processing classification prediction data
  • Analyzing confusion matrices
  • Assessing performance with confusion matrices
  • Performance metrics beyond accuracy
  • The kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance trade-offs
  • ROC curves
  • Predicting future performance
  • The holdout method
  • Cross-validation
  • Bootstrap sampling

7. Optimizing standard models for enhanced performance

  • Automated parameter tuning with caret
  • Developing simple tuned models
  • Customizing the tuning workflow
  • Boosting model performance via meta-learning
  • Concepts of ensembles
  • Bagging techniques
  • Boosting techniques
  • Random forests
  • Training random forests
  • Evaluating random forest performance

MINOR TOPICS

8. Understanding classification via nearest neighbors

  • The kNN algorithm
  • Distance calculation
  • Selecting an appropriate k value
  • Data preparation for kNN
  • The lazy nature of the kNN algorithm

9. Understanding classification rules

  • The separate and conquer strategy
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Understanding regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlations
  • Multiple linear regression

11. Understanding regression trees and model trees

  • Integrating regression into tree structures

12. Understanding association rules

  • Apriori algorithm for association rule learning
  • Measuring rule interest: support and confidence
  • Generating rule sets using the Apriori principle

Extras

  • Spark, PySpark, MLlib, and Multi-armed bandits

Requirements

Knowledge of Python

Number of participants


Price per participant

Testimonials (7)

Upcoming Courses

Related Categories