Machine Learning and Big Data Track

Abstract The ultimate guide to career transformation—go from novice to professional! In this class, you will learn the different techniques of machine learning and how to apply the data science life cycle on data sets. You will also get the practical knowledge of implementing the most effective methods by yourself. Moreover, you will get both […]

16 students enrolled

Abstract

The ultimate guide to career transformation—go from novice to professional! In this class, you will learn the different techniques of machine learning and how to apply the data science life cycle on data sets. You will also get the practical knowledge of implementing the most effective methods by yourself. Moreover, you will get both the theoretical underpinnings of learning along with the practical know-how needed to strongly apply these methodologies and techniques to new problems.

Instructor

Prof. Hazem Shatila, Virginia Tech University, USA
Eng. Ahmed Yehia, Markov
Eng. Amr Helal, Data Scientist, Markov

Duration

5 Months

Sessions

Sundays & Wednesdays, 7:00 to 9:30PM

Location

Live Streaming (on-line)

Prerequisites

None

1- Machine learning and Python (36hrs)

Module 1: Statistics Linear Algebra & Probability for Machine Learning

1
Descriptive Statistics
  • Introduction
  • Sampling Techniques
  • Measures of Central Tendency
  • Measures of Variability
  • Skewness and Outliers
2
Probability
  • Introduction to Probability
  • Probability Laws
  • Bayesian Theorem
  • Probability Distribution
  • Gaussian Distribution
  • Sampling Distribution
  • Central Limit Theorem
3
Normalizatio
  • Z-score
  • Min-Max Method
  • Decimal Scaling Method
4
Inferential Statistics
  • T-Test and ANOVA
  • Chi-Square Test
  • Spearman Correlation Coefficient
  • Pearson Correlation Coefficient
  • Regression Analysis
5
Linear Algebra Review
  • Review on Matrices
  • Operations on Matrices
  • Eigen Values and Eigen Vectors.
  • Dimensionality Reduction (Principal Component Analysis - PCA)

Module 2: Machine Learning

1
Introduction to Artificial Intelligence
2
Introduction to Data Science
3
Data Science life cycle.
4
Introduction to Machine Learning & Data Mining
5
Machine Learning
6
Data Mining
7
Supervised and Unsupervised Learning
8
Types of Data
9
Data Preprocessing
10
Frequent Item Sets
11
Association Rules & Apriori Algorithm
12
Regression for Data Science
  • Linear Regression
  • Polynomial Regression
13
Bias and Variance
14
Base Classifiers for Data Science
  • Logistic Regression
  • Decision Tree based Methods
  • K-Nearest Neighbor
  • Neural Networks
  • Naïve Bayes
  • Support Vector Machines
15
Clustering
  • K-mean Clustering
  • Hierarchical Clustering
  • Cluster Evaluation
16
Evaluation of Learning Models for a Data Scientist
  • F1-Score
  • ROC
  • Lift Curves

Module 3: Python for Machine Learning

1
Python Basics
  • General Syntax
  • Python objects
  • Data Types
2
Python Data Structures
  • Lists
  • Tuples
  • Sets
  • Dictionaries
3
Python Programming Fundamentals
  • Functions
  • Methods
  • Loops
  • Conditional Statements
  • Classes and Objects
4
Data Science Libraries
  • Numpy
  • Pandas
  • Matplotlib
  • Seaborn
  • Sklearn

Module 4: Class Projects

1
Project 1: Market Basket Analysis (Apriori)
  • Importing Packages
  • Loading the Data
  • Data Preprocessing
  • Build Frequent Item Set
  • Crating Association Rules
2
Project 2: Automotive Price Prediction (Linear Regression Algorithim)
  • Importing packages
  • Loading the data
  • Date Preprocessing
  • Creating Linear Model
  • Evaluating the Model
3
Project 3: Fraud Detection (Logistic Regression Algorithm)
  • Importing Packages
  • Loading the Data
  • Data Exploration
  • Date Preprocessing
  • Split the Data (train & Test)
  • Train Logistic Regression Algorithm
  • Test the Trained Model
  • Evaluating the Model
4
Project 4: Customer Churn Prediction (Neural Networks Algorithm)
  • Importing Packages
  • Loading the Data
  • Date Preprocessing
  • Split the Data (Train & Test)
  • Train NN Algorithm.
  • Test the trained Model
  • Evaluating the Model
5
Project 5: Customer Segmentation (K-Means Clustering)
  • Importing Packages
  • Loading the Data.
  • Date Preprocessing
  • Choose the Optimum Number of Clusters
  • Apply K-Means
  • Visualize the Output

2- Big Data Fundamentals (32hrs)

Module 1: Apache Hadoop essentials

1
Big Data
  • Traditional Large Scale Computing
  • What & Why Big Data?
  • Big Data Characteristics
  • Big Data Engineer career journey
2
Markov Data Cluster Overview
  • Cloudera Platform
  • Cloudera Manager UI
  • Cluster Architecture
3
Hadoop
  • Introduction to Big data and Hadoop Ecosystem
  • HDFS and YARN
  • MapReduce
  • Writing MapReduce jobs
4
Hive & SQOOP
  • Intro to SQL
  • Hive architecture on Hadoop
  • Hive analytics 2
  • SQOOP big Data Tool
5
Labs
  • Run wordCount Example MapReduce job
  • Move data from database to Hadoop cluster use case
  • Run hive SQL analytics queries on hive

Module 2: Hadoop Processing Engine [ Apache Spark ]

1
Introduction to Apache Spark
2
Spark fundamentals
3
Working With RDDs
4
Writing Spark applications
5
Spark Data aggregation
6
Spark data Persistence
7
Sparkling queries with Spark SQL
8
Running Spark on a Spark standalone cluster
9
Running Spark on a Spark multi-node cluster
10
Machine Learning in spark
11
Labs
  • Use Spark shell
  • Use Key Value Pair RDDs
  • Use RDD Persistence
  • spark machine learning [Uber trips clustering Use Case ]
  • spark SQL dataframe operations (join,groop by,reduce by)

Module 3: Hadoop Cluster Data Ingestion [Apache KAFKA]

1
Overview of ZooKeeper
2
Cluster, Nodes, Kafka Brokers
3
Consumers, Producers, Logs, Partitions, Records, Keys
4
Replicas, Followers, Leaders
5
Kafka Connect API
6
Labs
  • Create a kafka topic
  • Produce and consume messages using kafka
  • Using kafka connect

Module 4: Class Projects

1
Project 1: Website Activity Tracking

Build a user activity tracking pipeline as a set of real-time site activity (page views, searches, or other actions users may take)

2
Project 2: Building ETL PipeLine along with RDBMS

Build ETL jobs from RDBMS to Hadoop using spark batch engine

3
Project 3: Analyze website logs using spark

Analyze have huge amount of data for customer activities for your

website

4
Project 4: Twitter Data

process millions of tweets over big data engine environment

3- Advanced Big Data (40hrs)

Module 1: Bash Scripting Language

1
What is a Bash Script
2
Variables
3
Input – Different ways to supply data
4
Arithmetic
5
If Statements
6
Loops
7
Functions

Module 2: Data Warehousing with Apache Hive2

1
Data base fundamentals
  • Intro to SQL
  • Data base tables
2
Data Warehouse Concepts
  • ACID (atomicity, consistency, isolation, durability)
  • Data Warehouse
  • Data Warehouse Schemas
  • ETL
  • ELT2
3
Apache Hive
  • hive architecture on Hadoop
  • hive analytics
  • Moving data into Hive
  • Hive Operators and Functions
  • Data Warehouse ETL
4
Labs
  • Load data to Hadoop cluster use case
  • Run SQL analytics queries on hive
  • Ask and answer questions on movies datasets
  • Answer questions on covid-19 and get insights

Module 3: Hadoop Processing Engine [ Apache Spark 2]

1
Spark streaming
2
Spark Streaming Architecture and Advantages
3
Goals of Spark Streaming
4
How does Spark Streaming works
5
Spark Streaming Sources
6
Spark Streaming Operations
7
Streaming applications with kafak
8
Consume data from kafka using spark
9
produce data to kafka using spark
10
spark with yarn configurations
11
Labs
  • Produce and consume messages spark and kafka
  • Twitter streaming use case

Module 4: NO SQL Data Bases [Apache Hbase]

1
World Before NoSQL
  • RDBMS
  • Scaling Issues
2
WHAT IS HBASE
  • Habse and CAP Theorem
  • Hbase & Hadoop
3
HBASE TABLES
  • RowKey
  • Column Families ,Qualifier, Cell
  • HBASE SHELL OPERATIONS
  • Create
  • PUT (Insert)
  • GET (Select)
4
Merge And Compaction
5
DDI Concepts
  • Denormalization
  • Duplication
  • Intelligent Keys
6
Labs
  • hbase shell CRUD
  • Hbase Api

Module 5: Class Projects

1
Project 1: Twitter Online HashTag
  • Build a user activity tracking pipeline to collect new managed twitter hashtags and process big data using spark engine.
2
Project 2: Customer profile data management
  • Build ETL jobs from RDBMS files and social media to collect and aggregate customer 360 data to Hive DB.
3
Project 3: customer offers accumulators
  • Store, aggregate and accumulate customers sales websites logs to get triggered offers based on their activities.

4- Final project

Be the first to add a review.

Please, login to leave a review