Data Science & Python Track

Abstract • Build essential statistical knowledge and how to apply it on different business questions. • Will be able to analyze your data quickly and visualize your results using SPSS and Python. • You will learn how to use Python to analyze and visualize your data using different libraries. • You’ll explore the four crucial […]

5 students enrolled

Abstract

• Build essential statistical knowledge and how to apply it on different business questions.
• Will be able to analyze your data quickly and visualize your results using SPSS and Python.
• You will learn how to use Python to analyze and visualize your data using different libraries.
• You’ll explore the four crucial steps for any data analysis project. Reading, describing, cleaning, and visualizing data.
• You will work with the most common and popular tools that data analysts use every day.
• You will be able to confidently extract knowledge and answers from data.
• Build your expertise in the most widely used AI & ML tools and technologies.
• Understand the complete Data Science life cycle and its application on different Data Sets
• Acquire the ability as a data scientist to independently solve business problems using machine learning.
• Learn how data scientists exactly work in completing an end-to-end project. Starting from the understanding of a business problem then, data visualization, data preparation and moving through the whole data pipeline till deployment and reporting the results in the most effective way.
• Apply the whole data science life cycle on different real world use cases (end to end)

Instructor

Prof. Hazem Shatila, Virginia Tech University, USA
Eng. Ahmed Yehia, Markov
Eng.Engi Amin, Markov
Dr|sahar
Amr helal

Duration

4 Months

Sessions

Sunday & Tuesday 7.00PM -10.00pm

Location

Live Streaming (on-line)

Prerequisites

None

1- Introduction to Python (30hrs)

Part 1: Introduction to Programming

1
• Why Should you Learn to Write Programs?
2
• Understanding programming

1. Thinking like a programmer

2. Representing problems and solutions using algorithms, flowcharts, or

pseudocode

3
• Types of Programming Languages

1. Low-level languages

2. High-level languages

4
• Building Blocks and Terminologies

1. Interpreter vs compiler

2. Input, output, sequential execution, conditional execution, repeated execution, reuse

3. Types of errors

4. Debugging

5
• Installing and Running Python

Part 2: Python Basics

1
• Variables, Expressions, and Statements

1. Variables and variable types

2. Python statements (print & assignment)

3. Python operators (arithmetic & string)

4. Python expressions

5. Asking the user for input

6. Notes on python syntax and commenting

2
• Conditional Execution

1. Boolean expressions

2. Logical operators

3. Conditional executing using IF statement

4. Alternative execution

5. Chained conditionals

6. Nested conditionals

3
• Functions

1. Built-in functions

2. User-defined functions

3. Variable scope

4
• Iteration

1. What is an iteration or a loop?

2. Finite loops

• Definite (For loops)

• Indefinite (While loops)

3. Infinite loops

5
• Lists

1. List properties

2. Traversing a list

3. List operations

4. List methods

5. Deleting items

6. Lists and functions

6
• Strings

1. String properties

2. Traversal through string

3. String methods

7
• Dictionaries and Tuples

1. Dictionary/Tuples properties

2. Creating Dictionaries/tuples

3. Accessing items in a dictionary/tuple

8
• Files

1. Opening files

2. Reading files

3. Searching through a file

4. Letting the user choose file name

5. Writing files

Part 3: Exploratory Data Analysis using Python

1
• Essential Python Libraries

1. NumPy

2. Pandas

3. Matplotlib

4. Seaborn

2
• Jupyter Notebooks and Google Colaboratory
3
• Numpy Basics

1. Arrays

2. Numpy universal functions

4
• Pandas Basics

1. Pandas data structures

2. Indexing, selection, and filtering

3. Summarizing and computing descriptive statistics

5
• Data Loading, Storage, and File Formats
6
• Data Wrangling

1. Assessing Data

• Quality issues

• Structural issues

2. Data Cleaning and Tidying

• Sub-setting data

• Merging datasets

• Joining datasets

• Reshaping data

7
• Exploratory Data Analysis and Visualization

1. Data aggregation and group operations

2. Pivot tables and cross tabulations

3. Plotting with matplotlib/seaborn

2- Data Analysis & Visualization Using Python (30hrs)

Module 1: introduction to basics of Statistics.

1
Introduction
2
Types of Variables
3
Sampling Techniques
4
Sample size

Module 2: Data preparation and Descriptive Statistics using SPSS-software

1
Introduce SPSS
2
Data coding, and data entry
3
Importing data to SPSS from excel
4
Data manipulation in SPSS
5
Data checking and editing in SPSS Introduction to descriptive statistics using SPSS
6
Measures of central tendency for different types of data
7
Measures of Variability for different types of data
8
One-way Tabulation distinct types of variables
9
Two-way and 3-way tabulations for distinct types of variables

Module 3: Data Analysis & Visualization using Python Introduction to Python.

1
Descriptive Statistics
2
Python Programming
3
Data Types and Operators
4
Data Structures
5
Control Flow
6
Functions
7
Scripting
8
Working with Data in Python
9
Data Analysis
10
Pandas and NumPy
11
Data Analysis Process
12
Data Mining
13
Data Wrangling
14
Assessing and Cleaning Data
15
Exploratory Data Analysis
16
Anomaly Detection

Data Visualization

1
Introduction to Data Visualization Tools
2
Basic and Specialized Visualization
3
Matplotlib and Seaborn
4
Different Data Charts
5
Advanced Visualization Tools
6
Heat Maps Plotting
7
Word Cloud
8
Folium Maps
9
Intensity Maps
10
Geospatial Maps in Python
11
Choropleth Maps

Data Prediction Models

1
Linear Regression
2
Logistic Regression

Use Case-1

1
Explore US Bikeshare Data Use python to understand US bikeshare data. Calculate statistics and build an interactive environment where a user can choose the data and filter for a dataset to show.

Use Case-2

1
House Sales in King County, USA Analyze and predict housing prices using attributes or features such as square footage, number of bedrooms, number of floors and so on

Use Case-3

1
Audience Interest in Data Science Topics Use python to generate visualization plots to summarize the results of a survey that was conducted to gauge an audience interest in different data science topics

Use Case-4

1
San Francisco Incidents Distribution Use python Folium maps to generate a Choropleth map of the crime rate in San Francisco, based on data for one year, and show distribution of different crimes’ type ities

3- Data Science and Python (36hrs)

Module 1: Statistics

Linear Algebra & Probability for Data Science

1
• Descriptive Statistics

1- Measures of Central Tendency

2- Measures of Variability

3- Skewness and Outliers

2
• Probability

1. Introduction to Probability

2. Probability Laws

3. Bayesian Theorem

4. Probability Distribution

5. Gaussian Distribution

6. Sampling Distribution

7. Central Limit Theorem

3
• Normalization.

1. Z-score

2. Min-Max Method

3. Decimal Scaling Method

4
• Inferential Statistics

1. T-Test and ANOVA

2. Chi-Square Test

3. Spearman Correlation Coefficient

4. Pearson Correlation Coefficient

5. Regression Analysis

5
• Linear Algebra Review

1. Review on Matrices

2. Operations on Matrices

3. Eigen Values and Eigen Vectors.

4. Dimensionality Reduction (Principal Component Analysis - PCA)

Module 2: Data Science and Machine Learning

1
• Introduction to Artificial Intelligence
2
• Introduction to Data Science
3
• Data Science life cycle.
4
• Introduction to Machine Learning & Data Mining
5
• Machine Learning
6
• Data Mining
7
• Supervised and Unsupervised Learning
8
• Types of Data
9
• Data Preprocessing
10
• Frequent Item Sets
11
• Association Rules & Apriori Algorithm
12
• Regression for Data Science

1. Linear Regression

2. Polynomial Regression

13
• Bias and Variance
14
• Base Classifiers for Data Science

1. Logistic Regression

2. Decision Tree based Methods

3. K-Nearest Neighbor

4. Neural Networks

5. Naïve Bayes

6. Support Vector Machines

15
• Clustering

1. K-mean Clustering

2. Hierarchical Clustering

3. Cluster Evaluation

16
• Evaluation of Learning Models for a Data Scientist

1. F1-Score

2. ROC

3. Lift Curves

Module 3: Python for Data Science & Machine Learning

1
• Python Basics

1. General Syntax

2. Data Types

2
• Python Data Structures

1. Lists

2. Tuples

3. Sets

4. Dictionaries

3
• Python Programming Fundamentals

1. Functions

2. Methods

3. Loops

4. Conditional Statements

5. Classes and Objects

4
• Data Science Libraries

1. Numpy

2. Pandas

3. Matplotlib

4. Seaborn

5. Sklearn

Module 4: Class Projects

1
• Project 1: Market Basket Analysis (Apriori)

1. Importing Packages

2. Loading the Data

3. Data Preprocessing

4. Build Frequent Item Set

5. Crating Association Rules

2
• Project 2: Automotive Price Prediction (Linear Regression Algorithim)

1. Importing packages

2. Loading the data

3. Date Preprocessing

4. Creating Linear Model

5. Evaluating the Model

3
• Project 3: Fraud Detection (Logistic Regression Algorithm)

1. Importing Packages

2. Loading the Data

3. Data Exploration

4. Date Preprocessing

5. Split the Data (train & Test)

6. Train Logistic Regression Algorithm

7. Test the Trained Model

8. Evaluating the Model

4
• Project 4: Customer Churn Prediction (Neural Networks Algorithm)

1. Importing Packages

2. Loading the Data

3. Date Preprocessing

4. Split the Data (Train & Test)

5. Train NN Algorithm.

6. Test the trained Model

7. Evaluating the Model

5
• Project 5: Customer Segmentation (K-Means Clustering)

1. Importing Packages

2. Loading the Data.

3. Date Preprocessing

4. Choose the Optimum Number of Clusters

5. Apply K-Means

6. Visualize the Output

Be the first to add a review.

Please, login to leave a review