1- Introduction to Python (30hrs)
Part 1: Introduction to Programming
1. Thinking like a programmer
2. Representing problems and solutions using algorithms, flowcharts, or
pseudocode
1. Low-level languages
2. High-level languages
1. Interpreter vs compiler
2. Input, output, sequential execution, conditional execution, repeated execution, reuse
3. Types of errors
4. Debugging
Part 2: Python Basics
1. Variables and variable types
2. Python statements (print & assignment)
3. Python operators (arithmetic & string)
4. Python expressions
5. Asking the user for input
6. Notes on python syntax and commenting
1. Boolean expressions
2. Logical operators
3. Conditional executing using IF statement
4. Alternative execution
5. Chained conditionals
6. Nested conditionals
1. Built-in functions
2. User-defined functions
3. Variable scope
1. What is an iteration or a loop?
2. Finite loops
• Definite (For loops)
• Indefinite (While loops)
3. Infinite loops
1. List properties
2. Traversing a list
3. List operations
4. List methods
5. Deleting items
6. Lists and functions
1. String properties
2. Traversal through string
3. String methods
1. Dictionary/Tuples properties
2. Creating Dictionaries/tuples
3. Accessing items in a dictionary/tuple
1. Opening files
2. Reading files
3. Searching through a file
4. Letting the user choose file name
5. Writing files
Part 3: Exploratory Data Analysis using Python
1. NumPy
2. Pandas
3. Matplotlib
4. Seaborn
1. Arrays
2. Numpy universal functions
1. Pandas data structures
2. Indexing, selection, and filtering
3. Summarizing and computing descriptive statistics
1. Assessing Data
• Quality issues
• Structural issues
2. Data Cleaning and Tidying
• Sub-setting data
• Merging datasets
• Joining datasets
• Reshaping data
1. Data aggregation and group operations
2. Pivot tables and cross tabulations
3. Plotting with matplotlib/seaborn
2- Data Analysis & Visualization Using Python (30hrs)
Module 1: introduction to basics of Statistics.
Module 2: Data preparation and Descriptive Statistics using SPSS-software
Module 3: Data Analysis & Visualization using Python Introduction to Python.
Data Visualization
Data Prediction Models
Use Case-1
Use Case-2
Use Case-3
Use Case-4
3- Data Science and Python (36hrs)
Module 1: Statistics
Linear Algebra & Probability for Data Science
1- Measures of Central Tendency
2- Measures of Variability
3- Skewness and Outliers
1. Introduction to Probability
2. Probability Laws
3. Bayesian Theorem
4. Probability Distribution
5. Gaussian Distribution
6. Sampling Distribution
7. Central Limit Theorem
1. Z-score
2. Min-Max Method
3. Decimal Scaling Method
1. T-Test and ANOVA
2. Chi-Square Test
3. Spearman Correlation Coefficient
4. Pearson Correlation Coefficient
5. Regression Analysis
1. Review on Matrices
2. Operations on Matrices
3. Eigen Values and Eigen Vectors.
4. Dimensionality Reduction (Principal Component Analysis - PCA)
Module 2: Data Science and Machine Learning
1. Linear Regression
2. Polynomial Regression
1. Logistic Regression
2. Decision Tree based Methods
3. K-Nearest Neighbor
4. Neural Networks
5. Naïve Bayes
6. Support Vector Machines
1. K-mean Clustering
2. Hierarchical Clustering
3. Cluster Evaluation
1. F1-Score
2. ROC
3. Lift Curves
Module 3: Python for Data Science & Machine Learning
1. General Syntax
2. Data Types
1. Lists
2. Tuples
3. Sets
4. Dictionaries
1. Functions
2. Methods
3. Loops
4. Conditional Statements
5. Classes and Objects
1. Numpy
2. Pandas
3. Matplotlib
4. Seaborn
5. Sklearn
Module 4: Class Projects
1. Importing Packages
2. Loading the Data
3. Data Preprocessing
4. Build Frequent Item Set
5. Crating Association Rules
1. Importing packages
2. Loading the data
3. Date Preprocessing
4. Creating Linear Model
5. Evaluating the Model
1. Importing Packages
2. Loading the Data
3. Data Exploration
4. Date Preprocessing
5. Split the Data (train & Test)
6. Train Logistic Regression Algorithm
7. Test the Trained Model
8. Evaluating the Model
1. Importing Packages
2. Loading the Data
3. Date Preprocessing
4. Split the Data (Train & Test)
5. Train NN Algorithm.
6. Test the trained Model
7. Evaluating the Model
1. Importing Packages
2. Loading the Data.
3. Date Preprocessing
4. Choose the Optimum Number of Clusters
5. Apply K-Means
6. Visualize the Output