Skip to the content.

Foundation of Data Science and Analytics

Syllabus

Lectures: 4 Teaching Hours per week

Tutorial: 0 Teaching Hours per week

Practical: 0 Teaching Hours per two weeks

Year: I

Part: I

Course Type: Core

Course Objectives

The objective of the course is to provide the fundamental principles of Data Science and analytics specifics to the students. The course also provides the basic concepts on association analysis, model building processes steps and matrix factorization methods. In addition to the required basic mathematics and exploratory data analysis related statistics, the course also deals with DBMS, relational algebra, SQL and NoSQLs.

Course Outline

  1. Introduction to Data Science (2 Hrs) Data Science Hype, Why data science, Getting Past the Hype, The Current Landscape, Role of Data Scientist

  2. Data Types and Data Science Processes (7 Hrs) Facets of data: Structured data,Unstructured data, Natural language, Machine-generated data , Graph-based or network data, Audio, image, and video, Streaming data Process Overview, Defining goals, Retrieving data, Data preparation, Exploratory Data Analysis, Data Wrangling & Cleaning, Data Integration and Transformation, Data Reduction, Data modeling and Result Presentation

  3. Mathematical Foundation for Data Science (10 hours) Introduction and Descriptive Statistics : An overview of probability and statistics, Pictorial and tabular methods in descriptive statistics, Measures of central tendency, dispersion, and direction, Joint and conditional probabilities (3 Hrs) Random Variables and Probability Distributions: Random variables, Probability distributions for random variables, Expected values of discrete random variables and continuous distributions, The binomial probability distribution, The Poisson probability distribution (3 Hrs) Central limit theorem and Sample distribution concepts, Normal approximation; Hypothesis Testing Procedures: Tests about the mean of a normal population, The t-test, Z-tests for differences between two populations means, The two-sample t-test, Confidence interval for mean of normal population (4 Hrs)

  4. Regression and associated Models (11 Hrs) Empirical Models, Simple Linear Regression, MLE and Least Square Estimator, Logistic Regression, Hypothesis tests in simple linear regression, t-tests and ANOVA, Confidence intervals, Residual Analysis, Coefficient of Determination, Correlation (3 Hrs) Multiple Linear Regression, Matrix approach to Multiple Linear Regression, Polynomial Regression Models, Categorical Regressors, Indicator variables, Selection of variables and Model building (4 Hrs) Matrix Factorization (MF), Probabilistic and, Non-Negative MF, Industry Applications (4 Hrs)

  5. Modeling and validation processes for Machine Learning Techniques (8 Hrs) Supervised learning algorithms & Unsupervised learning algorithms. Modeling Process,Training /Validating model, Cross Validation methods, Predicting new observations Interpretation Measures for Model Performance and Evaluation: Classification accuracy, Confusion matrix, Sensitivity, Specificity, Precision, Recall, F-score, ROC curve, Clustering performance measures, other measures

  6. Association and Other types of Analysis (12 Hrs) Market Basket Analysis using frequent itemset, Association rules generation from transactional dataset, Apriori and other algorithms, Correlation analysis Outlier Analysis, Trend analysis, Time series analysis, Social network analysis

  7. Database and Datawarehousing (6 Hrs) DBMS fundamentals, Relational Algebra and SQL, OLTP, Datawarehouse, Multidimensional data model, Data Cubes, NoSQL, OLAP Operations

  8. Ethics and Recent Trends (4 hours) Data Science Ethics, Doing good data science, Owners of the data, Privacy aspects, Social impact, Getting informed consent, The Five Cs, Future Trends

Evaluation Scheme

a. Internal Examination

Type Weightage
Minor tests 50%
Assignments 50%

b. Final Examination There will be five units of questions carrying 12 marks each. The question will cover all chapters of the syllabus. The evaluation Scheme will be as indicated in the table.

S.N. Chapter Hours Marks Distrubution
1 1,3 2+10 12
2 2,7 7+6 12
3 4 11 12
4 5,8 8+4 12
5 6 12 12
Total     60

References

1.  Introducing Data Science: Big Data. Machine Learning and More, Using Python Tools. Cielen D, Meysman AD, Ali M. Manning, 2016
2. An Introduction to Statistical Learning: with Applications in R, Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, Springer, 1st edition, 2013
3. Applied Statistics and Probabilty for Engineers, Doglas C. Montgomery, Goerge C Runger, Wiley, 2014
4. Ethics and Data Science, D J Patil, Hilary Mason, Mike Loukides, O’ Reilly, 2018
5. Applied Data Science with Python and Jupyter: Galea A., Packt Publishing Ltd; 2018.
6. Adhikari A, DeNero J. Computational and Inferential Thinking: The Foundations of Data Science., 2017

Attributions to the Contributors:

Krischal Khanal