Skip to the content.

Big Data Analytics

Syllabus

Lectures: 4 Teaching Hours per week

Tutorial: 0 Teaching Hours per week

Practical: 0 Teaching Hours per two weeks

Year: I

Part: I

Course Type: Core

Course Objectives

The objective of the course is to familiarize students with big data and latest trends in big data analytics with the introduction of various technologies for handling big data including Hadoop and components in Hadoop Platform. Afte rcompleting the course, students will be able to perform basic exploration of large, complex datasets and scalable big data analysis. In addition, the students will also be able to apply big data tool for advanced analytics disciplines such as predictive analytics, data mining, text analysis and statistical analysis.

Theory

  1. Fundamentals of Big Data Analytics (4 hours)
    1. Big Data and the V’s of Big Data
    2. Handling and Processing Big Data
    3. The Big Data landscape
    4. Big Data Analytics
    5. Examples of real world big data problems
  2. Technologies for Handling Big Data (8 hours)
    1. GFS, HDFS, Google Big Table
    2. Intdocution of Hadoop, and its functioning
    3. MapReduce
    4. RDDs Cloud Computing for big data
  3. Understanding Big Data Technologies Foundations (10 hours)
    1. Big data stack i.e. data source layer
    2. Ingestion layer
    3. Storage layer
    4. Security layer
    5. Visualization layer and approaches
    6. Architectural design patterns and programming models used for Real-World Applications
    7. Lambda Architecture
  4. Understanding Big data Ecosystem (12 hours)
    1. Hadoop and its ecosystem
    2. Introduction and Experimentation with Apache Flume
    3. Apache Kafka
    4. Apache Zookeeper
    5. Apache Spark
    6. Apache Mesos
    7. Apache Kudu
    8. Amazon and Google Cloud Platform
    9. Microsoft Azure
    10. Amazon Kinesis
  5. Using Big Data for Analytics (10 hours)
    1. Basic approaches to querying and exploring big data
    2. New Databases for Big data analytics
    3. Using Big Data for analytics - Classiciation Craracteristics and Comparison
    4. Apache Hbase
    5. Apache Hive
    6. Apache Cassandra
    7. Descriptive, Diagnostic, Predictive, Presciptive Analytics
    8. Stream Analytics and Location Analytics
    9. Case studies for big data analytics
  6. Machine Learning with Big Data (12 hours)
    1. Introduction to parallel, distributed and scalable machine learning;
    2. Using processing layer tools (Apache Spark, Apache Mahout) to train, evaluate and validate basic predictive models.
    3. Case Studies for application of machine learning in big data.
  7. Latest Trends and Research on Big Data Issues. (4 hours)

Evaluation Scheme

a. Internal Examination

Type Weightage
Minor tests 50%
Assignments 50%

b. Final Examination There will be five units of questions carrying 12 marks each. The question will cover all chapters of the syllabus. The evaluation Scheme will be as indicated in the table.

S.N. Chapter Hours Marks Distrubution
1 1,2 4+8 12
2 3 10 12
3 4 12 12
4 5,7 10+4 12
5 6 12 12
Total     60

References

1. Tom White. Hadoop: The Definitive guide, Storage and Analysis at Internet Scale, O'Reilly Media, 4th Edition, 2015
2. Nathan Marz andJames Warren. Big Data: Principles and best practices of scalable real time data systems, Manning Publications, First Edition, 2015
3. Mark Grover, Ted Malaska, Jonathan Seidman, Gwen Shapira, Hadoop Application Architectures: Designing Real-World Big Data Applications, O'Reilly Media, First Edition, 2015
4. Holden Karau, Andy Konwinski, Patrick Wendell, MateiZaharia. Learning Spark: Ligntning - Fast  Big Data Analysis, O'Reilly Media, 2015
5. Nataraj Dasgupta, Practical Big Data Analytics: Hands-on techniques to improve enterprise analytics and machine leaerning using Hadoop, Spark, NoSQL and R, Packt Publishing, 2018

Attributions to the Contributors:

Krischal Khanal