About This Course
In today's data-driven landscape, processing massive datasets requires specialized distributed architectures beyond traditional databases. This course offers hands-on experience in building scalable ingestion pipelines, processing batch and real-time data streams, and managing distributed storage environments.
What You Will Learn
- ✦ Distributed computing architectures (Hadoop HDFS & YARN)
- ✦ In-memory data processing using Apache Spark & PySpark
- ✦ Large-scale data querying with Hive and Spark SQL
- ✦ Data ingestion pipelines using Sqoop and Kafka
- ✦ Scalable analytics and distributed machine learning workflows
Course Instructor
Dr. Yamina Azzi
Assistant Professor in Computer Science / Intelligent Systems and Machine Learning at Setif 1 Ferhat Abbes University