Apache Spark Machine Learning Blueprints

上QQ阅读APP看书，第一时间看更新

Chapter 2. Data Preparation for Spark ML

Machine learning professionals and data scientists often spend 70% or 80% of their time preparing data for their machine learning projects. Data preparation can be very hard work, but it is necessary and extremely important as it affects everything to follow. Therefore, in this chapter, we will cover all the necessary data preparation parts for our machine learning, which often runs from data accessing, data cleaning, datasets joining, and then to feature development so as to get our datasets ready to develop ML models on Spark. Specifically, we will discuss the following six data preparation tasks mentioned before and then end our chapter with a discussion of repeatability and automation:

Accessing and loading datasets
- Publicly available datasets for ML
- Loading datasets into Spark easily
- Exploring and visualizing data with Spark
Data cleaning
- Dealing with missing cases and incompleteness
- Data cleaning on Spark
- Data cleaning made easy
Identity matching
- Dealing with identity issues
- Data matching on Spark
- Data matching made better
Data reorganizing
- Data reorganizing tasks
- Data reorganizing on Spark
- Data reorganizing made easy
Joining data
- Spark SQL to join datasets
- Joining data with Spark SQL
- Joining data made easy
Feature extraction
- Feature extraction challenges
- Feature extraction on Spark
- Feature extraction made easy
Repeatability and automation
- Dataset preprocessing workflows
- Spark pipelines for preprocessing
- Dataset preprocessing automation

本周热推：

计算机网络 AI 3.0 Windows环境下32位汇编语言程序设计天才与算法：人脑与AI的数学思维学会提问，驾驭AI：提示词从入门到精通