Skip to content

YunhanWangUCSD/CSE255-BigDataWithSpark-16SP

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

2 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CSE255-BigDataWithSpark-16SP

Big Data mini-projects using Apache Spark

  • Extracted top10 relatively popular tokens from 100 GB twitter raw JSON data within 80s using Spark RDD
  • Applied PCA to analyze 38,436 pieces 1-year-length weather station records and accomplish reconstruction
  • Implemented Linear Regression, K-means++ to predict weather data coefficients using the top eigenvectors
  • Built Gradient Boosted Trees and Random Forest model on RDD to train a Tree Cover Type classifier

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

No releases published

Packages

 
 
 

Contributors