Get the App
SLTechnology News&Howtos  ›  Servers  › 

How to realize Principal component Analysis in data dimensionality reduction in spark mllib

Shulou Source: shulou.com Published: 2022-05-31 19:05:58 09月18日 Update

Editor to share with you how to achieve principal component analysis of data dimensionality reduction in spark mllib. I hope you will gain something after reading this article. Let's discuss it together.

The running code is as follows: package spark.DataDimensionReductionimport org.apache.spark.mllib.linalg.Vectorsimport org.apache.spark.mllib.linalg.distributed.RowMatriximport org.apache.spark. {SparkConf, SparkContext} / * * data dimensionality reduction * Principal component analysis PCA * tries to reassemble the indicators with certain related rows (such as P indicators) into a new set of independent comprehensive indicators to replace the original indicators. In order to achieve the purpose of data dimensionality reduction * Created by eric on 16-7-24. * / object PCA {val conf = new SparkConf () / / create the environment variable .setMaster ("local") / / set the localization handler .setAppName ("PCA") / / set the name val sc = new SparkContext (conf) Def main (args: Array [String]) {val data = sc.textFile (". / src/main/spark/DataDimensionReduction/a.txt") .map (_ .split (") .map (_ .toDouble)) .map (line = > Vectors.dense (line)) val rm = new RowMatrix (data) val pc = rm.computePrincipalComponents (3) / / extract principal components Set the number of principal components to 3 val mx = rm.multiply (pc) / / create principal component matrix mx.rows.foreach (println)}}

A.txt

1 2 3 45 6 7 89 0 8 76 4 2 1 the results are as follows

After reading this article, I believe you have a certain understanding of "how to achieve principal component analysis in data dimensionality reduction in spark mllib". If you want to know more about it, you are welcome to follow the industry information channel. Thank you for reading!

Tags: Composition data indicators analysis articles number code variables names done more environment purpose knowledge matrix results industry information information channels channels Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno MySQL Linux Shulou Technology Redmi MariaDB