標籤:
Mahout 的安裝
Mahout是Hadoop的一種進階應用程式。運行Mahout需要提前安裝好Hadoop,Mahout只在Hadoop叢集的NameNode節點上安裝一個即可,其他資料節點上不需要安裝
1.下載
2.配置環境變數
3.mahout --help
檢查Mahout是否安裝完好,看是否列出了一些演算法
當然,這種方法並不準確,可以通過接下來的步驟進行驗證。
4.mahout使用準備
a.下載一個檔案synthetic_control.data,http://archive.ics.uci.edu/ml/databases/synthetic_control/synthetic_control.data,並把這個檔案放在$MAHOUT_HOME目錄下。
b. 查看hadoop 狀態,要啟動hadoop
c.
c.建立測試目錄testdata,並把資料匯入到這個tastdata目錄中(這裡的目錄的名字只能是testdata)
[email protected]:~/$ hadoop fs -mkdir testdata #
[email protected]:~/$ hadoop fs -put /home/hadoop/mahout-distribution-0.7/synthetic_control.data testdata
d.使用kmeans演算法(這會運行幾分鐘左右)
[email protected]:~/$ hadoop jar /home/hadoop/mahout-distribution-0.7/mahout-examples-0.7-job.jar org.apache.mahout.clustering.syntheticcontrol.kmeans.Job
e.查看結果
[email protected]:~/$ hadoop fs -lsr output
如果看到以下結果那麼演算法運行成功,你的安裝也就成功了。
clusteredPoints clusters-0 clusters-1 clusters-10 clusters-2 clusters-3 clusters-4 clusters-5 clusters-6 clusters-7 clusters-8 clusters-9 data
Mahout 的安裝