Machine mac book, virtualbox4.3.6, and virtualbox are installed with ipvt13.10. In the multi-point distribution environment, After configuring a machine, clone the other two machines. There are three machines in total.
1. Configure the Environment
Bash language: sudo apt-get install-y openjdk-7-jdk openssh-server
Sudo addgroup hadoop
Sudo adduser-ingroup hadoop# Create password
Sudo shortdo
Hadoop ALL = (ALL) ALL# Hadoop user can use sudo
Su-hadoop# Need password
Ssh-keygen-t rsa-P ""# Enter file (/home/hadoop/. ssh/id_rsa)
Cat/home/hadoop/. ssh/id_rsa.pub>/home/hadoop/. ssh/authorized_keys
Wget http://apache.fayea.com/apache-mirror/hadoop/common/hadoop-2.3.0/hadoop-2.3.0.tar.gz
Tar zxvf hadoop-2.3.0.tar.gz
Sudo cp-r hadoop-2.3.0 // opt
Cd/opt
Sudo ln-s hadoop-2.3.0 hadoop
Sudo chown-R hadoop-hadoop hadoop-2.3.0
Sed-I '$ a \ nexport JAVA_HOME =/usr/lib/jvm/java-7-openjdk-amd64 'hadoop/etc/hadoop/hadoop-env.sh
2. Configure hadoop single Node environment
Cp mapred-site.xml.template mapred-site.xml
Vi mapred-site.xml
<Property>
<Name> mapreduce. cluster. temp. dir </name>
<Value> </value>
<Description> No description </description>
<Final> true </final>
</Property>
<Property>
<Name> mapreduce. cluster. local. dir </name>
<Value> </value>
<Description> No description </description>
<Final> true </final>
</Property>
Vi yarn-site.xml
<Property>
<Name> yarn. resourcemanager. resource-tracker.address </name>
<Value> 127.0.0.1: 8021 </value>
<Description> host is the hostname of the resource manager and port is the port on which the NodeManagers contact the Resource Manager.
</Description>
</Property>
<Property>
<Name> yarn. resourcemanager. schedager. address </name>
<Value> 127.0.0.1: 8022 </value>
<Description> host is the hostname of the resourcemanager and port is the port on which the Applications in the cluster talk to the Resource Manager.
</Description>
</Property>
<Property>
<Name> yarn. resourcemanager. schedager. class </name>
<Value> org. apache. hadoop. yarn. server. resourcemanager. schedity. capacity. capacityschedity </value>
<Description> In case you do not want to use the default scheduler </description>
</Property>
<Property>
<Name> yarn. resourcemanager. address </name>
<Value> 127.0.0.1: 8023 </value>
<Description> the host is the hostname of the ResourceManager and the port is the port on which the clients can talk to the Resource Manager. </description>
</Property>
<Property>
<Name> yarn. nodemanager. local-dirs </name>
<Value> </value>
<Description> the local directories used by the nodemanager </description>
</Property>
<Property>
<Name> yarn. nodemanager. address </name>
<Value> 0.0.0.0: 8041 </value>
<Description> the nodemanagers bind to this port </description>
</Property>
<Property>
<Name> yarn. nodemanager. resource. memory-mb </name>
<Value> 10240 </value>
<Description> the amount of memory on the NodeManager in GB </description>
</Property>
<Property>
<Name> yarn. nodemanager. remote-app-log-dir </name>
<Value>/app-logs </value>
<Description> directory on hdfs where the application logs are moved to </description>
</Property>
<Property>
<Name> yarn. nodemanager. log-dirs </name>
<Value> </value>
<Description> the directories used by Nodemanagers as log directories </description>
</Property>
<Property>
<Name> yarn. nodemanager. aux-services </name>
<Value> mapreduce_shuffle </value>
<Description> shuffle service that needs to be set for Map Reduce to run </description>
</Property>
Supplemental Configuration:
Mapred-site.xml
<Property>
<Name> mapreduce. framework. name </name>
<Value> yarn </value>
</Property>
Core-site.xml
<Property>
<Name> fs. defaultFS </name>
<Value> hdfs: // 127.0.0.1: 9000 </value>
</Property>
Hdfs-site.xml
<Property>
<Name> dfs. replication </name>
<Value> 1 </value>
</Property>
Bash language: cd/opt/hadoop
Bin/hdfs namenode-format
Sbin/hadoop-daemon.sh start namenode
Sbin/hadoop-daemon.sh start datanode
Sbin/yarn-daemon.sh start resourcemanager
Sbin/yarn-daemon.sh start nodemanager
Jps
# Run a job on this node
Bin/hadoop jar share/hadoop/mapreduce/hadoop-mapreduce-examples-2.3.0.jar pi 5 5 10
3. Running Problem
14/01/04 05:38:22 INFO ipc. client: Retrying connect to server: localhost/127.0.0.1: 8023. already tried 9 time (s); retry policy is RetryUpToMaximumCountWithFixedSleep (maxRetries = 10, sleepTime = 1 SECONDS)
Netstat-atnp # found tcp6
Solve:
Cat/proc/sys/net/ipv6/conf/all/disable_ipv6 #0 means ipv6 is on, 1 means off
Cat/proc/sys/net/ipv6/conf/lo/disable_ipv6
Cat/proc/sys/net/ipv6/conf/default/disable_ipv6
Ip a | grep inet6 # have means ipv6 is on
Vi/etc/sysctl. conf
Net. ipv6.conf. all. disable_ipv6 = 1
Net. ipv6.conf. default. disable_ipv6 = 1
Net. ipv6.conf. lo. disable_ipv6 = 1
Sudo sysctl-p # have the same effect with reboot
Sudo/etc/init. d/networking restart
4. Cluster setup
Config/opt/hadoop/etc/hadoop/{hadoop-env.sh, yarn-env.sh}
Export JAVA_HOME =/usr/lib/jvm/java-7-openjdk-amd64
Cd/opt/hadoop
Mkdir-p tmp/{data, name} # on every node. name on namenode, data on datanode
Vi/etc/hosts # hostname also changed on each node
192.168.1.110 cloud1
192.168.1.112 cloud2
192.168.1.114 cloud3
Vi/opt/hadoop/etc/hadoop/slaves
Cloud2
Cloud3
Core-site.xml
<Configuration>
<Property>
<Name> fs. defaultFS </name>
<Value> hdfs: // cloud1: 9000 </value>
</Property>
<Property>
<Name> io. file. buffer. size </name>
<Value> 131072 </value>
</Property>
<Property>
<Name> hadoop. tmp. dir </name>
<Value>/opt/hadoop/tmp </value>
<Description> A base for other temporary directories. </description>
</Property>
</Configuration>
It is said that dfs. datanode. data. dir needs to be cleared; otherwise, datanode cannot be started.
Hdfs-site.xml
<Property>
<Name> dfs. namenode. name. dir </name>
<Value>/opt/hadoop/name </value>
</Property>
<Property>
<Name> dfs. datanode. data. dir </name>
<Value>/opt/hadoop/data </value>
</Property>
<Property>
<Name> dfs. replication </name>
<Value> 2 </value>
</Property>
Yarn-site.xml
<Property>
<Name> yarn. resourcemanager. address </name>
<Value> cloud1: 8032 </value>
<Description> ResourceManager host: port for clients to submit jobs. </description>
</Property>
<Property>
<Name> yarn. resourcemanager. schedager. address </name>
<Value> cloud1: 8030 </value>
<Description> ResourceManager host: port for ApplicationMasters to talk to schedager to obtain resources. </description>
</Property>
<Property>
<Name> yarn. resourcemanager. resource-tracker.address </name>
<Value> cloud1: 8031 </value>
<Description> ResourceManager host: port for NodeManagers. </description>
</Property>
<Property>
<Name> yarn. resourcemanager. admin. address </name>
<Value> cloud1: 8033 </value>
<Description> ResourceManager host: port for administrative commands. </description>
</Property>
<Property>
<Name> yarn. resourcemanager. webapp. address </name>
<Value> cloud1: 8088 </value>
<Description> ResourceManager web-ui host: port. </description>
</Property>
<Property>
<Name> yarn. resourcemanager. schedager. class </name>
<Value> org. apache. hadoop. yarn. server. resourcemanager. schedity. capacity. capacityschedity </value>
<Description> In case you do not want to use the default scheduler </description>
</Property>
<Property>
<Name> yarn. nodemanager. resource. memory-mb </name>
<Value> 10240 </value>
<Description> the amount of memory on the NodeManager in MB </description>
</Property>
<Property>
<Name> yarn. nodemanager. local-dirs </name>
<Value> </value>
<Description> the local directories used by the nodemanager </description>
</Property>
<Property>
<Name> yarn. nodemanager. log-dirs </name>
<Value> </value>
<Description> the directories used by Nodemanagers as log directories </description>
</Property>
<Property>
<Name> yarn. nodemanager. remote-app-log-dir </name>
<Value>/app-logs </value>
<Description> directory on hdfs where the application logs are moved to </description>
</Property>
<Property>
<Name> yarn. nodemanager. aux-services </name>
<Value> mapreduce_shuffle </value>
<Description> shuffle service that needs to be set for Map Reduce to run </description>
</Property>
<! --
<Property>
<Name> yarn. nodemanager. aux-services.mapreduce_shuffle.class </name>
<Value> org. apache. hadoop. mapred. ShuffleHandler </value>
</Property>
-->
Mapred-site.xml
<Property>
<Name> mapreduce. framework. name </name>
<Value> yarn </value>
</Property>
<Property>
<Name> mapreduce. jobhistory. address </name>
<Value> cloud1: 10020 </value>
</Property>
<Property>
<Name> mapreduce. jobhistory. webapp. address </name>
<Value> cloud1: 19888 </value>
</Property>
Cd/opt/hadoop/
Bin/hdfs namenode-format
Sbin/start-dfs.sh
# Cloud1 NameNode SecondaryNameNode, cloud2 and cloud3 DataNode
Sbin/start-yarn.sh
# Cloud1 ResourceManager, cloud2 and cloud3 NodeManager
Jps
View Cluster status bin/hdfs dfsadmin-report
View File Block Composition bin/hdfs fsck/-files-blocks
NameNode view hdfs http: // 192.168.1.110: 50070
View RM http: // 192.168.1.110: 8088
Bin/hdfs dfs-mkdir/input
Bin/hadoop jar./share/hadoop/mapreduce/hadoop-mapreduce-examples-2.3.0.jar randomwriter input
5. Questions:
Q: 14/01/05 23:59:05 WARN util. NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
A:/opt/hadoop/lib/native/the dynamic link library below is 32bit and should be replaced with 64-bit
Q: Are you sure you want to continue connecting (yes/no) displayed during ssh logon )? Solution
A: Modify/etc/ssh/ssh_config and change the # StrictHostKeyChecking ask to StrictHostKeyChecking no.
Q: The DataNode of two slaves cannot be added to the cluster system,
A: Delete the content lines of 127.0.1.1 or localhost in/etc/hosts.