Get the App
SLTechnology News&Howtos  ›  Servers  › 

Running Hadoop and Mirror making method in Docker

Shulou Source: shulou.com Published: 2022-06-02 01:16:06 09月30日 Update

Repeat the wheel, here use repackaging to generate a Docker-based Hadoop image

   Hadoop cluster relies on software such as jdk, ssh, etc., so as long as these two items are packaged in the image with Hadoop correlation.

Profile preparation

1. Hadoop-related configuration files: core-site.xml, hdfs-site.xml, mapred-site.xml, yarn-site.xml, slaves, hadoop-env.sh

2. Ssh configuration file: ssh_config

3. Hadoop cluster startup file: start-hadoop.sh

Make a mirror image

1. Installation dependency

RUN apt-get update & &\ apt-get install-y openssh-server openjdk-8-jdk wget

2. Download the Hadoop package

RUN wget http://mirror.bit.edu.cn/apache/hadoop/common/hadoop-2.10.0/hadoop-2.10.0.tar.gz & &\ tar-xzvf hadoop-2.10.0.tar.gz & &\ mv hadoop-2.10.0 / usr/local/hadoop & &\ rm hadoop-2.10.0.tar.gz &\ rm / usr/local/hadoop/share/doc-rf

3. Configure environment variables

ENV JAVA_HOME=/usr/lib/jvm/java-8-openjdk-amd64 ENV HADOOP_HOME=/usr/local/hadoop ENV PATH=$PATH:/usr/local/hadoop/bin:/usr/local/hadoop/sbin

4. Generate SSH key for secret-free login of nodes

RUN ssh-keygen-t rsa-f ~ / .ssh/id_rsa-P''& &\ cat ~ / .ssh/id_rsa.pub > > ~ / .ssh/authorized_keys

5. Create Hadoop-related directories, copy relevant configuration files, add execution permissions to related files, and finally format namenode nodes. When each node starts, start the ssh service

RUN mkdir-p ~ / hdfs/namenode & &\ mkdir-p ~ / hdfs/datanode & &\ mkdir $HADOOP_HOME/logsCOPY config/* / tmp/# copy ssh, Hadoop configuration related RUN mv / tmp/ssh_config ~ / .ssh/config & &\ mv / tmp/hadoop-env.sh / usr/local/hadoop/etc/hadoop/hadoop-env.sh & &\ mv / tmp/hdfs-site.xml $HADOOP_HOME/etc/hadoop/hdfs-site.xml & &\ mv / tmp/core-site.xml $HADOOP_HOME/etc/hadoop/core-site.xml & &\ mv / tmp/mapred-site.xml $HADOOP_HOME/etc/hadoop / mapred-site.xml & &\ mv / tmp/yarn-site.xml $HADOOP_HOME/etc/hadoop/yarn-site.xml & &\ mv / tmp/slaves $HADOOP_HOME/etc/hadoop/slaves & &\ mv / tmp/start-hadoop.sh ~ / start-hadoop.sh & &\ mv / tmp/run-wordcount.sh ~ / run-wordcount.sh# add execution permission RUN chmod + x ~ / start-hadoop.sh & &\ chmod + x ~ / run-wordcount.sh & &\ chmod + x $HADOOP_HOME/sbin/start-dfs.sh & &\ chmod + x $HADOOP_HOME/sbin/start-yarn.sh # format namenodeRUN / usr/local/hadoop/bin/hdfs namenode-format

Run Hadoop Cluster in Docker

After    generates an image through the above Dockerfile, you can use the image generated above to build a Hadoop cluster. Here, one master and two slave nodes are started.

Add a bridged network:

Docker network create-driver=bridge solinx-hadoop

Start the Master node:

Docker run-itd-- net=solinx-hadoop-p 10070 net=solinx-hadoop 50070-p 8088 name solinx-hadoop-master-- hostname solinx-hadoop-master solinx/hadoop:0.1

Start the Slave1 node:

Docker run-itd-net=solinx-hadoop-name solinx-hadoop-slave1-hostname solinx-hadoop-slave1 solinx/hadoop:0.1

Start the Slave2 node:

Docker run-itd-net=solinx-hadoop-name solinx-hadoop-slave2-hostname solinx-hadoop-slave1 solinx/hadoop:0.1

Enter the Master node and execute the script to start the Hadoop cluster:

Summary

The above is the editor to introduce to you in the Docker running Hadoop and mirror production, I hope to help you, if you have any questions, please leave me a message, the editor will reply to you in time. Thank you very much for your support to the website!

If you think this article is helpful to you, you are welcome to reprint it, please indicate the source, thank you!

Tags: Nodes images files configurations clusters generation production permissions help two provenance variables formats environments questions directories websites networks scripts wheels Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno MariaDB Shulou Technology Apple Shulou Tech Info macOS