上一篇部落格說了怎樣建立一個 Local Server 的叢集,今天說說怎樣建立一個真正的分布式叢集。
我們準備了兩個機器,如下:
192.168.0.192192.168.0.193
我們將使用這兩個機器來組成一個叢集,然後把 tensorflow task 扔到其中的某個節點上運行。 我們準備了兩個 server 程式,用來分別在兩個機器上啟動來組成一個叢集,並接收task。 建立一個 client 程式,用來向叢集提交 task。 server1.py
import sysimport timeimport tensorflow as tftry: worker1 = "192.168.0.192:8881" worker2 = "192.168.0.193:8881" worker_hosts = [worker1, worker2] cluster_spec = tf.train.ClusterSpec({ "worker": worker_hosts}) server = tf.train.Server(cluster_spec, job_name="worker", task_index=0) server.join()except KeyboardInterrupt: sys.exit()
server2.py
import sysimport timeimport tensorflow as tftry: worker1 = "192.168.0.192:8881" worker2 = "192.168.0.193:8881" worker_hosts = [worker1, worker2] cluster_spec = tf.train.ClusterSpec({ "worker": worker_hosts}) server = tf.train.Server(cluster_spec, job_name="worker", task_index=1) server.join()except KeyboardInterrupt: sys.exit()
client.py
import tensorflow as tfwith tf.Session("grpc://192.168.0.192:8881") as session: with tf.device("/job:worker/task:0"): matrix1 = tf.constant([[3., 3.]]) matrix2 = tf.constant([[2.],[2.]]) product = tf.matmul(matrix1, matrix2) result = session.run(product) print result
測試 在 192.168.0.192 上運行 “python server1.py” 在 192.168.0.193 上運行 “python server2.py” 在任意一台機器上運行 “python client.py”