In addition to configuring resource information through the rank table file, you can also achieve so by combining environment variables described in this section.
This method applies only to communicator initialization of the TensorFlow network on the following products:
To configure resource information, set the following environment variables on every AI server node where training is needed. The following is an example:
1 2 3 4 5 6 | export CM_CHIEF_IP=192.168.1.1 export CM_CHIEF_PORT=6000 export CM_CHIEF_DEVICE=0 export CM_WORKER_SIZE=8 export CM_WORKER_IP=192.168.0.1 export HCCL_SOCKET_FAMILY=AF_INET |
The value of this environment variable must be an integer within the range of [0, Maximum number of devices in the server – 1].
For example, if HCCL_SOCKET_FAMILY is set to AF_INET6 but only IPv4 NICs are available on the device, IPv4 will be used instead.
export HCCL_NPU_SOCKET_PORT_RANGE="auto"
For details about the environment variable HCCL_NPU_SOCKET_PORT_RANGE, see Collective Communication in Environment Variables.
Assume that there are two AI server nodes and 16 devices (that is, eight on each AI server node) for distributed training. Before starting training processes on each device, configure the following environment variables in the corresponding shells to configure resource information.
1 2 3 4 5 | export CM_CHIEF_IP=192.168.1.1 export CM_CHIEF_PORT=6000 export CM_CHIEF_DEVICE=0 export CM_WORKER_SIZE=16 export CM_WORKER_IP=192.168.1.1 |
1 2 3 4 5 | export CM_CHIEF_IP=192.168.1.1 export CM_CHIEF_PORT=6000 export CM_CHIEF_DEVICE=0 export CM_WORKER_SIZE=16 export CM_WORKER_IP=192.168.2.1 |