MindCluster

MindCluster is a deep learning system that supports NPUs (Ascend AI Processors), providing robust cluster-level solutions for both training and inference jobs. By leveraging MindCluster, partners can rapidly deploy deep learning platforms while significantly reducing the development overhead associated with underlying resource scheduling.

Key Features of MindCluster 26.1.0

The key updates in this MindCluster version are as follows:

  • ToolBox: Add the DSA stress test, P2P stress test, and PRBS stream test to Atlas 350 accelerator card; adapt existing functions to Atlas 850E, Atlas 650E server, and Atlas 950 SuperPoD, including information query, performance test, and fault diagnosis.
  • Cluster scheduling components: Support device management, affinity scheduling, metric monitoring, fault detection, and resumable training for Atlas 850E, Atlas 650E server, and Atlas 950 SuperPoD; support load-based elastic scaling and container snapshot capabilities in inference scenarios; support the RDMA device plugin for the 1825 NIC.
  • Fault diagnosis tools: Add the fault mode library for Ascend 950 products and the pyMotor + vLLM fault mode library.

For details about the complete version mapping, version compatibility, updates, and upgrade impact, see the Release Notes.

MindCluster Documents

View MindCluster documents from the dimensions of performance test, cluster scheduling, fault diagnosis, and reference.

Table 3 Fault diagnosis

Fault Diagnosis

Ascend FaultDiag

Ascend FaultDiag Toolkit