From the Atlas 300I inference card to the Ascend 950PR: A comprehensive analysis of Huawei's Ascend computing power products in 2026 | Changfan Industrial Control
Huawei Ascend Computing Product Overview: From Accelerator Cards to Supernodes
Huawei Ascend is a core system for artificial intelligence computing in China. It’s necessary to clarify a common misconception: strictly speaking, Ascend is an NPU (Neural Processing Unit), not a traditional GPU. The product is divided into four levels based on deployment scale: Accelerator Cards (Atlas 300I series inference cards, Atlas 350 accelerator cards), AI Servers (Atlas 800I A2/A3 inference servers, Atlas 800T A2 training servers), Supernodes (Atlas 900 A3 / Atlas 950 SuperPoD), and Clusters (SuperCluster). The first step in procurement is to determine which level your business belongs to: choose accelerator cards for single-card inference, inference servers for enterprise-level inference, and supernodes for large-scale model training and large-scale clusters.
Ascend’s core competitiveness lies not only in its hardware, but also in its full-stack ecosystem of “chip + CANN heterogeneous computing architecture + MindSpore framework + industry applications,” and its compliance advantage of independent control in domestic IT innovation scenarios—this is the fundamental reason why governments, financial institutions, operators, the power industry, and other sectors prefer Ascend solutions.
Inference Accelerator Card: Atlas 300I Series
The Atlas 300I series is the most widely used Ascend inference card series. It features a standard PCIe card form factor and can be directly inserted into x86 or Kunpeng servers, focusing on high-performance inference and video analytics:
• Atlas 300I Pro: 140 TOPS INT8 / 70 TFLOPS FP16, 24GB LPDDR4X, suitable for OCR recognition, content moderation, lightweight model inference, and other scenarios.
• Atlas 300I Duo: Dual-chip design, 280 TOPS INT8, 48GB/96GB LPDDR4X (408GB/s bandwidth, ECC supported), 150W low power consumption, high energy efficiency of 1.86 TOPS/W, supports real-time analysis of 256 channels of 1080P HD video, and is the primary card for smart city and intelligent transportation video structuring projects.
• Atlas 300I A2: The series flagship, 560 TOPS INT8 / 280 TFLOPS FP16, 32GB (0.8TB/s bandwidth) or 64GB (1.6TB/s bandwidth) on-chip memory with ECC support, suitable for generating large model inference, content moderation, and other high-performance computing scenarios.
In addition, there is the Atlas 350 accelerator card (112GB HBM, 1.4TB/s bandwidth, PCIe 5.0, supports mxFP4/FP8 low-precision quantization acceleration) for higher performance needs. This card can be integrated into the server as a dual-purpose card for training and inference. Selection tip: Inference cards have relatively limited memory bandwidth; running large models above 32B requires quantization compression and multi-card parallelism. Be sure to conduct actual business model testing before purchasing.
Inference Server: Atlas 800I A2 / A3
The Atlas 800I A2 is a mainstream enterprise-grade Ascend inference model: a 2U rackmount server based on four Kunlun 920 processors + eight Ascend 910B NPUs (32GB HBM per chip, 256GB HBM graphics memory), equipped with 32 DDR4 memory slots and flexible SATA/NVMe storage combinations, iBMC out-of-band management supporting IPMI and KVM over IP. It is widely used in public cloud, internet, telecom operators, government affairs, finance, universities, and the power industry, and is the standard solution for deploying large-scale inference models ranging from 7B to 70B in China.
The next generation of the Atlas 800I A3 further increases the overall HBM capacity to 768GB, achieving a computing power of 12.4 PFLOPS@mxFP4, and supports dual-machine 16 NPU full interconnection based on the Lingqu (UB) protocol, handling larger parameter models and higher concurrent inference scenarios. For training, the Atlas 800T A2 training server handles model training tasks. For enterprises, purchasing an inference server essentially means “purchasing a proven complete system”—the CPU/NPU ratio, cooling, power supply, and management capabilities are already optimized, making it far more reliable than assembling a computer themselves.
Huawei Ascend 950PR Powered by Atlas 950 SuperPoD Supernode
At the Huawei Connect 2025 conference, Huawei announced a clear roadmap for its Ascend chips: the Ascend 950PR (optimized for inference pre-filling and recommendation scenarios) will launch in Q1 2026, and the 950DT (optimized for training and inference decoding) will launch in Q4 2026. The subsequent 960 and 970 will iterate in Q4 2027 and Q4 2028 respectively – following this pace, the Ascend chip will maintain an evolution speed of ‘one generation per year, doubling computing power’.
The Atlas 950 SuperPoD supernode was released simultaneously with the chip, supporting non-blocking full interconnection of 8192 Ascend cards, achieving 8 EFLOPS FP8 / 16 EFLOPS FP4 computing power, 1152TB of memory capacity, and 16.3PB/s all-optical interconnect bandwidth. It adopts a fully liquid-cooled and orthogonal architecture design and can be further combined into an Atlas 950 SuperCluster cluster of over 500,000 cards. The significance of the supernode lies in its ability to transform thousands of cards into “logical computers” through the Lingqiao interconnect protocol, completely solving the bottleneck of inter-machine communication in large-scale model training. The implication for buyers is that when planning computing power, it is not necessary to build supernodes all at once; instead, one can start with the Atlas 800I series and smoothly expand as business grows.
Key Considerations for Ascend Computing Power Procurement
• First, determine the tier and then select the model: For edge/lightweight inference, choose the Atlas 300I series accelerator cards; for enterprise-level large-scale model inference, choose the Atlas 800I A2/A3; for training and ultra-large-scale cluster evaluation, choose the SuperNode solution.
• Computational memory and bandwidth billing: For large model inference, first estimate the memory required for the number of parameters + KV cache. Ascend cards have significantly different memory bandwidths (from LPDDR4X 408GB/s to HBM tens of TB/s), which directly determines inference throughput.
• Pay attention to software ecosystem compatibility: Ascend uses the CANN + MindSpore system. Before purchasing, confirm that the target model has compatible case studies and allow time for model migration and optimization.
• Information creation compliance verification: Confirm the localization rate and procurement catalog requirements for government affairs, finance, and operator projects. Ascend solutions have a natural advantage in information creation scenarios.
• Overall Reliability: Verify redundant power supplies, hot-swappable fans, out-of-band management of the iBMC, and complete system aging test reports. Complete system assembly is not permitted in the production environment.
• Supply and Roadmap: The Ascend chip iteration cycle is clear (950PR→950DT→960). Confirm the delivery cycle and compatibility path for subsequent upgrades at the time of purchase.
Why Choose Changfan Industrial Control?
Changfan Industrial Control provides domestic computing power system hardware platforms and customized services for AI companies, research institutions, and system integrators. Its core advantages include:
• Full-Scenario Integrated Platform: Provides 2U/4U AI server chassis platforms, integrating services that can accommodate Ascend accelerator cards, supporting on-demand configuration of CPU, memory, storage, and network.
• Industrial-Grade Reliability: All products meet 7x24 design standards, including redundant power supplies, optimized airflow cooling, out-of-band management, complete system aging tests, and full-load stress tests.
• Deep OEM/ODM Capabilities: Chassis structure, front panel screen printing, BIOS/BMC, port layout, pre-installed drivers, and inference environment images can all be customized, helping integrators quickly build their own branded AI integrated machines.
• Comprehensive Certifications and Project Delivery: This product has passed 3C, CE, FCC, RoHS, and other certifications, providing bilingual specifications and test reports, supporting government and enterprise bidding and overseas project delivery.
• Project-Level Support: Tiered pricing, supply guarantee agreements, and personalized technical support for large-volume purchases are provided, with prototype testing conducted first to reduce selection risk.
Whether you need to deploy a domestic video analytics cluster or plan an enterprise-level large-scale model inference platform, Changfan Industrial Control can provide matching complete hardware and professional selection advice. Welcome to contact Changfan Industrial Control’s sales engineers for solutions and quotations.
Frequently Asked Questions (FAQ)
Q: What is the difference between Ascend NPU and NVIDIA GPU? A: Ascend is an AI computing NPU that works with the CANN+MindSpore full-stack ecosystem, offering compliance advantages in domestic and international innovation and creation scenarios; NVIDIA’s GPU ecosystem is more universal. The choice depends on the business scenario, compliance requirements, and existing software stack.
Q: Can the Atlas 300I Duo run large models? A: It’s suitable for lightweight models under 10 bytes and video analytics; 32-byte models require parallel quantization across multiple cards. For enterprise-level large model inference, the Atlas 800I A2 inference server is recommended.
What is the VRAM size of the Atlas 800I A2? The entire machine is equipped with eight Ascend 910B chips, each with 32GB HBM, totaling 256GB HBM VRAM; the newer Atlas 800I A3 generation has a total HBM capacity of 768GB.
When will the Ascend 950PR be released? A: According to the roadmap announced at Huawei Connect 2025, the Ascend 950PR will launch in Q1 2026, and the 950DT in Q4 2026. Please refer to Huawei’s official release details.
Q: Can Windows be installed on the Ascend server? A: No. The Ascend server is based on the Kunpeng ARM architecture and a Linux system similar to OpenEuler, and applications need to be migrated according to national software stack planning.
Q: Can Changfan Industrial Control provide customized Ascend complete system solutions? A: Yes. We can provide full-process customization from case structure and panel appearance to pre-installed drivers and inference environment, and provide supply guarantees for batch projects.
- Previous article Essential Reading for Enterprise Server Procurement: General-Purpose Servers vs. AI GPU Servers – 10 Dimensions Explaining Differences and Selection Logic | Changfan Industrial Control
- Next article How to Choose Hygon CPUs/DCUs? Introduction to the 3000/5000/7000 Series and the K100-AI Accelerator Card | Changfan Industrial Control
