Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

DeepSeek releases Ascend kernels as TileLang adds Huawei support

DeepSeek and TileLang published Ascend 950 support on September 30. The code brings new hardware options and important compatibility limits.

Listen to this article

DeepSeek has released inference components for Huawei’s Ascend platform, while TileLang has added official support for the Ascend 950 processor. The September 30 updates give developers public code and technical documentation for running important parts of DeepSeek’s inference workload on another accelerator family.

The concrete development is in the software. Efficient AI hardware needs compilers, kernels and tested integrations before application teams can make useful comparisons. These releases give engineers more of that foundation to inspect and benchmark.

What DeepSeek actually released

The official FlashMLA repository dates its Ascend attention kernel release to September 30. It includes sparse attention operations for the prompt processing and token generation stages of inference on Huawei Ascend 950 processors.

DeepSeek reports peak results of 410 teraflops during prompt processing and 360 teraflops during token generation, corresponding to 95 percent and 83 percent of theoretical hardware peak in its stated measurements. These are company reported kernel benchmarks. They do not measure the throughput, cost or reliability of an entire production service.

A separate technical explanation describes how the sparse attention implementation selects relevant stored information and organizes computation around the Ascend hardware. Publishing that design gives other engineers a way to examine the work behind the headline performance figures.

TileLang adds the compiler path

TileLang’s own release record also marks September 30 as the arrival of its Ascend 950 backend. The language lets developers express performance sensitive kernels using Python style syntax, with compiler machinery handling lower level execution details.

The new backend provides native code generation, scheduling and synchronization for Ascend. Its setup documentation requires compatible Huawei drivers and the CANN toolkit, along with the relevant PyTorch integration. A build configured for Ascend can disable CUDA, but it still depends on a hardware specific software stack.

That is the distinction to keep in view when comparing accelerator ecosystems. A higher level language can reduce some of the work needed to target different hardware. It does not remove the need to check supported operations, installation requirements and the behavior of the complete application.

Existing users should read the compatibility notice

FlashMLA’s September 30 release also removes support for NVIDIA Hopper and earlier DeepSeek model generations, and changes its key and value cache format. The maintainers direct users who need the older configuration to a previous revision. Upgrading an existing deployment therefore requires more care than replacing a package version.

A useful evaluation should pin the software revision and record the model, hardware, numerical format and workload. Measure time to the first token, sustained throughput, output correctness and failure recovery under realistic load. A fast kernel can help a service while leaving bottlenecks elsewhere in the system.

The release is a meaningful addition to the public software available for Ascend. The next question is how the components behave when independent teams assemble them into dependable serving systems, rather than whether one benchmark settles the competition between entire hardware platforms.

DeepSeek logo from LobeHub Icons, used under its MIT license.

Marcus Reid
Marcus Reid

Marcus Reid is focused on covering the money, rules, and institutional choices shaping AI. He runs from funding rounds and chip deals to regulation, lawsuits, leadership changes, and the business of building enormous computing systems. Marcus follows the incentives behind the announcement. Who pays, who gains leverage, and what changes for everyone else? The voice is direct, measured, and occasionally dry, especially when a grand promise arrives with very little detail.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile