Is CUDA's Moat Being Filled? DeepSeek Open-Sources TileLang for Ascend, Mirroring NVIDIA's Toolchain Piece by Piece
On September 30, DeepSeek open-sourced six infrastructure components for Huawei Ascend: TileLang for Ascend, DeepGEMM, DeepEP, TileKernels, FlashMLA, DeepSelect — each mirroring a piece of NVIDIA's toolchain. TileLang wraps Ascend C so the same Python APIs run on NVIDIA GPUs and Huawei NPUs. These components already power most operators behind DeepSeek V4 training, MIT licensed. Caveats: numbers were measured on unreleased PoC hardware; some communication primitives are still being optimized.

On September 30, DeepSeek's official WeChat account pushed a message that looked unremarkable but carried enormous weight: it had added Huawei Ascend backend support to all six of the high-performance infrastructure components used in its own training and inference stack — and open-sourced them. The same day, the README of the deepseek-ai/TileKernels GitHub repository gained a line under "News": "[2026-09-30] Huawei Ascend support." A week later, coverage from CNMO and Red Star Capital Bureau pushed the story into the headlines. But beyond the buzz, what deserves a careful read is the technical shape of the whole thing: DeepSeek didn't release a model — it published a "migration dictionary" for Chinese developers. Whatever CUDA can do, Ascend can now do too, with barely any code changes.
What exactly was open-sourced: six components, each mirroring an NVIDIA staple
Let's lay out the inventory first, because it's the skeleton of this whole story. Each of the six components DeepSeek adapted for Ascend has a direct counterpart in the NVIDIA ecosystem:
- TileLang for Ascend — the counterpart of CUDA itself, or more precisely of Triton. TileLang is DeepSeek's open-source domain-specific language for writing kernels: write once, compile and run on multiple hardware backends. This Ascend support wraps Huawei's low-level Ascend C instruction set, so developers can write high-performance kernels without wrestling Ascend C directly.
- DeepGEMM for Ascend (deepseek-ai/DeepGEMM-Ascend) — the counterpart of cuBLAS/cuBLASLt and DeepGEMM itself on NVIDIA. GEMM (general matrix multiplication) is one of the most compute-hungry operations in LLM training; DeepGEMM is DeepSeek's homegrown FP8 high-performance GEMM library. The Ascend version keeps API compatibility with the existing NVIDIA version while supporting BF16, FP8, and FP4 matrix operations.
- DeepEP for Ascend (deepseek-ai/DeepEP-Ascend) — the counterpart of the hardest part of NVIDIA's NCCL. DeepEP is a communication library purpose-built for MoE (mixture-of-experts) architectures, handling expert-parallel All-to-All traffic — "which expert should this token go to." That's the lifeline of training and inference for MoE models like DeepSeek-V3/V4. Its public interface is deliberately aligned with the NVIDIA version, so developers can keep using the call patterns they already know.
- TileKernels (deepseek-ai/TileKernels) — the counterpart of an entire collection of hand-written CUDA kernels. It's a library of dozens of deeply optimized kernels implemented in TileLang, covering the usual suspects in LLM training and inference: MoE routing, FP8/FP4 quantization, RoPE, Engram gating, manifold hyper-connections. The README is blunt about it: most kernels perform close to the hardware's compute or bandwidth ceiling, and every kernel has already been battle-tested in DeepSeek's internal workloads.
- FlashMLA (Ascend port) — the counterpart of the FlashAttention family on NVIDIA. MLA (multi-head latent attention) is the attention innovation DeepSeek has ridden since V2; FlashMLA is its high-performance implementation. With an Ascend version, Ascend cards can now push MLA inference throughput to the max.
- DeepSelect (deepseek-ai/DeepSelect) — the counterpart of sampling/selection-side kernel optimizations, completing the last mile of the inference pipeline.
Once you see this mapping, the ambition is clear: DeepSeek isn't casually "adding support for domestic chips." It took its entire high-performance toolchain — proven on NVIDIA cards — and replicated it wholesale on Ascend. From the language you write kernels in (TileLang), to the most compute-intensive matmuls (DeepGEMM), to the trickiest communication (DeepEP), to a ready-made kernel library (TileKernels), to inference optimization (FlashMLA, DeepSelect) — six pieces, each with an NVIDIA-ecosystem doppelganger. That's why people say it's "filling in CUDA's moat." My take: whether the moat is filled is debatable, but this "dictionary" is thicker than anything we've seen before.
Why TileLang is the real killer move
If you remember only one name out of the six, make it TileLang. It's not another kernel library — it's a language. A "write once, run on many kinds of hardware" DSL for operators.
To grasp its weight, you first have to understand what CUDA's moat actually is. Many people think the moat is that NVIDIA's cards sell well — that's the result, not the cause. The real cause: over the past decade-plus, researchers and engineers worldwide have written mountains of high-performance CUDA code, from cuDNN to countless hand-tuned kernels. That code is an asset, and assets don't travel. Want to switch chips? Fine — but your codebase, your tuning experience, every pitfall you've mapped, all start over. That's what makes CUDA so hard to dislodge: it was never the hardware that was hard to beat, it was the decade of accumulated software ecosystem.
TileLang aims precisely at that weak point. Its bet is to decouple "operator logic" from "hardware backend": developers write kernels in TileLang, and the same Python APIs run on both NVIDIA GPUs and Huawei NPUs, with the backend selected automatically at runtime — that's a near-verbatim quote from the README. It means a team that writes TileLang kernels on NVIDIA cards today can move to Ascend tomorrow, most likely without rewriting. This isn't "compatibility" — it reduces migration cost from "rewrite the code" to "swap the backend."
And the sharper edge: TileLang's Ascend backend wraps Ascend C — Huawei's lowest-level NPU programming interface, with the highest performance ceiling and the steepest learning curve. Mastering Ascend C is no easier than CUDA C was in the early days; people who can do it well are rare. TileLang packages that layer away, dropping the bar for high-performance Ascend programming from "expert-level" to "can write Python." And an ecosystem moat is, after all, built out of countless such bars. DeepSeek just knocked down one of the tallest bricks.
TileKernels: DeepSeek just opened its own toolbox
Back to the most visible repository in this news cycle: TileKernels. Created on April 22 this year, it picked up 1,940 stars in a little over a week after the Ascend support landed on September 30 — for a low-level kernel library, that velocity says it's solving a real pain point.
Three passages in the README deserve a word-by-word read. The first is the feature list: Top-k expert selection and scoring for MoE routing, per-token/per-block/per-channel FP8/FP4 casting and dequantization, fused SwiGLU+quantization ops, Engram gating (forward/backward passes and weight-gradient reduction), RoPE rotary embeddings — all the "performance-sensitive dirty work everyone has to write themselves" in LLM training and inference. The second is the line "Most kernels achieve performance close to the hardware's compute or memory bandwidth limits" — near the hardware roof, meaning this isn't tutorial-grade demo code; it's production-grade. The third is "All of these kernels have already been used in our internal training and inference workloads" — everything has run on real internal workloads.
Attentive readers may have noticed two unfamiliar names in the feature list: Engram and Manifold HyperConnection. They aren't generic terms — they're new architectural components introduced with the DeepSeek V4 series: Engram is a gated memory mechanism, and manifold hyper-connections generalize residual connections. Their presence in this open-source library corroborates the official announcement's claim that these components carry most of the operator implementations behind V4 training. In other words, DeepSeek handed the community the very "weapons" it honed training its latest-generation models. That "arm everyone with the latest gear" move is rare in open-source history: the usual script is to open-source the previous generation and keep the newest as the moat.
It's also worth noting the Acknowledgement section, where DeepSeek explicitly thanks Huawei for "technical support and engineering expertise throughout the development of Tile Kernels' Ascend backend." That's not boilerplate — the Requirements section makes it plain: the Ascend backend needs an Ascend 950 NPU and CANN 9.2.0 or above. That kind of port doesn't happen without deep collaboration from Huawei. The two sides also announced joint deep optimization for a 128-card Ascend 950 supernode. Note that number: a 128-card supernode is aimed squarely at the form factor NVIDIA pioneered with the NVL72 rack — "dozens or hundreds of cards fused into one big computer." The goal was never just "it runs"; it's "it runs at full tilt."
Why a "migration dictionary" matters more than a new model
Now we can answer the key question: why do I say this matters more than releasing a new model?
Because model competition is competition over points; toolchain competition is competition over the whole plane. A new model release holds attention for two weeks; a toolchain, once it spreads, shapes the next five years for everyone building on it. That's exactly how CUDA's ecosystem was built: not on any single card, but on making developers worldwide able to write high-performance code only in CUDA — and then that code locked in hardware choices in return.
What DeepSeek is doing this time is replaying that playbook in reverse: making sure developers don't have to be locked in. Once TileLang plus the six-piece set matures far enough, the way a startup picks hardware changes. The old math was "Ascend cards are cheaper but nobody knows how to tune them — too risky." The new math becomes "same code — pick whichever is cheaper." Hardware differentiation ultimately converges on software-ecosystem competition, and DeepSeek just handed everyone a ticket into that ecosystem, under the MIT license, free for commercial use.
There's also an easily missed angle: why would DeepSeek do this at all? It doesn't sell chips. The most direct explanation is cost and autonomy: its own training increasingly depends on domestic compute, and rather than re-porting everything from scratch each time, it's better to turn the porting layer into an open standard the whole community maintains. The cleverer explanation is strategic: when the entire Chinese-language AI ecosystem runs on "the toolchain DeepSeek defined," DeepSeek becomes the de facto standard-setter — just as Google open-sourced TensorFlow not to sell a framework but to define the rules of the game. DeepSeek skipped the framework entirely and went straight for the operator layer.
A dose of sobriety: the caveats in the README worth stating plainly
A solid report can't only tell the good news. TileKernels' README actually hides a few caveats worth laying out honestly — and interestingly, they make the whole thing more credible, not less.
First, the performance numbers were measured on "futures hardware." According to community write-ups of the official repositories' READMEs, those numbers were produced on a proof-of-concept hardware development kit (PoC HDK — early test hardware Huawei supplied to DeepSeek, not publicly released) plus an undisclosed manual configuration. The README states that full bandwidth requires Huawei's commercial Atlas 850E HDK, officially slated for around mid-October — but that's a vendor roadmap, and the real date may still move. In other words, the "near the hardware ceiling" we've seen so far is a ceiling that's still sitting in futures. Any interpretation stated in absolutes deserves a question mark until measurements on commercial hardware appear.
Second, communication and some operators are still "under construction." The distributed-communication part of the README is admirably direct: with expert-parallel scale (EP — how many cards split the work across different experts at once) at 32 or below, DeepEP's dispatch bandwidth reaches roughly 90–95% of the physical limit; beyond that scale, and for the combine (result-merging) path, the file says "still being optimized." Some collective-communication operations are still in development, and some interfaces remain experimental. The layer that claims to rival CUDA has gaps too: FlashMLA's fused kernel — which fuses Q-norm, RoPE, attention computation, and type conversion into a single call — currently supports only the NVIDIA platform, with no Ascend version yet; DeepSelect's Ascend build only accepts bf16, not fp32. Teams wanting onto this toolchain also need the matching chips plus a CANN 9.2.0-or-newer software stack and firmware — the bar is not low.
Third, the hardware bar is real. The Requirements section is explicit: the Ascend backend needs an Ascend 950 NPU and CANN 9.2.0+. For the vast majority of individual developers, that card simply isn't on their desk — the Ascend 950 is a data-center NPU, not something you buy at retail. So the direct beneficiaries of this release are first cloud vendors, universities, and institutions with Ascend compute, and only indirectly the individual developers who reach it through cloud instances. Hoping to "run it on my laptop tonight" isn't realistic.
Laying out these three caveats isn't about pouring cold water — it's about making the point that this story's value lies not in "today" but in "direction." Get the direction right and the numbers and ecosystem will follow; get it wrong and even the prettiest numbers are a flash in the pan. My judgment is that the direction is right — because "lowering migration cost" is a real, rigid demand that won't go away as long as domestic compute exists.
What it means for vibe coders and indie developers
By now you might ask: I build apps, I don't train models — what does any of this have to do with me? More than you'd think, on two levels.
The first level is inference cost. App developers don't write kernels directly, but every model API you call sits on top of inference cost. As inference-optimization components like FlashMLA and DeepSelect mature, inference services running on Ascend could end up significantly cheaper than NVIDIA-based ones — and those savings are real money. When inference on domestic compute gets cheap enough and stable enough, the call volume indie developers can afford — and the product shapes they can attempt — both expand. Compute democratization eventually reaches everyone who calls an API.
The second level is about who gets "first-hand toolchains." For a long time, Chinese developers in high-performance computing were second-class citizens: new hardware arrived, and the toolchain lagged; when it arrived, the docs were in English and the examples ran on NVIDIA cards. This time it's reversed: DeepSeek's six-piece set got its Ascend support open-sourced in sync with the official announcement, giving Chinese developers first-hand high-performance tooling on domestic compute for the first time — MIT licensed, with a bilingual README no less. That "day-one experience" does as much for ecosystem confidence as any performance number.
More concretely: if you're building an indie product around LLM inference — a local knowledge base, an agent app — it's time to start watching Ascend instance pricing from cloud vendors. When "Ascend inference" stops being a marketing phrase and becomes an option backed by a complete toolchain, the first people to ship it into products and drive costs down will pocket a very real dividend. The window for a technology dividend is never long; those who see it first, eat first.
One question worth chasing: why does Huawei need DeepSeek to do this
Finally, an open question — and to me the most thought-provoking layer of this news cycle: Ascend is Huawei's chip, Ascend C is Huawei's interface, CANN is Huawei's software stack. So why is the "CUDA alternative" for Ascend being built by DeepSeek, not by Huawei itself?
One reading is division of labor: Huawei excels at hardware and low-level software, but the people with the most authority on "what operator libraries does LLM training actually need" are the ones training models every day. DeepSeek happens to be one of the teams in the world that understands MoE training best — it knows where the communication bottlenecks are and where the quantization pitfalls lie. That kind of "demand definition from the front line" is something no hardware team can buy.
The harsher reading is that it shows exactly how hard software ecosystems are to build. Huawei has pushed Ascend for years — MindSpore, CANN, all of it — but whether developers buy in is another matter. An ecosystem needs a killer app to pull it forward: deep learning did it for CUDA back in the day, and large models may do it for Ascend. DeepSeek is that "killer tractor" — bringing its own traffic and its own developers. Huawei provides engineering support; DeepSeek defines the toolchain. It's a textbook "hardware vendor + marquee user" alliance.
But the flip side of the coin: how far this alliance goes depends on whether DeepSeek's investment is a one-off. The thing open-source repositories fear most is "peak at launch," followed by nobody maintaining it. If TileLang's Ascend backend keeps being updated, if the 128-card supernode optimization actually lands, if the community starts contributing third-party kernels — then today's news is a historic starting point for China's AI infrastructure. If not, it's just pretty PR. The commit history will give us the answer; we'll be watching.
There's one more long-term variable: US export controls. Half of Ascend's strategic significance comes from performance; the other half comes from being "the option when you can't buy NVIDIA." This toolchain's value is cost optimization in a world with choices — and a lifeline in a world without them. In either script, its strategic weight only grows, never shrinks — which is why I dare put the word "historic" on it.
One sentence to close: what DeepSeek open-sourced this time isn't code — it's choice. When "CUDA or nothing" becomes "CUDA or Ascend," a moat stops being a moat and becomes just a river. And rivers can be bridged.
Sources
Related articles

On October 6, Cursor's changelog announced Remote Control: the iOS app can now see local agents running on your computer and message them. Compute stays local; only your line of sight moves into your pocket. Within a week of Conductor's iPhone app, two vendors bet on the same interaction — agent programming is shifting from 'terminal dialogue' to 'pocket foreman,' and the workflow is becoming launch-leave-intervene.

Vercel added stealth reasoning model Glyph Cluster to AI Gateway (Oct 7): coding plus long-context analysis, function calling, streaming. Pro/Enterprise teams with purchased credits use it free during stealth (stealth/glyph-cluster), selectable in Claude Code, Codex, Cursor. Why indie developers should benchmark it, the gateway's failover and budget value, and the limits: text-only input, no structured outputs, no ZDR — prompts may train the model, so set a data boundary first.

On Oct 8, Docker's docker CLI plugin Docker Agent hit the HN front page (290 points / 133 comments). It defines agents in declarative YAML (agent.yaml) — 'no code required' — with container-style commands. Model-agnostic (7 providers), native MCP tools, built-in think/todo/memory, full RAG stack, multi-agent delegation. Killer move: agents push/pull to any OCI registry like images, making definitions versionable, PR-reviewable artifacts. Repo dates to Sep 2025: 10,735 commits, 263 releases.