The AI inference market is entering a genre shift. On December 18, 2024, Moore Threads co-founder Wang Dong made a statement that should echo across both Silicon Valley and crypto-native compute networks: “There is no universal chip for inference. What we need is a combination of solutions.”
This isn't a technical footnote. It's a pivot point—one that aligns perfectly with the structural logic underpinning decentralized physical infrastructure networks (DePIN) and tokenized compute marketplaces. If Wang Dong is correct, the monolithic GPU narrative that has dominated the last two years is about to fracture, creating a vacuum that crypto infrastructure can fill.
Context: The Fragmentation Reality
The inference market is not training. Training demands brute-force matrix multiplication on massive clusters—NVIDIA’s natural monopoly. Inference, however, is a fragmented beast: low-latency chatbots, high-throughput batch generation, streaming code completion, video synthesis. Each scenario optimizes for a different performance vector—latency, throughput, memory bandwidth, power efficiency.
Today, NVIDIA’s H100/B200 still commands ~90% of AI workloads. But beneath the surface, the market is already splintering. Groq’s LPU excels at low-latency text generation. Cerebras’ wafer-scale chip dominates sparse inference. AMD’s MI300X offers competitive throughput for batch inference. And dozens of Chinese players—including Moore Threads, Huawei, and Cambricon—are building chips tailored to specific constraints of the domestic ecosystem: lower cost, easier compliance, and political preference for homegrown hardware.
Wang Dong’s thesis—“no universal chip”—is not novel in isolation. But his framing as a combination of solutions is. It captures the emerging engineering reality: inference is becoming a multi-architecture problem, not a single-vendor solution.
Core: The Incentive Structure Behind the ‘Combination’ Narrative
From my experience auditing 50+ ICO whitepapers during the 2017 frenzy, I learned one iron law: every market narrative is underwritten by an incentive structure. Wang Dong’s argument is no exception.
Moore Threads is a latecomer to a GPU market dominated by two giants (NVIDIA and, in China, Huawei). Their MTT S4000 series is competitive on paper—7nm process, solid FP16 throughput—but it lacks the software ecosystem maturity of CUDA or the government-blessed vertical integration of Ascend. A “combination of solutions” narrative does three things for them:
- Shrinks the competition into a component of a larger system. By saying no single chip is universal, Wang Dong implicitly relegates NVIDIA’s dominance to just one node in a heterogeneous cluster. This opens the door for Moore Threads to be the “best-in-slot” for a specific scenario—say, cost-sensitive edge inference with custom quantized models.
- Positions software orchestration as the real moat. “Through software-hardware co-optimization, every model can find its optimal hardware combination.” This is code for: the value capture is shifting from the chip itself to the compiler, scheduler, and abstraction layer that stitches multiple chips together. This is exactly the layer where decentralized compute networks (Render Network, Akash, io.net) are building their own moats—using token incentives to aggregate fragmented hardware supply.
- Reframes cost advantage as a structural phenomenon. Wang Dong’s claim that “Chinese frontier base models have a cost advantage” is really a statement about hardware arbitrage. Chinese labs are forced to use less advanced chips, so they innovate via aggressive quantization, sparsity, and distillation. The result: cheaper inference per token, even if not per floating-point operation.
This last point resonates with the crypto-native compute thesis. Decentralized GPU networks have long argued that the marginal cost of inference in a permissionless, globally distributed hardware pool is structurally lower than centralized cloud GPU. Wang Dong’s China-specific hardware arbitrage is a microcosm of that same principle: forced heterogeneity leads to lower cost.
Decoding the signal from the narrative noise—the real insight is not about chips. It’s about the emergence of a composable hardware orchestration layer that can dynamically route inference requests to the cheapest, fastest, or most compliant hardware. That abstraction layer is where next-generation value will accrue.
Contrarian: The Bear Case for the ‘Combination’ Thesis
Every narrative has a blind spot. For Wang Dong’s combination-of-solutions thesis, it’s threefold:
First, the engineering complexity is understated. Running the same model across different architectures—NVIDIA, AMD, Moore Threads—requires exhaustive operator library support, consistent numerical precision, and unified memory management. Even with tools like Triton Inference Server, achieving production-grade performance uniformity is non-trivial. The cost of integration may outweigh the hardware savings for all but the largest AI workloads.

Second, the ISP (Inference Service Provider) model faces a trust issue. Wang Dong predicts a wave of independent ISPs that orchestrate multi-vendor hardware. But the largest inference workloads today are still run by hyperscalers (AWS, Azure, GCP) or model providers (OpenAI, Anthropic). These incumbents have no incentive to fragment their supply chain. An independent ISP would need massive upfront capital to provision heterogeneous compute, and then must compete on price against integrated stacks. The margins are razor-thin.
Third, Moore Threads’ own position is precarious. Wang Dong’s pitch is a survival strategy, not a victory lap. If NVIDIA releases a next-gen chip that achieves 2x price-performance improvement while maintaining software dominance, the “combination” argument loses its urgency. Moreover, Moore Threads has not disclosed a single major ISP partner or large customer deployment. The lack of hard evidence suggests the pitch is still aspirational.
Unearthing the logic within the speculative fog—the crypto analog is sobering. Many DePIN compute networks have promised “the world’s largest distributed GPU cluster” but have yet to demonstrate consistent, low-latency inference at scale for demanding workloads. The combination thesis is seductive, but execution is everything.
Takeaway: What This Means for Crypto AI
Wang Dong’s statements, while not about crypto, validate a core narrative that has been building in the intersection of blockchain and AI. The thesis that inference hardware is becoming a heterogeneous, orchestrated market is exactly what decentralized compute networks are designed to serve.
If the combination-of-solutions becomes the dominant paradigm, then the infrastructure layer that abstracts hardware heterogeneity—rather than the hardware itself—becomes the primary value capture point. That abstraction layer could be built on tokenized incentives, open-source order books, or on-chain reputation systems. The projects that solve this orchestration problem will be the ones that capture the next narrative cycle.
Building frameworks for the next narrative cycle—the question for investors and builders is not whether Moore Threads or any single chipmaker wins. It’s: who builds the layer that lets any model run on any chip, with optimal cost and latency? That layer is the true bottleneck. And it might not be built by a hardware company at all.
Based on my experience tracking DeFi Summer liquidity and NFT genre shifts, I’ve learned that the most valuable infrastructure is the one that connects fragmented supply to fragmented demand—without owning either side. Wang Dong’s speech just provided the clearest signal yet that inference demand is about to splinter. The question is whether crypto’s supply side is ready to assemble the pieces.
