Startups Are Strongest for AI Infra… AI recommendations | Parse
Which startups are strongest for AI infrastructure that reduces GPU dependence?
Data as of Sep 24, 2026 · Based on 50 AI responses from ChatGPT Search and Google AI Mode · See how Parse measures this
Sources AI cites for this prompt
Source
Citation share
Breaking the GPU stronghold: emerging competition in AI infrastructurehttps://www.kearney.com/industry/technology/article/breaking-the-gpu-stronghold-emerging-competition-in-ai-infrastructure
2%
Top 30+ AI Chip Makers: NVIDIA & Its Competitors - AIMultiplehttps://aimultiple.com/ai-chip-makers
1%
The Inference Economy Arrives: AI Chip Rules Are Being Rewrittenhttps://insights.trendforce.com/p/ai-inference-chip-architecture
1%
GPU efficiency: which startup is ahead? - New Market Pitchhttps://newmarketpitch.com/blogs/news/ai-infrastructure-gpu-efficiency-startup
1%
Groq, Cerebras, Tenstorrent: The Inference-Chip Challengers | SiliconReporthttps://www.siliconreport.com/inference-chip-startups-groq-cerebras-tenstorrent-51b3c733
1%
[2604.10852] The xPU-athalon: Quantifying the Competition of AI Accelerationhttps://arxiv.org/abs/2604.10852
1%
AI Infrastructure: what are the top startups now? - New Market Pitchhttps://newmarketpitch.com/blogs/news/ai-infrastructure-top-startups
1%
These 70 Companies Are Building the Core Infrastructure for Artificial Intelligence (AI)https://probabilitylens.substack.com/p/these-70-companies-are-building-the
1%
Has the hunt for AI compute uncovered the next Cerebras? | TechCrunchhttps://techcrunch.com/2026/05/28/has-the-hunt-for-ai-compute-uncovered-the-next-cerebras/
As the AI industry hits severe power , copper-interconnect, and cost bottlenecks associated with scaling traditional GPU clusters, a powerful class of infrastructure and semiconductor startups is emerging to bypass Nvidia dependency. These companies focus heavily on frontier inference optimization, wafer-scale integration , and domain-specific architectures (ASICs) that process tokens drastically faster and at a fraction of the power budget.
The strongest and most heavily capitalized players driving this shift include:
Etched : Easily one of the most explosive names in the non-GPU landscape, Etched builds application-specific processors hardcoded exclusively for transformer models, completely dropping support for general-purpose computing to maximize output. By abandoning flexibility for raw execution speed, their chips deliver massive leaps in tokens-per-watt and tokens-per-dollar. They secured massive market validation and multi-billion dollar valuations backed by heavyweights like Jane Street and Kleiner Perkins.
Fractile : Operating out of the UK, this startup interweaves memory and compute directly on the same silicon die to eliminate the data-transfer bottlenecks that plague standard GPU setups. Fractile has generated major momentum—including major supply discussions and deals with frontier labs like Anthropic—sending its valuation soaring into the billions as it targets extreme inference acceleration.
Cerebras Systems : Moving past the traditional notion of discrete chips, Cerebras builds massive Wafer-Scale Engines (the size of an entire dinner plate) that integrate compute and massive on-chip memory together. Having completed a major public listing on Nasdaq under the ticker CBRS , Cerebras provides an alternative architecture capable of training and running giant models locally without clustering thousands of individual tiny GPUs.
2. Alternative Architecture & Memory Orchestration
SambaNova Systems : SambaNova utilizes Reconfigurable Dataflow Units (RDUs) rather than fixed-instruction GPU cores. Their architecture is engineered with multi-tiered memory systems designed to handle trillion-parameter enterprise generative models efficiently within smaller footprints, reducing the sprawling infrastructure requirements typical of standard GPU data centers.
Lightmatter : Instead of relying purely on electronic transistors and copper wires which heat up and bottleneck at scale, Lightmatter uses light (photons) to perform the heavy matrix math required by neural networks. This drastically slashes power consumption and latency, providing a systemic infrastructure layer that detours around traditional electronic scaling limits.
If you're exploring these for an upcoming deployment or investment thesis, let me know:
Are you focusing on training or inference workloads?
Do you need an alternative for on-premise hardware or cloud-accessible APIs?
If by “reduces GPU dependence” you mean infrastructure that can replace or materially reduce Nvidia GPUs for inference, serving, or memory-heavy AI workloads, the most interesting startup/scale-up set in 2026 is:
sambanova.ai — Probably the most directly relevant private company. Its dataflow architecture is designed specifically around inference, with its SN50 targeting existing data centers and lower power/cooling requirements than conventional GPU systems.
d-matrix.ai — Particularly interesting because its compute-in-memory approach attacks the memory/bandwidth bottleneck of LLM inference rather than simply trying to build another GPU. Its architecture can also be deployed alongside GPUs, potentially making adoption easier.
etched.com — A much more aggressive ASIC bet. Its Sohu chip is designed specifically for Transformer inference. The tradeoff is flexibility: specialized silicon can be extremely efficient, but architectural changes in models can make the bet riskier.
taalas.com — Perhaps the most radical efficiency approach: hard-wiring model weights into silicon. Taalas reported extremely high per-user inference throughput on Llama 3.1 8B, illustrating how far specialized hardware can move from the GPU model. The catch is that the hardware is correspondingly less general-purpose.
axelera.ai — Worth watching for edge and enterprise inference rather than frontier-model training. Its second-generation Europa chip launched in September 2026, with Dell and Supermicro integration and more than 600 customers reported by the company.
tenstorrent.com — A broader accelerator play rather than a pure inference ASIC. The appeal is a more open alternative accelerator/software ecosystem, potentially reducing dependence on both Nvidia silicon and CUDA.
One important caveat: Cerebras and Groq
Cerebras is arguably the most commercially validated example of the thesis, but it is no longer a startup in the conventional sense after going public in 2026. Its wafer-scale architecture is aimed heavily at inference and has demonstrated very high token throughput.
Groq would historically have been near the top of this list, but its situation changed dramatically: Nvidia obtained its inference technology and much of its core team in a $20 billion transaction in late 2025. Groq subsequently shifted toward an AI-cloud business, so I wouldn't treat it today as an independent GPU alternative in the same sense as SambaNova or d-Matrix.
How I'd segment the opportunity
Company
Main GPU-reduction thesis
Best fit
SambaNova
Dataflow accelerator
Enterprise / cloud inference
d-Matrix
Compute-in-memory
High-volume LLM inference
Etched
Transformer-specific ASIC
Very high-volume inference
The broader technical thesis is important: GPU displacement is much more plausible in inference than in frontier-model training. Inference is increasingly constrained by memory movement, KV-cache capacity/bandwidth, latency and power rather than simply raw matrix-compute throughput. Recent accelerator research also finds that the optimal architecture varies substantially with batch size, sequence length and model size—so there probably won't be one universal “GPU killer.”
If you're looking at these as potential startup investments, rather than as technology vendors, I can also break them down by funding/valuation, customer traction, technical moat, TAM, and likely exit paths using current 2026 data.
For reducing dependence on Nvidia GPUs, the strongest startup/scale-up names to investigate are:
Cerebras — wafer-scale accelerators; particularly differentiated for high-throughput inference.
d-Matrix — inference-focused chips using in-memory compute; its next-generation Raptor is being designed to integrate with Nvidia systems via NVLink Fusion.
SambaNova — reconfigurable dataflow processors with a more flexible memory architecture for inference.
Positron — memory-first inference processors; it recently raised $875M at a $5B valuation, signaling substantial commercial momentum.
Axelera AI — inference accelerators targeting enterprise/edge deployments, with its second-generation Europa chip now entering the market.
Important:Groq belongs in the historical comparison, but Nvidia acquired/licensed its core chip technology in 2025; Groq is now primarily operating as an inference-cloud company rather than an independent chip challenger.
If you're evaluating these as venture investments, the criteria differ substantially from evaluating them as technology suppliers.
As the market shifts from raw, single-vendor hardware accumulation to software-defined compute abstraction , a specialized tier of AI infrastructure startups has emerged. Rather than building proprietary silicon to compete with hardware giants, these companies focus on cross-chip portability, orchestration layers, and inference efficiency —allowing enterprises to run workloads seamlessly across non-NVIDIA accelerators (like Google TPUs , AMD Instinct, and Amazon Trainium ) or drastically reduce the raw amount of compute needed.
The strongest and most strategically relevant startups addressing GPU dependence and compute optimization include:
1. Cross-Chip Portability & Abstraction Layers
Decart: Leading the abstraction layer by engineering optimization stacks that actively support multi-hardware environments, spanning NVIDIA GPUs alongside alternative silicon like Google TPUs and Amazon Trainium . By focusing on algorithmic efficiency (such as radical cost and token-generation optimization in generative and real-time video models), Decart decouples heavy workloads from traditional hardware constraints.
FlexAI: Builds a software-defined cloud infrastructure and unified control plane designed to let enterprises deploy AI inference and fine-tuning across heterogeneous fleets—operating across mixed NVIDIA and AMD environments without locking the developer into a single hardware ecosystem . It targets the exact friction point of hardware scarcity and vendor lock-in.
2. Inference Efficiency & Workload Optimization
FriendliAI: Noted for strong production backing and benchmarks in inference serving . By maximizing the efficiency of how models execute at the serving layer, FriendliAI reduces the overall volume of concurrent GPU hours needed to serve high-throughput commercial LLM applications.
Groq / Rebellions: While technically operating full-stack hardware-software alternatives rather than software-only abstractions, companies like Groq (with their LPU architecture) and South Korea's Rebellions provide specialized inference silicon designed specifically to bypass GPU bottlenecks for high-speed, low-latency token generation.
Would you like to explore how cross-chip abstraction layers handle performance trade-offs compared to native NVIDIA CUDA deployments?
Are you looking at these startups from a cost-optimization/inference perspective or for large-scale training portability?
Cast AI and ScaleOps are currently the strongest infrastructure startups for maximizing efficiency and reducing raw GPU waste, while alternative hardware-software stack builders like Decart, FriendliAI , and specialized chip-adjacent innovators like d-Matrix and SambaNova Systems lead the push toward alternative accelerators.
The market divides cleanly into orchestration/efficiency platforms that squeeze more out of existing silicon and alternative architectures designed to bypass traditional NVIDIA bottlenecks.
Top AI Infrastructure Startups for GPU Optimization and Independence
Cast AI : Leads the infrastructure-level efficiency race with an automated Kubernetes-native platform that optimizes and dynamically right-sizes GPU and CPU clusters, minimizing idle compute waste across cloud environments.
ScaleOps : Closely follows as a real-time automated resource management platform that continuously tunes pod resource allocations for AI workloads, drastically reducing the physical footprint and number of active GPUs required in production.
Decart : Focuses closely on the runtime layer, optimizing workloads to run efficiently and fluidly across diverse non-NVIDIA chip families and heterogeneous hardware.
If you want to narrow this down, let me know:
Are you looking to optimize cloud costs for existing models or explore alternative chips?
Are your workloads focused on training or inference?
The strongest GPU-reduction bets today are d-Matrix, Etched, SambaNova, and Positron; Tenstorrent is the strongest open-architecture play. The opportunity is mainly inference, where specialized silicon can beat GPUs on cost, latency, and power.
d-Matrix — strongest overall specialized-inference bet; compute-in-memory architecture and now integrating with NVIDIA systems via NVLink Fusion.
If by “reduces GPU dependence” you mean startups building infrastructure that can materially replace NVIDIA GPUs—not merely make GPU usage more efficient—the strongest group in 2026 is fairly concentrated.
My shortlist
Startup
What they’re betting on
My take
Cerebras Systems
Wafer-scale inference/AI compute
Strongest overall
SambaNova Systems
Reconfigurable dataflow processors + inference systems
Strongest flexible alternative
Tenstorrent
Open RISC-V + AI accelerators
Most interesting platform/ecosystem bet
d-Matrix
Memory-centric inference silicon
Very interesting architecture
Etched
Transformer-specific ASIC
Highest-risk/high-upside
Lightmatter
Optical interconnect/compute
Longer-term infrastructure bet
Taalas
Hard-wired model inference
Radical cost/efficiency play
One important caveat: Groq would have been near the top, but its situation has changed substantially—NVIDIA's $20B transaction around Groq's technology makes it less compelling as an independent startup investment thesis.
1. Cerebras — clearest current winner
Cerebras is probably the strongest answer if you're asking “who has actually demonstrated that GPUs don't have to be the default for inference?”
Its wafer-scale architecture attacks the problem very differently from NVIDIA: enormous on-chip SRAM and a wafer-scale processor rather than thousands of conventional GPU processors. The company has moved beyond the pure startup phase with its 2026 IPO, and recent industry comparisons put it substantially ahead of private inference-chip competitors on commercial maturity.
Why I like it: the architecture is genuinely differentiated, and inference is increasingly a memory/data-movement problem rather than simply a FLOPS problem.
2. SambaNova — perhaps the most strategically interesting
SambaNova's approach is less extreme than Cerebras: its reconfigurable dataflow architecture combines substantial on-chip memory with HBM/DDR, making it better suited to large and changing models.
That matters because the winning non-GPU architecture probably can't be too specialized. Enterprises want to run many models, long contexts and increasingly agentic workloads. Recent reporting suggests SambaNova's upcoming systems are specifically targeting that flexibility.
My view: if Cerebras is the strongest proof point, SambaNova may be the more interesting “general-purpose replacement for GPU inference” thesis.
Tenstorrent is interesting for a different reason: it's trying to build a broader computing platform around AI accelerators, RISC-V CPUs and an increasingly open software stack.
That gives it a potentially much larger addressable market than a chip optimized solely for LLM inference. The tradeoff is that it has to execute across silicon, compiler, software and ecosystem simultaneously.
Investment thesis: less “beat NVIDIA at tokens/sec” and more “create an alternative AI-compute stack.”
4. d-Matrix — one I'd watch closely
d-Matrix is attacking what may ultimately be the most important bottleneck: moving model data and KV-cache around efficiently.
That's particularly attractive because decode is increasingly memory-bandwidth constrained. The academic literature is converging on the idea that the “memory wall” is one of the central problems in LLM inference.
So I would put d-Matrix in the high-potential architecture category rather than the proven-winner category.
5. Etched — spectacularly asymmetric bet
Etched's thesis is basically:
Transformers are important enough that you can build an ASIC specifically for them.
If that bet works, the economics could be extraordinary. But specialization creates a corresponding risk: architecture shifts, new model families and increasingly heterogeneous workloads can make a beautifully optimized ASIC obsolete.
I'd classify Etched as venture-style asymmetric upside, rather than the safest current winner.
6. Taalas — the wild card
Taalas takes specialization even further by essentially hard-wiring models into silicon. Its HC1 launch demonstrated the extreme version of the thesis: enormous inference throughput from a chip designed around a specific model rather than a general-purpose GPU.
If AI inference becomes dominated by a relatively small number of stable, high-volume models, this could be extremely powerful.
If models continue changing rapidly, it's much harder.
But there's another category I'd take very seriously
GPU dependence can be reduced without replacing GPUs.
Companies such as Baseten, Together AI, and Fireworks AI attack the problem at the software/orchestration layer: better batching, quantization, model routing, utilization and heterogeneous infrastructure.
Baseten, for example, reports up to 225% better price-performance on certain model-serving workloads through optimization.
That's important because you don't necessarily need to eliminate the GPU to eliminate NVIDIA's economic leverage. Getting 2–4× more useful inference out of the same accelerator can be economically equivalent to adding a competing accelerator.
And the emerging inference-gateway layer—routing workloads among different providers rather than owning the silicon—is another way to reduce dependence on any one GPU supplier.
My ranking by investment thesis
If I were building a 2026 venture watchlist, I'd roughly rank them:
Etched — huge upside, much higher execution/technology risk
Taalas — fascinating extreme-specialization bet
Lightmatter — potentially enormous long-term opportunity, but further from mainstream replacement
The bigger point is that I wouldn't frame this as “who will replace NVIDIA?” The more interesting question is which layer of the AI stack captures the economics as inference explodes.
Today, the strongest thesis is probably specialized inference hardware + memory optimization + software orchestration, rather than a single magical “GPU killer.” Recent research explicitly points toward heterogeneous, memory-centric systems rather than one universal accelerator.
If you're asking this from an investment perspective, I can also rank these by funding/valuation, revenue traction, technical moat, likelihood of IPO/acquisition, and 5–10× upside.
Its wafer-scale architecture attacks the problem very differently from NVIDIA: enormous on-chip SRAM and a wafer-scale processor rather than thousands of conventional GPU processors. The company has moved beyond the pure startup phase with its 2026 IPO, and recent industry comparisons put it substantially ahead of private inference-chip competitors on commercial maturity.
Why I like it: the architecture is genuinely differentiated, and inference is increasingly a memory/data-movement problem rather than simply a FLOPS problem.
2. SambaNova — perhaps the most strategically interesting
SambaNova's approach is less extreme than Cerebras: its reconfigurable dataflow architecture combines substantial on-chip memory with HBM/DDR, making it better suited to large and changing models.
That matters because the winning non-GPU architecture probably can't be too specialized. Enterprises want to run many models, long contexts and increasingly agentic workloads. Recent reporting suggests SambaNova's upcoming systems are specifically targeting that flexibility.
My view: if Cerebras is the strongest proof point, SambaNova may be the more interesting “general-purpose replacement for GPU inference” thesis.
Tenstorrent is interesting for a different reason: it's trying to build a broader computing platform around AI accelerators, RISC-V CPUs and an increasingly open software stack.
That gives it a potentially much larger addressable market than a chip optimized solely for LLM inference. The tradeoff is that it has to execute across silicon, compiler, software and ecosystem simultaneously.
Investment thesis: less “beat NVIDIA at tokens/sec” and more “create an alternative AI-compute stack.”
4. d-Matrix — one I'd watch closely
d-Matrix is attacking what may ultimately be the most important bottleneck: moving model data and KV-cache around efficiently.
That's particularly attractive because decode is increasingly memory-bandwidth constrained. The academic literature is converging on the idea that the “memory wall” is one of the central problems in LLM inference.
So I would put d-Matrix in the high-potential architecture category rather than the proven-winner category.
5. Etched — spectacularly asymmetric bet
Etched's thesis is basically:
Transformers are important enough that you can build an ASIC specifically for them.
If that bet works, the economics could be extraordinary. But specialization creates a corresponding risk: architecture shifts, new model families and increasingly heterogeneous workloads can make a beautifully optimized ASIC obsolete.
I'd classify Etched as venture-style asymmetric upside, rather than the safest current winner.
6. Taalas — the wild card
Taalas takes specialization even further by essentially hard-wiring models into silicon. Its HC1 launch demonstrated the extreme version of the thesis: enormous inference throughput from a chip designed around a specific model rather than a general-purpose GPU.
If AI inference becomes dominated by a relatively small number of stable, high-volume models, this could be extremely powerful.
If models continue changing rapidly, it's much harder.
But there's another category I'd take very seriously
GPU dependence can be reduced without replacing GPUs.
Companies such as Baseten, Together AI, and Fireworks AI attack the problem at the software/orchestration layer: better batching, quantization, model routing, utilization and heterogeneous infrastructure.
Baseten, for example, reports up to 225% better price-performance on certain model-serving workloads through optimization.
That's important because you don't necessarily need to eliminate the GPU to eliminate NVIDIA's economic leverage. Getting 2–4× more useful inference out of the same accelerator can be economically equivalent to adding a competing accelerator.
And the emerging inference-gateway layer—routing workloads among different providers rather than owning the silicon—is another way to reduce dependence on any one GPU supplier.
As the AI industry hits physical power, cooling, and supply constraints tied to traditional graphics processors, the focus of AI infrastructure has shifted heavily toward alternative silicon architectures, specialized application-specific integrated circuits (ASICs), and software layers that abstract away hardware lock-in.
The strongest startups and alternative hardware players reducing or bypassing pure Nvidia GPU dependence include:
Cerebras Systems : Operating on a radically different scale, Cerebras builds massive wafer-scale processors (like the CS-3) rather than clustering thousands of smaller discrete GPUs. By integrating compute and memory onto a single giant chip, they dramatically reduce inter-chip communication bottlenecks, positioning themselves as high-throughput, cost-efficient alternatives for both training and dense inference.
Groq : Known for its deterministic Language Processing Unit (LPU) architecture, Groq bypasses traditional memory hierarchies by relying strictly on on-chip SRAM and pre-planned execution schedules via an advanced compiler. This design choice delivers ultra-low latency and massive token-per-second throughput for inference workloads, sidestepping the energy bloat and memory-bandwidth walls typical of standard GPU farms. (Note: Nvidia has moved to absorb and integrate aspects of Groq's technology and IP through recent strategic maneuvers, highlighting the threat the architecture posed to the GPU monopoly).
d-Matrix : Focused specifically on the generative AI inference bottleneck rather than heavy model training, d-Matrix builds digital in-memory computing (DIMC) architecture. Backed by heavyweights like Microsoft, their ASICs are tailored to make running transformer models radically more power-efficient and cost-effective once deployment scale outpaces training.
TensorWave : Operating more on the cloud infrastructure and deployment side rather than raw chip fabrication, TensorWave has positioned itself as a primary alternative by building high-performance cloud clusters explicitly around AMD's MI300X accelerators instead of defaulting to Nvidia. They act as a crucial supply-chain safety valve for enterprise buyers looking to de-risk their hardware dependencies.
If you'd like to narrow this down, let me know:
Are you looking for hardware alternatives for training or inference?
Are you interested in custom silicon/chipmakers or alternative cloud providers?
If by “reduces GPU dependence” you mean startups building infrastructure that can make AI inference cheaper/faster without relying entirely on NVIDIA GPUs, the strongest group is surprisingly concentrated.
My 2026 shortlist
Startup
What it does
GPU-displacement potential
My take
Cerebras Systems
Wafer-scale AI processors + systems/cloud
★★★★★
Strongest overall
Groq
Specialized LPU inference architecture/cloud
★★★★★
Exceptional inference economics
SambaNova Systems
Dataflow AI accelerators + inference systems
★★★★☆
Strong enterprise alternative
d-Matrix
Compute-in-memory inference chips
★★★★☆
Particularly interesting for token economics
Etched
Transformer-specific ASIC
★★★★☆
Very high upside, higher model-risk
Tenstorrent
RISC-V + AI accelerators
★★★½☆
Broader platform bet
Lightmatter
Photonic interconnect/computing
★★★½☆
More fundamental infrastructure bet
FuriosaAI
Efficient inference accelerators
★★★☆☆
Interesting efficiency play
The big strategic point is that inference is becoming the battleground, rather than training. Specialized accelerators attack the parts of inference where GPUs can be inefficient—memory movement, bandwidth, latency and token-by-token generation.
1. Cerebras — strongest overall
Cerebras has arguably the most differentiated architecture: instead of assembling enormous numbers of relatively small accelerators, its wafer-scale engine puts an enormous amount of compute and memory onto one piece of silicon.
It's also crossed an important commercialization threshold. Cerebras raised $1B in 2026 and subsequently went public, while signing a $10B multi-year compute agreement with OpenAI.
Why I like it: it isn't merely selling a theoretical accelerator; it's becoming an alternative compute platform.
Risk: wafer-scale manufacturing and the economics of its architecture are considerably more specialized than NVIDIA's ecosystem.
2. Groq — best pure inference story
Groq's architecture is almost the opposite philosophy: build silicon specifically around extremely predictable, low-latency inference.
There's an important wrinkle, though: NVIDIA acquired/licensed the core Groq technology in a $20B transaction and has incorporated it into its own inference roadmap. Groq itself continues as an inference-cloud company.
That makes Groq fascinating but less clean as a venture bet against NVIDIA.
Investment thesis: specialized inference silicon is so valuable that even NVIDIA decided it was worth owning.
3. SambaNova — strongest enterprise-oriented alternative
SambaNova's dataflow architecture is designed to handle large models while using a hierarchy of on-chip memory, HBM and external memory.
Its appeal is less “absolute maximum tokens/sec” and more flexibility + memory capacity + enterprise deployment. Its newer architecture is explicitly aimed at addressing some of the limitations of earlier inference accelerators.
I'd put it ahead of many newer chip startups because it has spent years building the complete hardware/software/system stack.
4. d-Matrix — one of the most interesting private bets
d-Matrix is particularly interesting because it attacks the memory problem using digital compute-in-memory.
That is potentially powerful because inference increasingly looks like a data-movement problem rather than simply a FLOPS problem. Its Corsair architecture is designed to keep computation close to the data rather than repeatedly shuttling weights around.
If you're specifically looking for “who could reduce the number of GPUs required per million tokens?”, d-Matrix deserves serious attention.
5. Etched — highest-risk/highest-upside architecture bet
Etched takes an unusually aggressive position: build an ASIC optimized around the Transformer architecture rather than trying to be a general-purpose AI accelerator.
The upside is enormous efficiency if Transformers remain dominant. The downside is equally obvious: AI architectures can change.
I'd therefore rank Etched below Cerebras/Groq/SambaNova on commercial maturity but potentially above them on asymmetric upside.
The broader opportunity
I'd actually divide the market into four ways of reducing GPU dependence:
Alternative silicon — Cerebras, Groq, SambaNova, d-Matrix, Etched.
Better memory/interconnect — Lightmatter and related photonics companies.
Inference software/orchestration — systems that extract much more utilization from each GPU.
Distributed/alternative compute — using idle or geographically distributed hardware. New startups are experimenting with this, although economics and reliability remain unproven.
The last category is much less compelling to me than specialized silicon. The current evidence increasingly suggests that architectural specialization, rather than simply finding more GPUs somewhere else, is where the biggest structural savings may come from.
If I were ranking the investment theses
Best combination of technology + traction: Cerebras
Best proof that specialized inference works: Groq
Best enterprise alternative: SambaNova
Most interesting memory-centric architecture: d-Matrix
Most asymmetric bet: Etched
Most interesting longer-term infrastructure bet: Lightmatter
One important caveat: “GPU dependence” isn't disappearing. NVIDIA still has an enormous software/ecosystem advantage, and the emerging architecture may actually be heterogeneous—GPUs for some workloads and specialized accelerators for others. Recent research on agentic inference, for example, points toward disaggregating prefill/decode and using heterogeneous hardware rather than replacing GPUs wholesale.
If your question is specifically for venture investing, I can also rank these by funding/valuation, customers, technical moat, revenue traction, and probability of becoming a $10B+ company.
Risk: wafer-scale manufacturing and the economics of its architecture are considerably more specialized than NVIDIA's ecosystem.
2. Groq — best pure inference story
Groq's architecture is almost the opposite philosophy: build silicon specifically around extremely predictable, low-latency inference.
There's an important wrinkle, though: NVIDIA acquired/licensed the core Groq technology in a $20B transaction and has incorporated it into its own inference roadmap. Groq itself continues as an inference-cloud company.
That makes Groq fascinating but less clean as a venture bet against NVIDIA.
Investment thesis: specialized inference silicon is so valuable that even NVIDIA decided it was worth owning.
3. SambaNova — strongest enterprise-oriented alternative
SambaNova's dataflow architecture is designed to handle large models while using a hierarchy of on-chip memory, HBM and external memory.
Its appeal is less “absolute maximum tokens/sec” and more flexibility + memory capacity + enterprise deployment. Its newer architecture is explicitly aimed at addressing some of the limitations of earlier inference accelerators.
I'd put it ahead of many newer chip startups because it has spent years building the complete hardware/software/system stack.
4. d-Matrix — one of the most interesting private bets
d-Matrix is particularly interesting because it attacks the memory problem using digital compute-in-memory.
That is potentially powerful because inference increasingly looks like a data-movement problem rather than simply a FLOPS problem. Its Corsair architecture is designed to keep computation close to the data rather than repeatedly shuttling weights around.
If you're specifically looking for “who could reduce the number of GPUs required per million tokens?”, d-Matrix deserves serious attention.
5. Etched — highest-risk/highest-upside architecture bet
Etched takes an unusually aggressive position: build an ASIC optimized around the Transformer architecture rather than trying to be a general-purpose AI accelerator.
The upside is enormous efficiency if Transformers remain dominant. The downside is equally obvious: AI architectures can change.
I'd therefore rank Etched below Cerebras/Groq/SambaNova on commercial maturity but potentially above them on asymmetric upside.
The broader opportunity
I'd actually divide the market into four ways of reducing GPU dependence:
The strongest startups reducing Graphics Processing Unit (GPU) dependence approach the problem from two angles: alternative hardware architecture (custom AI chips) and software-layer optimization (maximizing efficiency to use fewer chips).
As the industry shifts from training large frontier models to production inference, these startups are winning enterprise workloads by cutting costs and breaking the vendor lock-in of traditional GPU suppliers.
The leading AI infrastructure startups addressing GPU dependence are categorized below by their technological approach:
Instead of relying on general-purpose GPUs, these companies design Application-Specific Integrated Circuits (ASICs) optimized purely for AI workloads.
Cerebras Systems : Famous for its "Wafer-Scale Engine," which places an entire computer chip on a single silicon wafer. It provides massive compute density, significantly reducing the physical footprint and GPU count required for large-scale training and inference.
Groq : Built the Language Processing Unit (LPU) architecture. It offers incredibly fast, low-latency inference speeds specifically for Large Language Models (LLMs), removing the need for traditional GPU clustering in conversational AI.
SambaNova Systems : Utilizes a "Reconfigurable Dataflow Architecture" that handles complex AI workloads and enterprise-scale inference with higher efficiency and lower energy consumption than legacy chips.
These startups don’t build hardware; instead, they make software that squeezes maximum performance out of existing infrastructure, drastically lowering the total number of GPUs an enterprise needs to lease or buy.
Cast AI : A prominent leader in the GPU efficiency landscape. It automates Kubernetes cluster management, aggressively optimizing model-aware allocation and reducing architectural waste across multi-cloud environments.
ScaleOps : Rapidly closing in on infrastructure management by automating live runtime optimization and dynamically adjusting resource allocation so workloads don’t leave idle hardware running.
Decart : Focuses closer to the runtime execution layer. It specializes in helping AI workloads run seamlessly across different families of chips, allowing enterprises to easily swap GPUs for cheaper custom silicon.
High-Efficiency Inference Platforms
These software-layer platforms focus on serving open-weights models at commercial scale with hyper-optimized token delivery.
Together AI : Boasts massive scale with over $1.15 billion in annual bookings. Its software stack optimizes model fine-tuning and inference to dramatically lower the computing overhead required by its enterprise clients.
Fireworks AI : Serves over 40 trillion tokens per day, achieving massive hardware savings through advanced model-quantization and compilation techniques that let smaller clusters handle massive request volumes.
FriendliAI : Highly specialized in inference optimization, utilizing unique scheduling and memory management techniques to maximize token throughput per chip.
Direct Comparison Overview
Startup
Core Layer
Primary Mechanism to Reduce GPU Reliance
Groq
Hardware (ASIC)
Replaces GPUs entirely for ultra-fast, low-latency LLM inference.
Etched
Hardware (ASIC)
Hardcodes Transformer architecture into silicon for specialized maximum efficiency.
Cast AI
Software (Orchestration)
If you want to narrow down your search, let me know:
Are you evaluating these from an investment perspective or looking for a vendor solution?
Is your primary workload focus heavy model training or high-volume production inference?
Do you prefer a purely software-driven approach or migrating to alternative chip hardware?
I can tailor a deep dive into the performance benchmarks or architectural details of any of these players.
FriendliAI : Specializes in serving inference workloads at maximum throughput with minimal hardware overhead, lowering the overall GPU count needed to service high-volume user traffic.
d-Matrix & SambaNova Systems : Deliver alternative architectural approaches (such as memory-centric and reconfigurable dataflow architectures) targeted at high-efficiency generative AI inference without relying entirely on traditional monolithic GPU setups.
Etched : Specifically designs specialized chips built to run a single type of AI architecture—Transformers. By burning the Transformer algorithm directly into the silicon, it delivers unprecedented throughput compared to flexible GPUs.
Eliminates hardware waste through automated, model-aware Kubernetes scaling.
Together AI
Software (Inference Platform)
Optimizes how models compile and route to drastically lower overall hardware usage.
Decart
Software (Runtime)
Bridges the gap between chip families to allow easy migration off traditional GPUs.