If you need a single, production-ready browser ML library, choose TensorFlow.js for its balance of power and ecosystem. Pick ONNX Runtime Web when you need maximum client-side performance for exported PyTorch/Keras models. Use Transformers.js or WebLLM for in-browser NLP/LLM tasks, ml5.js for beginner-friendly prototyping, MediaPipe for vision, and Brain.js for very small networks.
Explore the market ranking
ML Deployment & Inference Optimization Tools
Brands AI recommends here
Mentioned in
1TensorFlow.jsBest general-purpose choice for web developers who want a production-ready, widely supported browser ML stack with broad model and tooling support; tradeoff: heavier than tiny libraries.88%
2ONNX RuntimeBest for performance-focused teams that export PyTorch/Keras models to ONNX and need near-native CPU/WebGPU speed in the browser; tradeoff: optimized around ONNX-format workflows.65%
3ml5.jsBest for educators, artists, and rapid prototyping where simplicity matters—high-level APIs built on TensorFlow.js; tradeoff: less control for advanced performance tuning.55%
4Transformers.jsBest when you need transformer/LLM or Hugging Face model inference in-browser for NLP applications and privacy-sensitive use without a backend; tradeoff: focused on transformers.48%
Does AI recommend your brand?
Recommendations by AI engine
Recommended in · Sep 5 – Sep 21, 2026
Transformers.jsChatGPT Search5 of 5 answersGoogle AI Mode5 of 5 answers
TensorFlow.jsChatGPT Search5 of 5 answersGoogle AI Mode5 of 5 answers
ONNX RuntimeChatGPT Search5 of 5 answersGoogle AI Mode4 of 5 answers
WebLLMChatGPT Search0 of 5 answersGoogle AI Mode3 of 5 answers
MediaPipeChatGPT Search0 of 5 answersGoogle AI Mode3 of 5 answers
Hugging FaceChatGPT Search1 of 5 answersGoogle AI Mode1 of 5 answers
The "best" JavaScript library for running machine learning client-side in the browser depends heavily on what kind of model you want to run and your ecosystem . Client-side ML has shifted away from heavy in-browser training toward high-performance inference powered by modern browser acceleration like WebGPU, WebNN , and WebAssembly (WASM).
The top options break down by use case:
Transformers.js (by Hugging Face) — Best for NLP, Audio, Vision, and Transformer models.
Why it’s great: It brings state-of-the-art transformer models directly to the browser with a high-level pipeline() API that feels identical to Python's transformers library. Under the hood, it uses ONNX Runtime Web and seamlessly supports WebGPU for hardware acceleration alongside a fallback. You can load hundreds of thousands of pre-converted models straight from the Hugging Face Hub.
Best for: Embeddings, sentiment analysis, object detection, speech-to-text, and small language models.
ONNX Runtime Web (by Microsoft) — Best for custom or framework-agnostic models.
Why it’s great: This is a lower-level execution engine. If you train a custom model in PyTorch, TensorFlow, or scikit-learn, export it to the open ONNX format, and need absolute control over input/output tensors and execution graphs, this is the gold standard. It boasts robust optimization across WebGPU, WebNN, WebGL, and WASM.
Best for: Production web apps running proprietary, highly-customized, or non-transformer architectures.
WebLLM (by MLC AI) — Best for running full-scale Large Language Models (LLMs) in-browser.
Why it’s great: If you want to run heavy generative text models (like Llama 3 or Mistral) locally in the user's browser via WebGPU with hardware acceleration, WebLLM acts like MLc-Caffe/TVM compiled for the web. It handles chat generation and streaming natively with minimal boilerplate.
Best for: In-browser local chatbots and AI assistants requiring heavy LLM execution.
TensorFlow.js (by Google) — Best if you need on-device training or legacy ecosystem support.
Why it’s great: It is the most mature and battle-tested framework for JavaScript ML. Unlike others that focus strictly on pre-trained inference, TensorFlow.js allows you to train models directly in the browser using client data. However, for standard inference tasks, it has largely been outperformed in developer ergonomics and speed by Transformers.js and ONNX.
Best for: Interactive web apps that require fine-tuning or training models client-side using user interactions/sensors.
MediaPipe JS (by Google) — Best for real-time computer vision.
Why it’s great: It isn't a general-purpose ML builder, but a suite of pre-packaged, highly optimized solutions for face mesh, hand tracking, pose estimation, and selfie segmentation. It runs blazingly fast in the browser via WebAssembly and WebGL/WebGPU.
Best for: Webcam effects, augmented reality (AR), and gesture controls.
Quick Recommendation
Choose Transformers.js if you want the easiest path to running modern pre-trained AI models (BERT, Whisper, CLIP, Vision Transformers) with minimal setup.
Choose ONNX Runtime Web if you are bringing your own custom-trained model.
Choose WebLLM if your primary goal is running a heavy open-weights LLM locally.
If you tell me what specific task you want the model to perform (e.g., text generation, image classification, audio transcription) and what framework you used to train it (if custom), I can recommend the exact implementation approach and performance optimizations you'll need.
If your goal is general-purpose ML inference directly in the browser, I’d choose Hugging Face Transformers.js for most new projects.
My recommendation
Library
Best for
Browser GPU
Model ecosystem
Ease of use
Transformers.js
Modern NLP, LLMs, vision, audio
WebGPU
Excellent
⭐⭐⭐⭐⭐
ONNX Runtime Web
Maximum control/performance with ONNX
WebGPU
ONNX
⭐⭐⭐⭐
TensorFlow.js
TensorFlow models / custom JS ML
WebGPU
TensorFlow.js
⭐⭐⭐⭐
Why Transformers.js: it gives you a high-level pipeline() API while using ONNX Runtime underneath, supports text, vision, audio and multimodal tasks, and can run models through WebAssembly on CPU or WebGPU on compatible devices.
For example:
import { pipeline } from "@huggingface/transformers";
const classifier = await pipeline(
"sentiment-analysis",
"Xenova/distilbert-base-uncased-finetuned-sst-2-english",
{ device: "webgpu" }
);
const result = await classifier("This is surprisingly good!");
console.log(result);
For maximum low-level control, I'd use ONNX Runtime Web directly. It lets you load an ONNX model and explicitly select WebGPU, WASM, etc.; it's particularly attractive if you already have a model exported to ONNX or need to optimize the inference pipeline yourself.
import * as ort from "onnxruntime-web/webgpu";
const session = await ort.InferenceSession.create(
"/model.onnx",
{ executionProviders: ["webgpu"] }
);
One important caveat: WebGPU isn't universal yet. Current browser support is substantial, but Transformers.js notes roughly 85% global support as of March 2026, with browser/platform differences. You should provide a WASM/CPU fallback if your application needs broad compatibility.
Rule of thumb
Building an AI feature quickly → Transformers.js
You control/own the ONNX model and care about performance → ONNX Runtime Web
Tiny conventional models where compatibility matters most → ONNX Runtime Web + WASM
Also, for browser inference, quantized models are usually worth prioritizing because model download size and client memory are major constraints. Transformers.js explicitly recommends quantized models for resource-constrained environments such as browsers.
If you tell me what model you want to run (LLM, embedding model, image model, Whisper, custom PyTorch model, etc.) and your target browsers, I can recommend the exact stack and model format.
If your goal is ML inference directly in a browser, my default recommendation today is huggingface.co—if your models are in the Hugging Face/Transformer ecosystem. It uses ONNX Runtime underneath and gives you a much higher-level API than working with tensors and execution providers directly.
Arbitrary/custom ONNX models, maximum runtime control
WebGPU
Low-level
TensorFlow.js
TensorFlow/Keras models and TF.js-specific workflows
WebGL/WebGPU
Medium
Transformers.js is particularly attractive for text, embeddings, image classification/object detection, speech recognition, and similar workloads. It can use WASM on CPU or WebGPU, and supports quantized models such as q8 and q4, which can substantially reduce browser download and memory requirements.
For example:
import { pipeline } from "@huggingface/transformers";
const classifier = await pipeline(
"sentiment-analysis",
"Xenova/distilbert-base-uncased-finetuned-sst-2-english",
{ device: "webgpu" }
);
const result = await classifier("This runs entirely in the browser!");
console.log(result);
When I'd use ONNX Runtime Web instead
Choose onnxruntime.ai if you have your own ONNX models or need fine-grained control over model execution. It supports WASM, WebGPU, WebGL, and WebNN depending on browser/platform.
import * as ort from "onnxruntime-web/webgpu";
const session = await ort.InferenceSession.create(
"/model.onnx",
{ executionProviders: ["webgpu"] }
);
const output = await session.run({
input: new ort.Tensor("float32", data, [1, 3, 224, 224])
});
WebGPU is generally the direction I'd target for compute-heavy browser inference; ONNX Runtime specifically recommends WebGPU over its older WebGL path for better performance.
Bottom line
Building an AI feature around existing Hugging Face models → Transformers.js
Deploying custom ONNX models / maximum control → ONNX Runtime Web
Already invested heavily in TensorFlow/Keras → TensorFlow.js
Need broad browser compatibility → use WASM fallback
Need good performance on capable devices → WebGPU + quantized models
If you tell me what model you want to run (LLM, image model, Whisper, embeddings, etc.) and your target browsers, I can recommend the specific runtime/model format and architecture.
The "best" JavaScript library for client-side browser inference depends entirely on what kind of model you are trying to run . There is no single silver bullet, but the ecosystem has consolidated around a few standout champions based on your specific use case.
Here is the breakdown of the top contenders and when to use them:
Transformers.js (by Hugging Face)
Best for: General-purpose AI (NLP, embeddings, computer vision, audio, and smaller generative tasks) directly using Hugging Face models.
Why it shines: It brings the Hugging Face ecosystem straight to the browser using ONNX Runtime Web under the hood. You can run thousands of pre-trained models (like BERT, Whisper, or small vision models) with just a few lines of JavaScript without setting up a backend server. It supports WebAssembly (WASM) and WebGPU for hardware acceleration.
Check it out: Transformers.js on Hugging Face
WebLLM (by MLC-AI)
Best for: Running large language models (LLMs) like Llama 3, Phi-3, or Mistral completely client-side.
Why it shines: If your goal is a high-performance in-browser chatbot or local assistant, WebLLM leverages WebGPU and optimized compilation (MLC) to deliver native-like token generation speeds directly in the browser tab. It handles model caching and memory management exceptionally well for heavy generative models.
TensorFlow.js (TF.js)
Best for: Traditional ML, custom-trained Keras/TensorFlow models, and fine-grained tensor operations in the browser.
Why it shines: It is the veteran framework of the bunch. It supports hardware acceleration via WebGL, WebAssembly, and WebGPU. Choose TF.js if you trained a custom model in Python using TensorFlow/Keras, need to retrain or adapt a model on the client side, or need low-level linear algebra operations.
Check it out: TensorFlow.js Documentation
Google MediaPipe (Web)
Best for: Real-time, ultra-fast computer vision and audio tasks (face mesh, pose estimation, hand tracking, object detection, segmentation).
Why it shines: If you need real-time 60fps video processing on a user's webcam feed, MediaPipe Solutions is heavily optimized out-of-the-box for low latency and minimal resource consumption on mobile and desktop browsers alike.
Quick Decision Matrix
Want to run a standard open-source Hugging Face model (embeddings, text classification, image-to-text)? Use Transformers.js.
Want to run a heavy conversational LLM locally in the browser? Use WebLLM.
Migrating an existing Python TensorFlow/Keras model to the web? Use TensorFlow.js.
Doing real-time webcam/audio tracking or segmentation? Use MediaPipe.
To help narrow this down, could you tell me:
What specific type of model or task are you trying to run (e.g., LLM chat, image classification, audio transcription)?
Do you already have a pre-trained model file ready, or are you starting from scratch?
Running machine learning inference directly in the browser depends heavily on your specific use case, what kind of model you plan to run, and whether you are loading a pre-trained transformer or rolling out a custom architecture.
The client-side JavaScript ML ecosystem centers around three primary libraries, each serving a distinct purpose:
Transformers.js (by Hugging Face)
Best for: Running state-of-the-art Hugging Face models (like BERT, Whisper, or small LLMs and vision transformers) directly in the browser with minimal friction.
How it works: It uses ONNX Runtime Web under the hood, backed by WebAssembly (WASM) and WebGPU for acceleration. It provides high-level pipelines matching the Python transformers library.
Pros: Drop-in access to tens of thousands of pre-trained models; excellent task-level APIs (sentiment analysis, text generation, object detection); robust hardware fallback options.
Cons: Less optimized if you want to build and train custom, non-transformer layers from scratch.
ONNX Runtime Web (onnxruntime-web)
Best for: Production-grade performance and direct, low-level control over custom or converted ONNX models (exported from PyTorch, JAX, or TensorFlow).
How it works: Microsoft’s high-performance inference engine compiled for the web. It features execution providers targeting WebGPU, WebNN, WebGL , and WebAssembly.
Pros: Incredible execution speed, excellent hardware acceleration (leveraging modern WebGPU/WebNN standards), and fine-grained control over tensors and memory.
Cons: Lower-level API means you have to handle tokenization, pre-processing, and post-processing yourself compared to task-based libraries.
TensorFlow.js (TF.js)
Best for: Training models directly in the browser, interactive educational web apps, or deploying legacy models originally trained in the TensorFlow ecosystem.
How it works: Google’s end-to-end machine learning framework for JavaScript supporting WebGL and WebGPU acceleration.
Pros: The most mature and historically established framework; allows client-side fine-tuning or training from scratch; pairs nicely with high-level wrappers like ml5.js for creative coding and education.
Cons: Heavier bundle sizes and a larger dependency tree; for pure inference on modern LLMs or transformer pipelines, it has largely been outperformed in developer preference by ONNX-backed runtimes.
Quick Decision Guide
Choose Transformers.js if you want to quickly integrate a standard NLP, audio, or vision model from the Hugging Face Hub.
Choose ONNX Runtime Web if you have a custom-trained model, need max throughput via WebGPU/WebNN, and want absolute control over tensor operations.
Choose TensorFlow.js if you need client-side model training or are working within an existing TensorFlow-heavy stack.
To help narrow down the best fit, could you share:
What type of model you plan to run (e.g., LLM, image classification, audio transcription)?
Are you using a pre-trained model or a custom model you trained yourself?
If your goal is ML inference entirely in the browser, my default recommendation today is Hugging Face Transformers.js.
Best default: Transformers.js
Hugging Face's Transformers.js is probably the easiest choice if you're working with modern NLP, vision, audio, embeddings, or multimodal models. It runs models directly in the browser and uses ONNX Runtime Web underneath.
npm install @huggingface/transformers
Then:
import { pipeline } from "@huggingface/transformers";
const classifier = await pipeline(
"sentiment-analysis",
"Xenova/distilbert-base-uncased-finetuned-sst-2-english",
{
device: "webgpu",
}
);
const result = await classifier("I love running models in the browser!");
console.log(result);
It supports WebGPU for GPU acceleration and can fall back to WASM/CPU. Quantized models such as q4 and q8 are particularly useful in browsers because model download size and memory matter a lot.
When I'd choose something else
Library
Best for
Transformers.js
⭐ Most general-purpose browser ML; Hugging Face models; NLP/audio/vision
ONNX Runtime Web
Maximum control/performance with your own ONNX models
TensorFlow.js
Existing TensorFlow/Keras models and TF.js ecosystem
WebGPU directly
Very specialized/custom ML workloads where you're willing to build more infrastructure
ONNX Runtime Web is the lower-level choice I'd use if you already have an ONNX model and care about controlling preprocessing, tensors, execution providers, memory, etc. It supports WebGPU as an execution provider.
My recommendation
For a new project, I'd generally use:
Model → ONNX → Transformers.js → WebGPU → Web Worker
The Web Worker is important: inference on the main JS thread can freeze the UI, and Hugging Face specifically recommends moving model loading/inference into a worker for production applications.
If you tell me what model you're trying to run (LLM, Whisper, embeddings, image classifier, YOLO, custom PyTorch model, etc.), I can recommend the specific library/model/runtime combination and expected browser performance.
Very easy API: its pipeline() API makes loading and running models straightforward.
Broad model support: NLP, LLMs, embeddings, computer vision, speech/audio, and multimodal models.
WebGPU support: you can run inference on the user's GPU with device: "webgpu".
WebAssembly fallback: browser inference can run through ONNX Runtime/WASM when GPU acceleration isn't available.
Quantization: useful for keeping model downloads and memory usage manageable in browsers; Transformers.js supports formats such as q4 and q8.
ONNX underneath: models can be converted from PyTorch, TensorFlow, or JAX to ONNX, giving you a fairly broad deployment path.
A minimal example:
import { pipeline } from "@huggingface/transformers";
const classifier = await pipeline(
"sentiment-analysis",
"Xenova/distilbert-base-uncased-finetuned-sst-2-english",
{
device: "webgpu",
}
);
const result = await classifier("This runs entirely in my browser!");
console.log(result);
When I'd choose something else
Library
Best choice when...
Transformers.js
You want pretrained NLP/LLM/vision/audio models with the least development work
ONNX Runtime Web
You have your own ONNX model and want maximum control over inference/runtime behavior
TensorFlow.js
You're already invested in TensorFlow/Keras or want to train/retrain models in JavaScript
WebGPU directly
You're building a highly specialized inference engine and need low-level GPU control
ONNX Runtime Web is particularly compelling if your team owns the model pipeline. It exposes the WebGPU execution provider directly and gives you lower-level control over tensors, execution, and GPU memory.
TensorFlow.js remains a good option for TensorFlow-centric applications and supports both running existing models and training/retraining in the browser.
My recommendation
If you're starting a new project, I'd use:
Transformers.js → WebGPU when available → WASM fallback → quantized models → Web Worker for inference.
That gives you a relatively high-level developer experience while still taking advantage of the browser GPU. For a production application, I'd also put inference in a Web Worker so large model execution doesn't freeze the UI; current Transformers.js integrations provide worker-oriented patterns for this.
If you tell me what model you're trying to run (LLM, embeddings, image model, Whisper, custom PyTorch model, etc.) and your target browsers, I can recommend the specific library/runtime and model format.
There is no single "best" library because the ideal choice depends heavily on what kind of model you want to run (e.g., standard computer vision/embeddings vs. large language models) and your target ecosystem.
The modern client-side AI landscape centers on a few standout libraries, heavily accelerated by modern web compute abstractions like WebAssembly (WASM) and WebGPU:
1. Best for LLMs & Generative AI: WebLLM
Best for: Running heavy Large Language Models (like Llama 3, Phi-3, or Qwen) locally in the browser.
How it works: It brings the MLC-LLM ecosystem to the web, leveraging WebGPU for high-throughput acceleration.
Pros: Unmatched performance for generative text models in the browser; supports a wide range of open-weight chat models; manages model weights and caching smoothly.
Cons: Strictly built for LLMs; requires WebGPU support (won't work gracefully on older hardware/browsers without fallback limitations).
2. Best for NLP, Audio, & Vision Pipelines: Transformers.js
Best for: Running standard Hugging Face models (sentiment analysis, object detection, embeddings, Whisper audio transcription) completely client-side.
How it works: Developed by Hugging Face, it is functionally equivalent to the Python transformers library, but runs in the browser using ONNX Runtime Web under the hood.
Pros: Direct access to tens of thousands of pretrained models on the Hugging Face Hub; high-level pipeline API makes implementation trivial; robust WebGPU and WASM fallback support.
Not designed for massive generative LLMs (though smaller ones are supported, heavy LLMs are better suited to ).
3. Best for Framework Agnosticism & Custom Models: ONNX Runtime Web
Best for: Maximum flexibility when you have custom models trained in PyTorch, TensorFlow, or scikit-learn and want to export them to ONNX format for web deployment.
How it works: Microsoft’s cross-platform inference engine compiled for the web. It automatically targets WebGPU, WebGL, or WebAssembly depending on client device capabilities.
Pros: Extremely fast and optimized; framework-agnostic (export from PyTorch/TF); handles hardware acceleration gracefully.
Cons: Verbose API. You have to handle your own tokenization, pre-processing, and post-processing logic compared to the plug-and-play nature of Transformers.js.
4. Best for Real-Time Computer Vision: MediaPipe (by Google)
Best for: Face tracking, hand-pose estimation, selfie segmentation, and object detection.
How it works: Highly optimized, production-ready pre-packaged solutions tailored specifically for real-time perception tasks.
Pros: Incredible out-of-the-box performance on mid-range and mobile hardware; handles end-to-end media pipelines smoothly.
Cons: Limited strictly to Google’s curated set of perception models; not a general-purpose ML runner.
(Note on TensorFlow.js : While historically the go-to standard, it has largely been superseded for high-performance inference by ONNX Runtime and WebGPU-native tooling, though it remains useful if you are deeply embedded in the TensorFlow ecosystem or need client-side training.)
Quick Decision Matrix
Choose WebLLM if you want an in-browser chatbot or text-generation assistant.
Choose Transformers.js if you want an easy, high-level API to pull standard models from Hugging Face for embeddings, audio, or vision.
Choose ONNX Runtime Web if you trained a custom model and want the fastest, rawest ONNX execution layer.
Choose MediaPipe if you are strictly building real-time camera/video tracking features.
To help narrow this down, tell me:
What specific ML task are you trying to accomplish (e.g., text generation, image classification, audio transcription)?
Are you bringing your own trained model , or looking for a pre-trained model?
The best JavaScript library for running machine learning models client-side depends entirely on your specific model architecture and use case, but Transformers.js by Hugging Face is the top recommendation for modern NLP, computer vision, and audio models, while ONNX Runtime Web (ORT-Web) or LiteRT.js are better suited for custom or lightweight architectures.
Running generic, custom-trained ONNX models across computer vision and tabular tasks.
WebGPU, WebGL, WebAssembly
.js
Key Library Breakdown
Transformers.js (Hugging Face): The definitive choice if you want to run modern pre-trained models (like BERT, Llama-style small language models, or Whisper speech-to-text) without building pipelines from scratch. It mirrors the Python transformers API and handles tokenization and execution seamlessly.
ONNX Runtime Web: The underlying engine for many modern web frameworks. Choose ORT-Web if you train custom models in PyTorch or JAX, export them to the ONNX format, and need fine-grained control over the inference session and memory management in the browser.
LiteRT.js (formerly TFLite Web): Google's high-performance native web runtime. It executes .tflite files via WebAssembly and taps straight into hardware accelerators like WebNN and XNNPACK, making it ideal for mobile-first web apps.
Hardware Acceleration Considerations
To achieve near-native speeds in client-side inference, ensure your implementation targets modern browser acceleration APIs:
WebGPU: The gold standard for client-side GPU acceleration, offering massive performance boosts over older WebGL contexts.
WebNN: An emerging web API designed to hook directly into device NPUs (Neural Processing Units) for energy-efficient local execution.
WebAssembly (WASM): Used as a high-performance fallback when GPU/NPU acceleration is unavailable, heavily optimized via multithreading and SIMD.
If you tell me what specific model or task (e.g., LLM text generation, image segmentation, speech recognition) you plan to run, I can recommend the ideal library configuration and quantization approach.
If your goal is ML inference entirely in the browser, my default recommendation today is huggingface.co.
It depends somewhat on what models you're running:
Library
Best for
Browser acceleration
Ease of use
Transformers.js
LLMs, embeddings, NLP, vision, audio
WebGPU + WASM
⭐⭐⭐⭐⭐
ONNX Runtime Web
Custom ONNX models / maximum control
WebGPU + WASM + WebNN
⭐⭐⭐⭐
TensorFlow.js
TensorFlow/Keras models, browser ML
WebGPU/WebGL/WASM
⭐⭐⭐⭐
My recommendation: Transformers.js
For modern AI applications, I'd start here. It lets you run pretrained Transformer models directly in the browser and provides a high-level pipeline() API, so you don't have to manually implement tokenization, preprocessing, inference, and postprocessing.
For example:
import { pipeline } from "@huggingface/transformers";
const classifier = await pipeline(
"sentiment-analysis",
"Xenova/distilbert-base-uncased-finetuned-sst-2-english",
{ device: "webgpu" }
);
const result = await classifier("This is fantastic!");
console.log(result);
Transformers.js uses ONNX-based execution under the hood, and its newer releases have significantly improved WebGPU support and performance.
When I'd choose ONNX Runtime Web instead
If you have your own model—for example, a PyTorch model that you've exported to ONNX—and you care about squeezing out performance or controlling the inference pipeline yourself, I'd use onnxruntime.ai directly.
import * as ort from "onnxruntime-web/webgpu";
const session = await ort.InferenceSession.create(
"/models/model.onnx",
{ executionProviders: ["webgpu"] }
);
const output = await session.run({
input: new ort.Tensor("float32", data, [1, 3, 224, 224])
});
ONNX Runtime Web supports WASM for broad compatibility and WebGPU for more compute-intensive workloads. Its documentation specifically recommends WebGPU when you want to take advantage of the client's GPU.
Practical decision
I'd use:
LLM / embeddings / speech / computer vision model from Hugging Face → Transformers.js
Your own ONNX model → ONNX Runtime Web
Existing TensorFlow/Keras model → TensorFlow.js
Maximum performance → ONNX Runtime Web + WebGPU, assuming your target browsers support it
Need to work on weaker/older browsers → WASM fallback
One important architectural point: WebGPU should generally be your preferred backend for heavier models, with WASM as the fallback. ONNX Runtime's current browser documentation notes that WebGL is now in maintenance mode and recommends WebGPU for better performance.
If you tell me what model you want to run (e.g. Llama, Whisper, YOLO, BERT, a custom PyTorch model), I can recommend the exact browser stack and model format.
Running optimized Google .tflite models with low latency and native device bindings.
XNNPACK (CPU), ML Drift (GPU), WebNN
TensorFlow.js
Legacy TensorFlow/Keras workflows and standard pre-trained browser models.
WebGL, WebAssembly, CPU
TensorFlow.js: Best if you are already deeply embedded in the TensorFlow ecosystem or require older specific layers and training-in-the-browser capabilities. However, for pure inference performance on newer architectures, it has largely been surpassed by ONNX/WebGPU-backed stacks.