Data as of Sep 18, 2026 · Based on 288 AI responses from ChatGPT Search and Google AI Mode · See how Parse measures this
Pharmaceutical research teams have access to various biology-specific foundation models depending on their focus. NVIDIA BioNeMo serves as a central hub hosting diverse models for drug discovery. For more targeted research, companies like EvolutionaryScale (ESM-3) and Google DeepMind (AlphaFold 3) are leaders in protein modeling, while Atomic AI and
Deep Genomics specialize in RNA-based discovery, and
provides multi-omics and interactome models.
Brands AI recommends here
Named in 45% of answers
Named in 44% of answers
Named in 39% of answers
Named in 35% of answers
Yes. As of **September 2026**, there is a fairly broad ecosystem of biology-specific foundation-model providers. The important distinction is that some provide **general biological foundation models**, while others specialize in **proteins, molecular structure, chemistry, cells, or drug-discovery workflows**. ###…
Yes. As of September 2026, there is a fairly broad ecosystem of biology-specific foundation-model providers. The important distinction is that some provide general biological foundation models, while others specialize in proteins, molecular structure, chemistry, cells, or drug-discovery workflows.
| Provider | Representative models/platform | Biological focus | Typical drug-discovery use |
|---|---|---|---|
| EvolutionaryScale / Biohub | ESM3, ESM Cambrian | Protein sequences, structure, function, generation | Protein engineering, therapeutic proteins, sequence design |
| Boltz | Boltz-2, BoltzGen, BoltzProt-1, BoltzMol-1 | Biomolecular structure, protein design, small molecules | Structure prediction, affinity estimation, hit discovery, biologics |
| NVIDIA | BioNeMo | Proteins, molecules, DNA/RNA, chemistry | Foundation-model development, customization, deployment, drug-design workflows |
| Recursion | Recursion's biology/chemistry foundation models | Cellular phenotypes, multi-omics, biology + chemistry | Target discovery, phenotypic screening, disease biology |
| Genentech/Roche | BRAID and associated models | Perturbation biology, genomics, spatial biology, multimodal biology | Target identification, biological mechanism, virtual screening |
| Chai Discovery | Chai models | Biomolecular structure and interactions | Protein–ligand/biomolecular modeling and design |
| Basecamp Research | EDEN family | Protein/DNA biology and biological design | Protein engineering and genetic medicine |
| OpenFold | OpenFold models | Protein structure | Structure prediction and downstream protein design |
| Terray Therapeutics | Structural/chemistry foundation models | Molecular binding and generative chemistry | Small-molecule discovery and optimization |
| Iambic | NeuralPLexer and related models | Protein–ligand structure | Structure prediction and molecular design |
NVIDIA's BioNeMo ecosystem is particularly broad: it supports 3D protein-structure prediction, de novo protein and small-molecule generation, property prediction and docking, and allows organizations to customize models with their own data.
For protein foundation models specifically, EvolutionaryScale is one of the most important names. Its ESM3 is a generative protein model, while ESM Cambrian is its newer representation-learning family. EvolutionaryScale notes that ESM3 models are now available through Biohub.
Boltz is particularly relevant if your objective is closer to actual pharmaceutical discovery than generic biology. Its current portfolio spans biomolecular structure, protein design and small-molecule hit discovery; in June 2026 it announced BoltzMol-1, BoltzProt-1 and an API.
Recursion takes a different approach: its foundation models are tied to very large experimental biological and chemical datasets, including cellular phenotyping. The company describes its NVIDIA collaboration as focused on scaling foundation models toward a "virtual cell."
Genentech's BRAID is another interesting category: rather than concentrating primarily on protein sequence/structure, it develops foundation models over perturbational, genomic, spatial and other biological data, with applications including target identification and drug discovery.
For a pharma AI strategy, I would think about the landscape as five model families rather than simply asking who has a "biology foundation model":
That distinction matters because the model provider is increasingly only one part of the moat. Recent work suggests proprietary pharmaceutical protein-structure data can materially improve foundation-model performance, which makes access to experimental data and the ability to close the prediction → experiment → retraining loop strategically important.
If you're evaluating vendors for a pharmaceutical discovery program, I can also give you a more useful vendor landscape of ~15–25 providers, broken down by protein, small molecule, target discovery, multimodal biology, virtual cell, licensing/API availability, open vs. closed weights, and suitability for pharma deployment.
Boltz is particularly relevant if your objective is closer to actual pharmaceutical discovery than generic biology. Its current portfolio spans biomolecular structure, protein design and small-molecule hit discovery; in June 2026 it announced BoltzMol-1, BoltzProt-1 and an API.
Recursion takes a different approach: its foundation models are tied to very large experimental biological and chemical datasets, including cellular phenotyping. The company describes its NVIDIA collaboration as focused on scaling foundation models toward a "virtual cell."
Genentech's BRAID is another interesting category: rather than concentrating primarily on protein sequence/structure, it develops foundation models over perturbational, genomic, spatial and other biological data, with applications including target identification and drug discovery.
The landscape for biology-specific foundation models has shifted rapidly from narrow molecular screening tools to general-purpose, multimodal "digital biology" architectures. These models reason natively across sequences, 3D structures, and chemical/biological…
The landscape for biology-specific foundation models has shifted rapidly from narrow molecular screening tools to general-purpose, multimodal "digital biology" architectures. These models reason natively across sequences, 3D structures, and chemical/biological functions.
Providers of biology-specific foundation models span specialized AI-native biotech labs, big tech spinouts, and open-science organizations:
If you'd like to narrow this down, tell me:
I can recommend the best integration path for your team.
Entering the space of generative AI for drug discovery means choosing whether to build on top of commercial/open-source **biology foundation models (BioFMs)** or partner with specialized AI-native biotech companies. The primary providers of biology-specific foundation models span specialized AI labs, big tech…
Entering the space of generative AI for drug discovery means choosing whether to build on top of commercial/open-source biology foundation models (BioFMs) or partner with specialized AI-native biotech companies.
The primary providers of biology-specific foundation models span specialized AI labs, big tech spinouts, infrastructure platforms, and biotech pioneers:
Are you looking to fine-tune an open-source model (like ESM or Nucleotide Transformer) on your own internal data, or are you searching for an end-to-end commercial discovery partner?
DNABERT
Yes. If by **“biology-specific foundation models”** you mean models pretrained on biological data and intended to be adapted for drug discovery, protein engineering, genomics, or cellular biology, the landscape in 2026 is fairly broad. ### Leading providers to know Provider | Representative models | Primary biological…
Yes. If by “biology-specific foundation models” you mean models pretrained on biological data and intended to be adapted for drug discovery, protein engineering, genomics, or cellular biology, the landscape in 2026 is fairly broad.
| Provider | Representative models | Primary biological focus | Relevance to pharma |
|---|---|---|---|
| EvolutionaryScale / Biohub | ESM-3, ESM Cambrian | Protein sequence, structure & function | Protein/antibody design, enzyme engineering, biologics |
| Boltz | Boltz-2, BoltzGen, BoltzProt-1, BoltzMol-1 | Biomolecular structure, affinity, protein & molecule design | Small-molecule discovery, biologics, target–ligand modeling |
| Arc Institute | Evo 2, State, Stack | Genomes, DNA/RNA, cellular responses | Genomics, variant effects, functional genomics, virtual-cell research |
| Profluent | Protein language models, E1 | Protein sequence/function | Protein and genome-editor design |
| Generate:Biomedicines | Generative protein models / Chroma | De novo protein and biologic design | Therapeutic protein and antibody discovery |
| Chai Discovery | Chai-1 and subsequent models | Biomolecular structure/design | Protein–ligand and protein–protein modeling |
| Google DeepMind / Isomorphic Labs | AlphaFold 3, drug-design systems | Biomolecular interactions | Structure prediction, target/ligand/biologic discovery |
| Xaira Therapeutics | Foundation-model stack including X-Cell and protein-design models | Cellular + molecular biology | Drug discovery and therapeutic design |
| Recursion | Large-scale biological/phenotypic models | Cellular phenomics | Target discovery, disease biology, compound screening |
| Insilico Medicine | Biology42 / PandaOmics and related models | Disease biology + target discovery | Target identification and drug discovery |
A few distinctions are important.
For protein foundation models, I'd put EvolutionaryScale, Profluent, Generate:Biomedicines, Boltz and Chai near the top of the evaluation list. EvolutionaryScale's ESM3 is particularly notable because it jointly represents protein sequence, structure and function and can generate proteins from combinations of those specifications.
For molecular/structural foundation models, Boltz is particularly interesting. Its current portfolio spans protein design, small-molecule hit discovery and biomolecular modeling; in 2026 it released BoltzMol-1 and BoltzProt-1, while Pfizer announced a collaboration to develop models using Pfizer's proprietary data.
For genomics, Arc Institute's Evo 2 is a major model to investigate. It was trained on more than 9 trillion nucleotide tokens from genomes spanning the tree of life and is designed to model and generate biological sequences. Arc is also developing models for cellular perturbation and virtual-cell applications.
For protein engineering specifically, Profluent is another important player. Its platform combines foundation models with an in-house wet lab, and its E1 model is a protein encoder available under a permissive license for research and commercial use.
If you're preparing an AI strategy for pharmaceutical discovery, I would not treat all of these as interchangeable “biology foundation model” vendors. I'd divide them into:
That distinction matters because a foundation model provider isn't necessarily a drug-discovery platform provider. For example, ESM3 is fundamentally a protein-generation model, whereas Recursion is built around large-scale experimental phenomics and drug-discovery workflows.
If you're deciding what to actually evaluate for a pharmaceutical R&D organization, I'd shortlist roughly 8–12 models and benchmark them against your modalities (small molecules, antibodies, peptides, targets, genomics, etc.), rather than simply choosing the vendor with the largest model.
If you tell me whether your focus is small molecules, biologics/antibodies, target discovery, or an end-to-end pharma platform, I can give you a more useful 2026 vendor/model landscape with model size, training data, licensing/API availability, pharma partnerships, and recommended use cases.
For genomics, Arc Institute's Evo 2 is a major model to investigate. It was trained on more than 9 trillion nucleotide tokens from genomes spanning the tree of life and is designed to model and generate biological sequences. Arc is also developing models for cellular perturbation and virtual-cell applications.
For protein engineering specifically, Profluent is another important player. Its platform combines foundation models with an in-house wet lab, and its E1 model is a protein encoder available under a permissive license for research and commercial use.
The ecosystem for biology-specific foundation models has expanded rapidly, moving from general sequence models to multimodal "world models" of biology that span genomics, pathology, and protein dynamics.[](https://www.bioptimus.com/youtube-videos/inside-bioptimus-training-and-tracking-foundation-models-for-biology)…
The ecosystem for biology-specific foundation models has expanded rapidly, moving from general sequence models to multimodal "world models" of biology that span genomics, pathology, and protein dynamics.
Providers and key players developing and deploying these specialized foundation models fall into several key categories:
If you tell me which therapeutic modality (e.g., small molecules, antibodies, gene therapy) or biological scale (e.g., genomics, proteomics, whole-cell simulation) you intend to focus on, I can help you narrow down which specific model or partnership strategy fits your pipeline best.
DNABERT
Yes. The market has moved well beyond generic LLMs: there are now several companies building **biology-native foundation models** trained on protein sequences, structures, DNA/RNA, cells, and molecular data. For pharmaceutical discovery, I’d divide the landscape into a few important groups. Provider | Representative…
Yes. The market has moved well beyond generic LLMs: there are now several companies building biology-native foundation models trained on protein sequences, structures, DNA/RNA, cells, and molecular data. For pharmaceutical discovery, I’d divide the landscape into a few important groups.
| Provider | Representative models/platform | Biological focus | Particularly useful for |
|---|---|---|---|
| EvolutionaryScale | ESM-3, ESM Cambrian | Protein sequence, structure & function | Protein design, biologics, protein engineering |
| NVIDIA | BioNeMo | Proteins, DNA/RNA, molecules, structure | Enterprise drug-discovery pipelines, customization |
| Profluent | Protein language/generative models | Protein sequences & design | Protein engineering and novel therapeutics |
| Chai Discovery | Chai-1 and newer biomolecular models | Protein–ligand and biomolecular structure | Structure prediction, molecular interactions |
| Boltz | Boltz models / Boltz Lab | Biomolecular structure and interactions | Structure prediction and molecular design |
| Basecamp Research | EDEN models | Protein/DNA biology | Protein and genetic engineering |
| Recursion Pharmaceuticals | Phenom/BioNeMo ecosystem | Cellular phenotypes, imaging, biology | Target discovery and phenotypic screening |
| Insilico Medicine | Chemistry/biology AI stack | Targets, molecules, drug design | End-to-end drug discovery |
| Schrödinger | ML/physics-based models | Structural biology & chemistry | Structure-based drug design |
1. EvolutionaryScale — strongest pure biology/protein-model option
Its ESM family is probably the most prominent example of a biology-native foundation-model family. ESM-3 is a generative model that jointly represents protein sequence, structure, and function, while ESM Cambrian focuses on protein representation learning.
This is particularly interesting if your pharmaceutical program involves antibodies, enzymes, protein therapeutics, protein engineering, or de novo protein design.
2. NVIDIA BioNeMo — strongest platform/ecosystem option
BioNeMo is somewhat different: NVIDIA isn't simply selling one biology foundation model. It provides a platform for using, fine-tuning and deploying multiple biomolecular models, including models for protein structure, molecular generation, docking, DNA/RNA and property prediction.
That makes it particularly attractive to a pharma organization that wants to bring proprietary experimental data into model development rather than simply consume an API.
NVIDIA has also substantially expanded BioNeMo in 2026, including RNAPro, ReaSyn v2 and tooling specifically intended for training/customizing biological foundation models.
3. Chai Discovery / Boltz — particularly interesting for molecular structure
These are worth evaluating if your emphasis is protein–ligand interactions, complexes, structure prediction and structure-based drug design, rather than just protein sequence modeling. NVIDIA identified both as part of its 2026 BioNeMo ecosystem.
I'd avoid treating "biology foundation model" as one category. For pharma, there are at least five distinct model families:
For example, NVIDIA's ecosystem already spans DNA, RNA, proteins, molecules and cellular data; its earlier BioNeMo models included DNABERT and scBERT, while the current platform has expanded considerably.
I would evaluate providers on six dimensions, rather than simply asking which model has the highest benchmark score:
For a pharmaceutical discovery program, my initial shortlist would therefore be EvolutionaryScale, NVIDIA BioNeMo, Chai Discovery, Boltz, Profluent, and Recursion, with the exact ranking depending heavily on whether you're doing small molecules, antibodies/biologics, target discovery, or cell-based discovery.
If you tell me which of those four areas you're targeting, I can give you a more comprehensive vendor landscape—including model names, licensing/API availability, modalities, approximate model sizes, strengths/weaknesses, and which ones are actually suitable for commercial pharma use.
2. NVIDIA BioNeMo — strongest platform/ecosystem option
BioNeMo is somewhat different: NVIDIA isn't simply selling one biology foundation model. It provides a platform for using, fine-tuning and deploying multiple biomolecular models, including models for protein structure, molecular generation, docking, DNA/RNA and property prediction.
That makes it particularly attractive to a pharma organization that wants to bring proprietary experimental data into model development rather than simply consume an API.
NVIDIA has also substantially expanded BioNeMo in 2026, including RNAPro, ReaSyn v2 and tooling specifically intended for training/customizing biological foundation models.
3. Chai Discovery / Boltz — particularly interesting for molecular structure
These are worth evaluating if your emphasis is protein–ligand interactions, complexes, structure prediction and structure-based drug design, rather than just protein sequence modeling. NVIDIA identified both as part of its 2026 BioNeMo ecosystem.
For pharmaceutical discovery, “biology-specific foundation models” generally refers to large pretrained models built on biological data (proteins, DNA/RNA sequences, single-cell data, microscopy, clinical data, molecular structures, etc.) and adapted for tasks such as target identification, protein design, antibody…
For pharmaceutical discovery, “biology-specific foundation models” generally refers to large pretrained models built on biological data (proteins, DNA/RNA sequences, single-cell data, microscopy, clinical data, molecular structures, etc.) and adapted for tasks such as target identification, protein design, antibody engineering, biomarker discovery, and drug candidate generation.
The landscape is split between model builders, platform providers, and AI-native biotech companies using proprietary models.
| Provider | Focus | Examples / capabilities |
|---|---|---|
| Google DeepMind | Protein structure and biological reasoning | AlphaFold and related models transformed protein structure prediction; used broadly in structural biology and drug discovery workflows. |
| EvolutionaryScale | Protein language models | ESM3 generates and designs proteins using sequence, structure, and function representations. Axomap Blog |
| NVIDIA | BioAI infrastructure and model ecosystem | NVIDIA BioNeMo provides pretrained models and tooling for protein, chemistry, and molecular generation workflows. NVIDIA Investor Relations |
| Amazon Web Services | Hosted bio foundation model access | aws.amazon.com provides access to biological AI models and workflows for drug discovery. Amazon News |
| Microsoft | AI-for-science models and infrastructure | Provides AI research infrastructure and partnerships for biology applications. |
| Meta Platforms | Open biological models | Developed protein-focused models such as ESMFold for fast structure prediction. |
| Model | Domain | Notes |
|---|---|---|
| AlphaFold | Protein structure | Widely used structural biology foundation. |
| ESM | Protein sequences | Learns protein sequence representations and supports structure/function prediction. |
| Geneformer | Transcriptomics | Learns gene expression patterns for cellular biology tasks. |
| scGPT | Single-cell biology | Transformer model for cell-state and perturbation modeling. |
| OpenFold | Protein structure | Open implementation inspired by AlphaFold methods. |
For a pharmaceutical discovery program, I would evaluate providers by biological modality rather than by brand:
For a pharma organization starting a generative AI discovery effort, the shortlist is usually not “which foundation model is best,” but rather which model provider can integrate with your proprietary biology, assays, and experimental feedback loops.
BoltzGen
`Major technology entities, specialized TechBio firms, and research institutes` provide biology-specific foundation models for pharmaceutical discovery. These models range from open-access structural predictors to proprietary generative suites spanning multi-omics, genomic sequences, and de novo protein design. Key…
Major technology entities, specialized TechBio firms, and research institutes provide biology-specific foundation models for pharmaceutical discovery. These models range from open-access structural predictors to proprietary generative suites spanning multi-omics, genomic sequences, and de novo protein design.
Key Providers of Biology Foundation Models
Provider Comparison
| Provider / Creator | Core Focus / Modality | Primary Application in Discovery |
|---|---|---|
| Isomorphic Labs | Multi-scale molecular structure & interaction | Target validation & de novo ligand design |
| Arc Institute (Evo) | Whole-genome DNA/RNA sequences | Genomic editing & mutation effect prediction |
| Recursion | High-throughput phenomics & imaging | High-content screening & target discovery |
| Generate Biomedicines | De novo protein design | Biologics & antibody creation |
To help narrow down the right approach for your pipeline, could you share:
For pharmaceutical discovery, “biology-specific foundation models” are being developed by a mix of **AI labs, biotech companies, cloud/platform providers, and open-source research groups**. The main providers fall into a few categories: ## Protein and biomolecular foundation model providers Provider | Key models /…
For pharmaceutical discovery, “biology-specific foundation models” are being developed by a mix of AI labs, biotech companies, cloud/platform providers, and open-source research groups. The main providers fall into a few categories:
| Provider | Key models / platforms | Primary use cases |
|---|---|---|
| Meta Platforms | ESM family (ESM-1, ESM-2, ESMFold) | Protein sequence understanding, embeddings, structure prediction, protein function prediction NVIDIA Developer |
| DeepMind | AlphaFold, AlphaFold-derived systems | Protein structure prediction, structural biology workflows NVIDIA Newsroom |
| NVIDIA | NVIDIA BioNeMo | Hosted and customizable biomolecular foundation models for proteins, molecules, docking, and generative chemistry NVIDIA |
| OpenAI | General-purpose frontier models with biology applications | Scientific reasoning, literature analysis, agentic workflows (not a dedicated protein foundation model provider) |
| Broad Institute and academic collaborators | Various genomic and biomedical models | Genomics, single-cell biology, disease biology research |
These companies generally combine foundation models with proprietary experimental data, assays, and discovery platforms:
For small-molecule discovery, important model builders include:
For an internal pharma AI program, these are commonly evaluated:
NVIDIA BioNeMo is one example of an ecosystem approach: it packages models such as AlphaFold-related tools, ESM models, protein generators, molecular generators, and docking models into a drug-discovery development environment.
For a pharmaceutical discovery program, the decision is usually less about “which model is biggest” and more about the biological modality:
A typical enterprise stack might combine:
For a pharma organization starting today, the most mature starting points are generally protein foundation models (ESM/AlphaFold/OpenFold families), NVIDIA BioNeMo, and AI-native biotech platforms that can incorporate proprietary assay data.
RoseTTAFold
The landscape of biology-specific foundation models spans open-science research labs, tech giants, and specialized tech-bio companies. These models process various biological modalities—such as amino acid sequences, genomic data, multi-omics, and 3D molecular…
The landscape of biology-specific foundation models spans open-science research labs, tech giants, and specialized tech-bio companies. These models process various biological modalities—such as amino acid sequences, genomic data, multi-omics, and 3D molecular structures.
Key providers of biology-specific foundation models are categorized below by their primary focus:
If you'd like to narrow this down, tell me: