Biodiversity AI Raises $140M to Deliver Gene Therapy Without Factories

September 23, 2026:

Biodiversity AI Raises $140M to Deliver Gene Therapy Without Factories
Biodiversity AI Raises $140M to Deliver Gene Therapy Without Factories
GIUSEPPE CACACE/AFP via Getty Images

Basecamp Research, a London- and Boston-based AI biotech startup, closed its oversubscribed Series C for $140 million today, using an approach most pharmaceutical AI companies have never attempted: training its models not on human biology, but on the genetic diversity of more than one million newly sequenced species from rainforest soils, Antarctic ocean floors, and volcanic terrain across 30-plus countries. The company’s EDEN biological foundation models are then applied to design a new class of gene-insertion enzymes — large serine recombinases — that could, in theory, reprogram a patient’s immune cells directly inside the body, eliminating the complex and costly manufacturing chain behind current approved cancer cell therapies that run $400,000 or more per patient.

That $400,000 figure is not a sticker price in the traditional sense. It is what the laboratory work costs before a hospital bill enters the picture — and it is why a substantial share of blood cancer patients who qualify for existing CAR-T treatments die before cells are ready on manufacturing waitlists. Basecamp’s pitch is that its platform, if it works in humans, removes the factory from the equation entirely.

The round was led by S32 venture capital, the California venture firm founded by Bill Maris, who previously built Google Ventures into one of the most successful corporate venture funds in Silicon Valley. Andy Conrad, former Verily CEO joining board as S32 General Partner, comes aboard as part of the deal. Participants include Anthropic’s Anthology Fund, Nvidia’s NVentures arm (a returning investor after a January 2026 SAFE note), the NATO Innovation Fund, the UK’s Sovereign AI Fund, The Rockefeller Foundation, King Philanthropies, and senior personal investors from the biopharma world including André Hoffmann, vice-chairman of Roche. The round brings Basecamp’s total capital raised to $225 million and values the company at approximately $800 million.

Why Mainstream Biological AI Has a Blind Spot

The public databases that underpin almost every competing biological AI model share a structural problem Basecamp has spent six years trying to solve. According to Basecamp’s own analysis, roughly 68 percent of the sequence volume in the Sequence Read Archive — the dominant global repository for genomic data — comes from just five species, with humans alone accounting for around 54 percent. As CTO Philip Lorenz has put it, training only on public genome data is like teaching a language model to understand all human communication using only newspaper articles from 1975.

The practical consequence is that any chemistry nature has evolved outside those five species — the antimicrobial compounds produced by competitive soil bacteria, the precise gene-splicing enzymes in bacteriophages attacking extreme-environment microbes, the protein structures used by organisms no laboratory has ever studied — simply doesn’t exist in the training data. EDEN was built to close that gap.

Basecamp’s Trillion Gene Atlas, the proprietary dataset at the center of EDEN’s training stack, currently holds approximately 15 trillion tokens of DNA sequence — broadly comparable in scale to the text corpora used to train frontier large language models, but comprising DNA building blocks rather than words. The company has set a target of one quadrillion tokens within the next 18 months, drawing on field partnerships across more than 30 countries on all seven continents.

The first generation of EDEN models was trained on 9.7 trillion nucleotide tokens representing genetic material from more than one million previously unknown species, with parameters reaching 28 billion — a scale comparable to some of the largest publicly released biological models. In internal tests, Basecamp found that proprietary data produced steeper scaling than public databases: as the model grew larger, its biological task performance improved more rapidly with Basecamp’s data than with publicly available genomic sequence collections.

Lorenz noted a counterintuitive finding from those training experiments: a model architecture with lower perplexity scores on standard AI benchmarks performed worse on biological tasks than a Llama-based architecture that scored higher on those standard metrics. The company now uses biological task performance at regular training checkpoints as its primary evaluation signal — effectively disregarding the standard proxy metrics that the broader AI field uses to compare models.

What Is a Large Serine Recombinase, and Why Does It Matter for Cancer Treatment?

The core technical bet at the center of Basecamp’s therapeutic pipeline is a class of enzymes called large serine recombinases (LSRs) — molecular machines that bacteriophages, the viruses that infect bacteria, evolved over billions of years to insert their own genetic material into bacterial genomes with extraordinary precision.

Unlike CRISPR-Cas9, which makes a double-strand cut in DNA and relies on the cell’s own error-prone repair machinery to incorporate new genetic material, LSRs insert DNA without double-strand breaks. The mechanism works through site-specific recombination: the recombinase recognizes two short attachment sequences — one on the donor DNA it is delivering and one on the target chromosome — and catalyzes a precise strand exchange, covalently linking broken ends and religating them in the correct configuration. The reaction is unidirectional without reversal protein and can deliver large payloads of DNA, potentially entire gene sequences spanning thousands of base pairs, at a single specified location.

The challenge is that naturally occurring LSRs were not designed for human cells. Many target sequences that exist only in bacteria or bacteriophage genomes; others integrate into human DNA but at too many different locations to be safe or useful therapeutically. The practical workaround — which the Arc Institute’s research group demonstrated in published gene therapy work and which Basecamp is now approaching via AI design — involves engineering or selecting LSRs that either target a specific naturally occurring sequence in the human genome, or can be preceded by the installation of a short “landing pad” sequence at the desired location.

Basecamp’s EDEN models are trained to learn from the enormous evolutionary diversity of bacteriophage-bacteria interactions — precisely the source of the most potent and precise naturally occurring LSRs. In laboratory experiments reported in its preprint, the company found that EDEN-designed recombinases showed T-cell activity in primary human T cells, with 50 percent of tested variants showing activity. In separate in vitro assays designed to model tumor-cell killing, the EDEN-designed CAR constructs cleared over 90 percent of tumors.

These are laboratory results — in vitro, meaning outside a living body. No Basecamp therapeutic program has yet entered animal efficacy studies, the step before human clinical trials can begin.

From Molecules to Medicines: What the Pipeline Actually Contains

Basecamp’s therapeutic pipeline spans six programs, all currently in lead optimization — the earliest formal stage of pre-clinical drug development, in which promising molecular candidates are chemically refined before animal testing begins. The company has not published specific data on all six programs.

The programs as disclosed include:

In vivo CAR-T cell therapies for blood cancer and an unnamed autoimmune disease. These are the programs most directly tied to the large serine recombinase platform — they aim to install CAR constructs into T cells while those cells remain inside the patient’s body, avoiding the manufacturing chain entirely.

A liver gene therapy for phenylketonuria (PKU), an inherited metabolic disorder in which the body cannot properly break down the amino acid phenylalanine, requiring patients to follow extremely restrictive low-protein diets for life.

Antimicrobial peptides targeting drug-resistant bacterial pathogens, including strains on WHO’s critical-priority list — organisms against which no reliably effective antibiotic currently exists.

Two earlier-stage programs: a CAR-T therapy for solid tumors and peptides targeting diabetes.

The realistic timeline from lead optimization to patient access, through animal efficacy and safety studies, IND application filing, Phase 1 human safety trials, Phase 2 and Phase 3 efficacy trials, and regulatory review, is a minimum of five to ten years for any individual program — and the majority of drug candidates that enter clinical trials do not complete them. The company plans to enter clinical trials for its most advanced programs, beginning with the in vivo CAR-T programs, within the next two to three years.

EDEN-7: An AI-Designed Antibiotic That Killed a Last-Resort Pathogen in Mice

Perhaps the most striking published result to date in Basecamp’s research program involves not cell therapy but antibiotic design. In the preprint accompanying the EDEN model family, Basecamp researchers report that 97 percent of AI-designed antimicrobial peptides showed biological activity against target pathogens in laboratory tests, conducted in collaboration with researchers at the University of Pennsylvania.

More specifically, a candidate peptide called EDEN-7, generated directly by the model without iterative human refinement, matched last-resort antibiotic efficacy in mouse models infected with multidrug-resistant bacteria. The work targets a public health problem of enormous scale: antimicrobial resistance killed 1.27 million people in 2019, according to the landmark GRAM study published in The Lancet, and the WHO projects the toll could rise to 10 million annual deaths by 2050 if new treatments are not developed.

“We prompt on a pathogen and then the model designs an antibiotic that kills it,” Lorenz told The Decoder in January 2026.

The mechanism draws directly on the same evolutionary logic that underpins the cell therapy programs. Bacteria in competitive soil ecosystems evolve potent antimicrobial molecules to defeat rivals; the genetic diversity of those interactions, encoded in the environmental DNA Basecamp samples globally, gives EDEN a far larger repertoire of functional designs to learn from than any model trained on laboratory culture collections.

Selected EDEN antibiotic and vaccine target design capabilities have been integrated into Claude Science platform, allowing researchers to describe a pathogen in plain language and receive AI-generated molecular candidates.

Where Is This Data Actually From, and Who Gets Paid?

The Trillion Gene Atlas draws on genetic samples collected by Basecamp through access and benefit-sharing (ABS) agreements with researchers, local communities, and national governments in more than 30 countries — a legal framework governed by the Nagoya Protocol on genetic resources, a supplementary agreement to the Convention on Biological Diversity adopted in 2010 and entered into force in 2014.

Under the Nagoya Protocol framework, parties providing genetic resources are entitled to prior informed consent and mutually agreed benefit-sharing terms before those resources may be commercially used. Basecamp says it has paid 52 beneficiaries across 19 countries as of the end of 2024, with some payment timelines as short as nine months from sampling to first distribution — an unusually rapid cycle in an industry where royalty payments have historically reached source countries, if at all, only after product commercialization.

The company’s benefit-sharing structure has also attracted criticism. A 2025 Financial Times report cited researcher Jim Thomas, who argued that payments amounting to roughly one percent of revenue do not adequately reflect the commercial value that biodiversity data generates for the company. A representative of Cameroon’s environment ministry called publicly for higher revenue shares.

Basecamp has pointed to its Cameroon partnership signed in 2024, covering four communities with provisions for local scientific training and laboratory support, as a model for what equitable field collaboration can look like. The company also screens sequences it generates against known pathogen databases and excludes certain viral genetic material from training, though Lorenz has acknowledged these measures are not a complete safeguard against misuse.

How Does Basecamp’s Approach Compare to Open-Source Competitors?

Basecamp’s most directly comparable peer in the biological foundation model space is the Arc Institute, whose Evo 2 model — developed in collaboration with Nvidia, Stanford University, UC Berkeley, and UC San Francisco — trained on 128,000 genomes across life representing all three domains of life. Evo 2 is publicly available via Nvidia’s BioNeMo platform.

Basecamp has taken the opposite strategy: EDEN is fully proprietary, with access gated behind pharmaceutical partnerships and the Claude Science platform. The company cites two reasons: biosafety concerns (limiting access to sequences with potential for misuse), and commitments to its data partners, many of whom prefer a model whose commercial use can be tracked and from which they can receive a revenue share.

The dual-use concern is not hypothetical. A May 2025 paper published in PLOS Computational Biology dual-use AI study by researchers at Johns Hopkins, Stanford, and Carnegie Mellon explicitly framed biological AI models as presenting dual-use capabilities analogous to the dual-use dilemmas in traditional life sciences research — tools that accelerate legitimate drug discovery while also potentially lowering barriers to engineering harmful agents. Anthropic itself, as Basecamp’s investor and platform partner, released an AI policy framework on biological weapons in June 2026 specifically addressing biological weapons as a priority risk category for AI governance.

Basecamp’s proprietary approach means independent researchers cannot audit what the model will and will not generate — a transparency trade-off that is not unique to Basecamp but is more significant when the training data is environmental microbiome samples including phage sequences selected specifically for their potency.

What Comes Next

Basecamp has appointed Richard Pearce, formerly of Biogen, as Chief Business Officer to build out pharmaceutical partnerships — a signal that the company intends to operate both as a platform licensing EDEN models and data to biopharma clients and as a drug developer advancing its own pipeline.

With $225 million in total capital, an investor base spanning Silicon Valley venture funds, AI model developers, defense alliances, sovereign wealth, and philanthropic institutions, and a technical team whose wet lab operations in Cambridge, Massachusetts are running reinforcement learning experiments on live human T cells around the clock, Basecamp has assembled an unusual array of institutional support for a company that has not yet dosed its first patient.

CEO Glen Gowers put the company’s goal plainly in the announcement statement: “We believe the future of medicine lies in reprogramming the body to repair itself. We design the models and the medicines to teach it how.”

The biology has produced results in the lab that warrant serious attention. Large serine recombinases, if they can be designed with sufficient precision and delivered effectively in vivo, represent a mechanistically cleaner approach to gene insertion than anything currently in widespread clinical use. The antimicrobial peptide results, if they replicate in larger animal models and eventually in humans, address one of the most urgent unmet needs in global medicine. The biodiversity dataset represents a genuinely different kind of training resource from anything a public-domain model can access.

The gap between laboratory validation and clinical approval remains wide, and it is a gap measured in years and billions of dollars, not months. The $140 million raised today begins to close that distance.


Frequently Asked Questions

What is in vivo cell therapy, and why does it cost so much less than current CAR-T?

Current approved CAR-T cancer therapies require extracting a patient’s T cells, shipping them to a central manufacturing facility, genetically engineering them using viral vectors, expanding them in bioreactors for up to five weeks, then shipping them back. That process involves more than 200 hours of specialized labor per batch and can cost upward of $400,000 before hospital charges. In vivo cell therapy — what Basecamp is developing — aims to deliver the genetic modification directly to T cells inside the patient’s body, skipping the manufacturing chain. If it works, researchers at institutions including Dana-Farber Cancer Institute estimate costs could fall by more than 97 percent. Basecamp’s in vivo programs are still in pre-clinical development and have not been tested in humans.

What makes large serine recombinases different from CRISPR for gene editing?

CRISPR-Cas9 makes a double-strand cut in DNA and relies on the cell’s repair machinery to insert new genetic material, which can cause off-target errors, deletions, and chromosomal rearrangements. Large serine recombinases insert DNA without making double-strand breaks: they recognize specific short sequences on both the incoming DNA and the target chromosome and catalyze a precise strand exchange without relying on cellular repair. This makes them particularly suited to delivering large payloads — full gene sequences rather than small edits — at defined genomic locations. They evolved in bacteriophages for exactly this purpose.

How does antimicrobial resistance connect to what Basecamp is doing?

Antimicrobial resistance — the ability of bacteria to survive antibiotics — caused 1.27 million deaths directly in 2019, more than HIV or malaria, according to a 2022 Lancet study. With projections of up to 10 million annual deaths by 2050 if no new treatments emerge, AI-designed antimicrobial peptides represent one potential route to new antibiotics that classical drug discovery has struggled to produce. Competitive soil bacteria have evolved potent antimicrobial molecules to defeat rivals over billions of years; Basecamp’s training data samples that evolutionary arsenal on a global scale. The company’s EDEN-7 peptide, designed by AI without human refinement, matched a last-resort antibiotic in mouse models — though, as with all the company’s programs, this result requires clinical validation before any conclusions about patient benefit can be drawn.

Is it ethical to build a commercial product from other countries’ biodiversity data?

This is actively contested. The 2010 Nagoya Protocol requires prior informed consent and benefit-sharing agreements before genetic resources from signatory countries can be commercially used. Basecamp says it has structured all of its partnerships to comply with these requirements, making payments within nine months of sampling and building local scientific capacity in partner countries. Critics, including researcher Jim Thomas (as reported by the Financial Times), have argued that approximately one percent of revenue is inadequate compensation for the commercial value these genetic resources create, and representatives of some partner governments have publicly called for larger shares. There is no industry-wide standard for what constitutes equitable compensation in this emerging category.

Source link