Editorial illustration for Basecamp's AI Training Dataset to Reach One Quadrillion Tokens
Basecamp's AI Reaches 1 Quadrillion Token Dataset
Basecamp's AI Training Dataset to Reach One Quadrillion Tokens
Basecamp Research has raised $140 million in a round led by investor S32, with Nvidia, Anthropic's Anthology Fund, the NATO Innovation Fund, and Redalpine also putting in money. The London company, founded in 2020, doesn't train its AI on scraped text or images. It trains on DNA collected from rainforests, oceans, and hot springs, using genetic material from microorganisms most researchers never bother sequencing. That data feeds a family of models called EDEN, which the company uses to design new antibiotics and biological tools capable of inserting genes at precise locations inside the human body.
The pitch is aimed at two of medicine's harder problems: drug-resistant bacteria and cell therapy. Early lab work and mouse studies suggest individual drug candidates can kill multidrug-resistant bacteria, though human trials haven't started, so safety and effectiveness in people remain unproven. Basecamp says the new funding will go toward expanding EDEN and pushing its own therapy candidates toward clinical development, starting with cell therapies that reprogram a patient's cells to fight disease like cancer.
CTO Philip Lorenz spoke with THE DECODER about why biology poses a much harder modeling problem than language, and why a model that scores well on paper doesn't always produce a molecule that works.
Lorenz sees EDEN's antibiotic designs as early proof of medical value. "We prompt on a pathogen and then the model designs an antibiotic that kills it," he says.
Why this matters
The bet here is that biology, not text scraped from the internet, is the next frontier for scaling laws, and Basecamp's investor list, Nvidia and Anthropic's Anthology Fund among them, suggests the big AI players believe it too. A hundredfold jump to one quadrillion tokens in eighteen months is an aggressive target, and the Trillion Gene Atlas partnership with Anthropic, Nvidia, PacBio, and Ultima Genomics reads like an attempt to industrialize genetic sequencing the way GPU clusters industrialized language modeling. For researchers, the interesting question isn't just scale, it's whether more environmental DNA actually produces better antibiotics and gene-insertion tools, or whether Basecamp is running the same "bigger dataset, bigger breakthrough" playbook that's had mixed results in language models.
Founders building in biotech should watch whether EDEN's outputs translate into actual lab validation, not just larger training runs. Tokens counted from microorganisms are not the same thing as therapies that work in humans, and that gap is where this story will actually get tested.
Common Questions Answered
What makes Basecamp Research's AI training dataset different from other AI companies?
Basecamp Research trains its AI models on DNA collected from microorganisms in rainforests, oceans, and hot springs rather than scraped text or images from the internet. This biological approach to training data represents a fundamentally different scaling strategy compared to traditional language models, positioning genetic material as the next frontier for AI development.
How does the EDEN model family apply AI to medical research?
The EDEN models are designed to create novel antibiotics by analyzing pathogenic organisms and generating antibiotic compounds that can kill them. According to Basecamp Research, this represents early proof that their biological AI training approach has direct medical value in combating infectious diseases.
What is the significance of Basecamp's one quadrillion token target?
Basecamp aims to scale its AI training dataset to one quadrillion tokens in eighteen months, representing a hundredfold increase that reflects the company's belief that biological data, not internet text, is the key to advancing AI scaling laws. This aggressive target has attracted major investors including Nvidia and Anthropic's Anthology Fund, suggesting industry confidence in this biological approach.
Who are the key investors backing Basecamp Research's $140 million funding round?
The funding round was led by S32, with significant participation from Nvidia, Anthropic's Anthology Fund, the NATO Innovation Fund, and Redalpine. This investor composition highlights major AI and technology players' belief in Basecamp's biological approach to AI training data.
What is the Trillion Gene Atlas partnership and why does it matter?
The Trillion Gene Atlas is a partnership between Basecamp, Anthropic, Nvidia, PacBio, and Ultima Genomics designed to industrialize genetic sequencing at scale. This collaboration represents an attempt to systematize and accelerate genetic data collection in the same way GPU clusters revolutionized AI training infrastructure.
Further Reading
- Basecamp Research raises $140M to advance AI-designed therapeutics - Yahoo Finance / PR Newswire
- AI-based drug developer Basecamp valued at $800 million after $140 million funding round - Reuters
- Anthropic-, Nvidia-Backed Basecamp Research Raise $140M Series C Financing Toward Advancing AI-Designed Drugs - GEN - Genetic Engineering & Biotechnology News
- Basecamp Research launches world-first AI models for programmable gene insertion - PR Newswire
- AI data under the microscope: Accelerating and securing the AI data supply chain for the health and biopharmaceutical sectors - Atlantic Council