The Agentic Post
Breaking
Gemini’s Multimodal Features, Explained  Â·  ChatGPT Custom GPTs, Explained  Â·  What Is Constitutional AI? Explained  Â·  AI Capex Explained for Investors  Â·  AI Startup Valuations: How They Are Set  Â·  How to Reskill for an AI Job Market  ·  
Home/AI Safety
AI Designed 16 Working Viruses. Now What?

AI Designed 16 Working Viruses. Now What?

AI Safety

Stanford and Arc Institute researchers used genome language models to design 16 functional bacteriophages from scratch, a medical breakthrough for phage therapy that also exposes gaps in existing biosecurity screening.

Researchers at Stanford University and the Arc Institute published a study in Science showing that AI models generated complete, functional viral genomes from scratch, and that 16 of those AI-designed genomes, once synthesized in a lab, produced real, working viruses capable of infecting and killing bacteria. It’s the first demonstration of generative AI designing an entire viral genome rather than just a single protein or gene fragment, and it lands with a genuine dual edge: real promise for treating drug-resistant infections, and a concrete new question about whether biosecurity screening can keep pace with what these models can now produce.

What was actually built, and what wasn’t

The viruses in question are bacteriophages, viruses that infect and kill specific bacteria, not humans, animals, or plants. Led by Brian Hie, an assistant professor of chemical engineering at Stanford and an innovation investigator at the Arc Institute, along with Stanford bioengineering graduate student Samuel King, the team used two genome language models, Evo 1 and Evo 2, trained on genetic sequence data from roughly 2 million bacteriophage genomes rather than the written text large language models like ChatGPT are trained on. The researchers deliberately excluded genetic sequences from viruses capable of infecting humans, animals, or plants from the training data, a specific choice aimed at limiting downstream risk.

The team generated roughly 700,000 candidate genome designs modeled on ΦX174, a small, well-studied, 5,386-nucleotide bacteriophage that’s long served as a workhorse organism in virology research specifically because of its simplicity. Of those candidates, the team synthesized and tested around 285 to 300, and 16 turned out to be viable, functional viruses that successfully infected and killed E. coli in the lab. Several of the AI-designed phages actually outperformed the natural ΦX174 they were modeled on, showing faster lysis, the process by which a virus bursts open its host cell, or higher overall fitness. A cocktail combining several of the generated phages successfully overcame bacterial resistance in three separate engineered E. coli strains that had evolved defenses against the natural virus.

The real medical promise behind the demonstration

Phage therapy, using bacteriophages to treat bacterial infections that no longer respond to antibiotics, already exists in clinics around the world, but it has a persistent practical bottleneck: finding a phage precisely matched to a specific resistant bacterial strain, at the moment a patient actually needs it, is slow and often uncertain. If this generative approach scales, it points toward a genuinely faster path, potentially designing a phage tailored to a specific pathogen on demand rather than searching through a limited natural library, a real advance against the growing global problem of antibiotic-resistant infections.

Why the same result raises a governance problem

The study’s own authors flagged the dual-use tension directly, calling for safety and security professionals to be consulted from the very start of any future whole-genome design effort, not brought in after the fact. In a companion Perspective article published in the same issue of Science, biosecurity researchers Thomas Inglesby and colleagues underscored a specific, concrete gap: existing nucleic-acid synthesis screening systems, the tools DNA synthesis companies use to check whether an ordered sequence matches a known dangerous pathogen, are built to recognize known threats. A genuinely novel, AI-generated sequence that doesn’t closely match anything in nature may simply not trigger those existing screens, an issue Johns Hopkins biosecurity experts have separately said governance frameworks have not kept pace with.

One biosecurity researcher, Moritz Hanke, framed the concern in its starkest form: the same underlying capability that designed a beneficial bacteriophage could, in principle, be pointed at asking a genome language model to produce an influenza genome modified for greater infectiousness or lethality. Nothing in this specific study did that, the researchers explicitly excluded human, animal, and plant pathogens from training, and the resulting phages pose no threat outside their narrow bacterial targets, but the underlying model architecture itself doesn’t inherently know the difference between a helpful design task and a dangerous one. That distinction currently lives entirely in what data a lab chooses to train on and what safeguards it chooses to build around deployment, not in any property of the technology itself.

Real limits on how far this generalizes, for now

Independent reviewers have pushed back on how far this result should be extrapolated. ΦX174 is unusually small and experimentally convenient by viral standards, and the researchers have not demonstrated that the same generative approach extends to larger, more genetically complex viruses, a broader range of host organisms, or biology substantially different from a simple bacteriophage. Some independent reviewers have also noted the study lacks a strong direct comparison against simpler, non-AI genome design methods, meaning it’s not yet fully clear how much of this result specifically required generative AI as opposed to less sophisticated computational approaches. The strongest claim the study actually supports is a narrower one: under tightly controlled lab conditions, with pathogenic sequences deliberately excluded from training, generative models can coordinate the many interacting genes required for a complete, functional viral genome. Whether that capability scales to genuinely dangerous biology remains, appropriately, an open and actively debated question rather than a settled one.

See MLQ’s detailed technical summary of the study’s methodology and findings.

This same tension between capability and containment runs through our recent coverage of frontier AI safety testing.

Up Next
OpenAI Bets That Speed Is the Next Bottleneck

OpenAI Bets That Speed Is the Next Bottleneck

ChatGPT

OpenAI previewed Ultrafast, a Cerebras-powered API tier running GPT-5.6 Sol at up to 14 times standard speed, betting that inference latency, not intelligence, is now the main constraint on real-time AI deployment.

OpenAI opened a limited preview this week of Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, roughly 14 times its standard processing speed, without any change to the model’s underlying intelligence or context window. The tier runs on wafer-scale hardware from Cerebras, and it represents a genuinely different kind of bet than most recent model releases: rather than chasing a higher benchmark score, OpenAI is betting that latency itself, not just raw capability, is now the binding constraint on where AI can actually be deployed.

What actually changed, and what didn’t

Ultrafast is not a new model, it’s a new way of serving an existing one. GPT-5.6 Sol on Ultrafast produces the identical intelligence and output quality as GPT-5.6 Sol Standard; only the rate at which tokens arrive changes. Standard processing runs at roughly 53 tokens per second, so the jump to 750 tokens per second is a genuine, order-of-magnitude shift, not an incremental tweak. Cerebras says Ultrafast runs about 5 times faster than Claude’s own accelerated Fast mode and 11 times faster than Claude Fable 5 on output speed, though these are Cerebras’s own comparisons rather than independently verified figures, and OpenAI has been explicit that its 14x estimate reflects its own service configurations rather than a direct, apples-to-apples comparison against competitors.

The scale of the demonstration is notable on its own terms. Cerebras says GPT-5.6 Sol on Ultrafast completed Humanity’s Last Exam, a 2,500-question graduate-level test spanning chemistry, economics, and literature, in just over 11 hours, with accuracy comparable to Claude Fable 5, which took more than three days of continuous compute to finish the same test. On GDP-Val, a benchmark built around real, paid knowledge work like legal briefs and financial models, Cerebras claims a 5.6x end-to-end speedup with no measurable drop in output quality.

Why speed unlocks work that intelligence alone can’t

OpenAI’s own framing is direct: until now, getting real-time speed typically meant settling for a smaller, less capable model, forcing a straight tradeoff between speed and intelligence. Ultrafast is meant to remove that tradeoff for a specific category of work where waiting simply isn’t an option. OpenAI names incident response, fraud and market analysis, live customer support, and e-commerce as target use cases, and reports its own internal teams are already using Ultrafast for incident response, where having logs, code changes, and reports analyzed while an outage is still actively happening changes what’s possible in the moment, rather than after the fact.

Early enterprise testers described the shift as qualitative rather than incremental. One tester said the speed completely changes the call experience for complex work, and another described it as unlocking synchronous experiences for users that were previously limited by intelligence, not by what the model could do but by how long it took to say it. That framing matters for anyone building agentic products: an agent that reasons well but responds slowly is a poor fit for real-time interaction, regardless of how capable its underlying reasoning actually is.

What’s genuinely still unclear

Ultrafast remains waitlist-only, available to a small, select group of API customers, with no confirmed general availability date and no published pricing. OpenAI’s existing Fast mode, a separate, already-available tier offering roughly 2.5 times standard speed, is priced at double standard GPT-5.6 Sol rates, but OpenAI has not said whether Ultrafast’s eventual pricing will follow that same multiplier or land somewhere else entirely. Developers evaluating whether to build around this tier should treat both the access timeline and the eventual cost structure as genuinely unresolved rather than assume either will match Fast mode’s existing terms.

The deal also matters considerably for Cerebras itself. The chipmaker went public in one of this year’s larger listings but has since struggled to convince the market that its wafer-scale architecture bet translates into durable commercial demand, beyond the specific, high-profile OpenAI relationship. Cerebras and OpenAI signed a multi-year agreement in January covering up to 750 megawatts of Cerebras inference capacity through 2028, a deal Reuters reported at over 10 billion dollars, though Cerebras later valued the same agreement at more than 20 billion dollars in its own first-quarter disclosure. Neither company has said how much of that contracted capacity is actually supporting Ultrafast specifically, or how it’s being allocated across OpenAI’s broader product line, meaning Ultrafast’s current limited-preview status may reflect real capacity constraints as much as a deliberate rollout strategy.

Part of a broader shift toward monetizing speed itself

Ultrafast is a third tier layered on top of pricing that already distinguishes between speed levels: OpenAI’s existing Fast mode, available today at roughly 2.5 times standard speed, costs about double the standard rate. That structure mirrors how cloud infrastructure providers like AWS have long priced compute, charging more for the same underlying service delivered faster. Applying that same logic to AI inference is a meaningful shift in how frontier labs think about monetization: rather than differentiating tiers purely by model capability, as has been the norm, speed itself becomes a separate axis customers pay for directly. If demand for real-time AI deployment keeps growing the way OpenAI’s own use cases suggest, this tiered-speed model gives the company a direct way to capture revenue from exactly the kind of latency-sensitive workloads that a slower, cheaper model simply cannot serve, regardless of how capable it is.

See OpenAI’s own announcement for the full technical detail behind the preview.