top of page

A CRISPR-Like Pattern in Phage DNA, Found by Claude AI Agents

5 minutes ago
6 min read
A CRISPR-Like Pattern in Phage DNA, Found by Claude AI Agents
A CRISPR-Like Pattern in Phage DNA, Found by Claude AI Agents

Anthropic reported on September 23, 2026 that Claude agents, given a high-level prompt to search a massive database of DNA sequences for interesting new reverse transcriptases (RTs), identified a previously uncharacterized enzyme system. The search ran for 21 hours, used roughly 950 agents and consumed 210 million tokens. Near the end, one agent flagged a repeating pattern of DNA sequences next to the gene for an unusual RT. Anthropic's scientists followed up in the lab and named the system array-associated reverse transcriptase, or ART.


The function of ART is unknown. The announcement describes early findings from one of the company's first research programs, and a pre-print carries the technical detail.


What the Agents Found


ART is found mainly in bacteriophages, the viruses that infect bacteria. It has three parts: an RT, a partner gene beside it, and a long array of evenly spaced DNA repeat sequences. The array resembles a CRISPR array, which holds a bank of different RNA sequences and is what makes CRISPR-Cas systems programmable. Anthropic's first experiments show that the ART array is expressed as a set of distinct short RNAs, which the team reads as a sign that something analogous to CRISPR may be happening.


The RT at the center of the system had already appeared in earlier studies, having been found in a jumbo phage. According to Anthropic, Claude appears to be the first to notice the system's defining features, an associated array of non-coding DNA sequences and an accessory protein of unknown function.


The team places ART alongside a small group of known systems. This combination of characteristics has been found together in only a handful of other systems, all of which are programmable and perform operations such as cutting, copying and pasting DNA. Several of those systems are now in development as tools. ART has not been shown to be programmable, and the announcement does not claim that it is.


Why Genome Mining Matters Here


The post opens with three earlier discoveries that began with someone noticing something odd in natural DNA or proteins. Restriction enzymes came from bacterial immune systems and made it possible to cut DNA at chosen places and splice genes between organisms. Taq polymerase came from a bacterium in a Yellowstone hot spring and became the basis for PCR. CRISPR was first seen as an unusual repeat sequence in certain bacteria and now underlies gene-editing medicines.


Reverse transcriptases have followed a similar path in recent years. Researchers have found many more of them, most in bacteria, where they act as part of the immune system. Nearly all of these families were found by genome mining, in which researchers search sequence databases for uncharacterized genes, pick out the unusual ones and work out what they do. That process depends on a person noticing an anomaly, and the ART result is an example of an agent doing the noticing.


How the Search Worked


The scientists gave Claude a prompt to look through the database for interesting new RTs. Claude agents then combed the sequences, investigated the distinct RT families and used their own judgment to pick candidates. Anthropic reports that the agents gathered over 200,000 RTs, picked out 3,500 new candidate systems and narrowed those to the 20 most compelling, each analyzed and written up as a human-readable report. The post says that an expert scientist would need weeks to months for this kind of analysis.


The agent that found ART was reading raw DNA next to the RT when it saw a tandem repeat array by eye and described the sequence as "spectacular." It then worked through the checks a human researcher would run. It counted the repeats, measured their spacing, compared the layout with known RT systems and searched the literature for any earlier report of the pattern. Once it was convinced the system was new, it filed a report for human review.


Anthropic's scientists then took over. They tested the candidate in their lab, and the analysis that followed led them to describe it as a previously uncharacterized system.


The Workflow Behind the Result


The post describes a general pattern the group uses for surveys of protein families. Claude first reads the relevant literature and reproduces established results from public data, which checks that its methods work. It then searches for family members or genomic neighbors that fit no described system. For each candidate it writes a short report that proposes a function and lays out the supporting evidence.


A second round of analysis follows, in which Claude critically evaluates that evidence. Most candidates are eliminated here, and a survey may end with a single candidate worth testing or with none. Candidates that survive go to the lab, where the protein is expressed in standard laboratory strains and characterized biochemically and structurally, with Claude helping to interpret the data. The team works in Claude Science and Claude Code, the tools available to any scientist, and sometimes uses its own harness to coordinate many Claude sessions running in parallel.


The volume of hypotheses has become a subject of study in its own right. A single campaign can produce hundreds to thousands of candidate reports, and the group has been examining what separates the proposals it judges worth testing from those it sets aside. What it learns is fed back into the instructions given to Claude, with the aim of teaching the model to reflect the team's own scientific judgment.


The Lab and the Team


The lab is in the Bay Area and looks like a typical molecular biology lab. Its research is limited to the lower biosafety levels, BSL-1 and BSL-2, and the team does not handle pathogens that can infect humans. Human scientists do all of the lab work. Anthropic says it has experimented with using AI to speed up lab work through initiatives such as the Model Hardware Standard, but found that approach less suited to the ad hoc workflows of its molecular biology research.


The team members have spent their careers on unusual proteins and on computational methods for reading DNA and interpreting its evolution. Their earlier work covered the evolution and regulation of CRISPR systems, new enzymes for cell and gene therapies, and tools for spotting anomalies in DNA such as human pathogenic variants. They sit within Anthropic's life sciences organization, next to teams working on drug discovery and on training Claude in biology and chemistry.


What Remains Open


Anthropic says experiments to determine how ART works are underway, and it chose to publish before they finish. The company's stated reasons are to show what Claude can do and to give the wider community a view of its work. The central question, what ART does, has no answer yet. The evidence so far is the structure of the system and the expression of the array as short RNAs.


Feng Zhang, a CRISPR genome editing pioneer and a professor at MIT and the Broad Institute, reviewed the pre-print. He called it an example of AI agents contributing to biological discovery, said the RNA-repeat arrays associated with reverse transcriptases are intriguing and merit further investigation, and said he hoped the work would lead more scientists to explore AI in their research. His comment is a statement of interest in the finding and does not verify a function for ART.


Readers should also keep in mind that the announcement comes from the group that built and ran the system. The post reports that 20 candidates reached the report stage and that ART emerged from further analysis and lab testing, but it does not say how many of the other reports were tested or what they showed. The post gives the token count but no dollar cost.


Practical Implications


For researchers with sequence data on protein families of unknown function, the post documents a workflow that others can examine and adapt: reproduce published results first, search for neighbors that fit no known system, require a written rationale for each candidate, and put every candidate through a round of critical evaluation before any lab work. Anthropic reports that most candidates fail that evaluation, and a workflow built this way should expect the same.


The company has also invited outside scientists to propose research questions, in genomics and other fields, for collaboration. Anyone considering that route should note the division of labor described in the post. Claude agents did the searching, analysis and report writing, and humans did the lab work and the judgment about which candidates to test.


The next thing to watch is the result of the ongoing experiments on ART's function. The pre-print is the place to check the methods and the evidence for the array's expression, and any follow-up will show whether the CRISPR comparison holds up.

 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page