Comment by Spacecosmonaut

16 hours ago

Current evolved Cas9 (CRISPR) variants are highly efficient and relatively unconstrained in terms of their human genome targeting coverage. Smaller nucleases and higher targeting specificity would be useful. But therapeutic use is mostly limited by delivery.

This seems revolve around a known retron-like reverse transcriptase. A sober framing would be something like: Claude identified a previously undescribed genomic arrangement around a known reverse transcriptase. Not all that sexy.

For now, this is mostly a story about how AI can be used to parse existing data to discover new biology (which is fantastic!).

I've been using Claude Science a lot and it is VERY good at finding patterns in the DNA around my binding sites - quite often it went 'you could put your primer here but that looks like an Alu repeat, so better not, the primer won't be specific' - it seems like the press release is one step above that pattern recognition? I.e., 'there's a recurring motif here that hasn't been described before', which is probably straightforward to pick up when your context window is 1 million tokens, i.e. within the range of entire bacterial genomes...

  • Isn't finding this out from an LLM somewhat... complex and non reproducible.

    • Everyone who survived 7 rounds of multi model reviews and they still keep finding mediums in their PRs is not in the least surprised. These things are not oracles - they miss stuff all the time even when told to look.

      1 reply →

    • Yep :) it's VERY non-reproducible - but I haven't asked it to look for Alu repeats in the first place, I can then go and reproduce the work it's doing

  • Do you find it a big improvement over the tools you used previously?

    • It's honestly harder - you have to do a lot of extra work to ensure your work is traceable. It's super fast in doing things you have no overview of - it makes pretty figures, it runs command line jobs etc. and I'm sure there are mistakes in there. For now I'm pretending Claude Science is an IDE, like Positron/VSCode, and I have to keep enforcing proper git usage etc. so I can reproduce this work

      Edit: compared to my tools before, it generally uses the same tools in the same way, just 20x faster than me and I mostly struggle to keep up and verify what it's doing

      2 replies →

> For now, this is mostly a story about how AI can be used to parse existing data to discover new biology (which is fantastic!).

I'd like to expand that: in my view, this is also a story of how agentic AI systems can come up with bioinformatics strategies to discover novel features. One would think such a task would be the ideal domain of the genome language models, which have learned the structure and functional relationships of DNA/RNA sequences. The agents instead relied on classical bioinformatics methods such as HMMs to make their discovery.

Note: I could not find the Supplementary Note 1 that was supposed to describe how exactly agents came to their solution, but I assume it was autonomous.

This is a good summary of what was going on. I kept reading the paper hoping for a cool wrinkle or function to be revealed, but it's just conserved, highly transcribed array sitting next to reverse transcriptases with a few possible partner genes.

A side note, Matt Durrant has hit on some pretty exciting recombinase activity previously (https://www.nature.com/articles/s41586-024-07552-4). If there's anyone who's well equipped to track down if ART is doing something cool, he's top of the list.

  • Sorry, but isn't a "conserved, highly transcribed array sitting next to reverse transcriptases" in itself the description of an unknown mechanism? If two parts are combined and conserved and we know what each means but not why they're combined and conserved then it's pretty intriguing, no?

Exactly the same mechanics was about astra decoding enigma encoded message: it's well-researched subject, with bunch of data and LLM created a breakthrough by identifying previously missed pattern/relation.

But its just PR so far. They haven't published a refereed science paper, in say Nature or Science. At this stage, its of little value to others until verified.

  • I guess the Poincare conjecture or the theory of relativity will continue to have "little value" until they get published in a peer reviewed journal

[flagged]

  • People involved in Anthropic will be catapulted to a new level of wealth for sure. The problem is the regular Joe investing his savings in Anthropic, thinking he is going to be catapulted as well ...

    • There's still time for them to be left holding the bag, much like OpenAI has been forced to for the time being. Even if people here believe that the economic activity in this sector doesn't represent a bubble, at least they can see the warning signs as a result of trade and literal war, 10-year yields are back over 5%, oil and diesel are going sky high, and there's no quick fix to any of it even if our leaders were willing and capable of trying.

      So personally I understand why Anthropic is only concerned about finding bag holders rather than ethics, decency, legality, responsibility, humanity, or a modicum of thought beyond their own selfish desires.

      7 replies →

    • Let's not pretend it's the regular joe that put all their savings into $FOO. It's degen gamblers that want those thousand percent gains.

  • Amen. One would hope that these companies, if they truly want to engage in scientific research, would pursue established routes in announcing results and having them validated/refereed independently. But no, this is PR. Its akin to former announcements of "cold fusion", until assessed and verified independently.