← Back to context

Comment by John7878781

1 day ago

Yep. Sequence-to-function models are still very limited. AlphaGenome Atlas, despite the flashy branding, is unlikely to provide significant benefit to researchers.

May be that's the reason for the alpha naming. We are waiting for a stable release (just kidding).

And it makes a lot of sense why they are limited. DNA is not an instruction set. It's more like a heavily encrypted dataset where the encryption key is the totality of physics and biology. The interactions with the physical world that result in the end product of life are enormously (it would seem hopelessly) complex.

For a machine intelligence to turn DNA sequences into organisim phenotype prediction requires modelling all that in latent space.

I imagine that is going to take a monumental amount of example data

  • It seems like we can skip much of the expensive modelling and use evolutionary conservation data to shortcut building a full latent space that captures all salient interactions. It's unclear to me whether we truly need to model the entire latent space. And given that biology develops in a generative way with feedback, it may be that attempting to model this using a static latent space is unproductive.