Comment by unrented7977
14 hours ago
Using a wiki for this is only one step above using stone tablets and messenger pigeons.
You'd want an enormous vector database at minimum. Text is just completely wrong for models at this scale, you must work in the latent space directly.
A global distributed network of vector databases. The web for agents but there's no text. Only vectors and knowledge graphs kind of like graph rag. Web 4.0 which is vector based and makes this arxiv paper come true and decentralized like the internet or the text based www?
Every existing text web server can voluntarily offer a vector version of their website and charge for it or inject advertisements onto their content so the ai labs don't have to do it all themselves as model pretraining off a dataset. The vector version can have many links to other websites in the knowledge graph. This would be decentralized so not a monopoly and everyone not just ai labs would contribute to ai development because the dataset would be open because it would come from the internet itself as it already is (except for synthetic or user data).
Humanity has built every capable institution, every scientific discovery, every technological innovation, without mind-reading. Quite the opposite: By putting thoughts from our latent spaces into words, we better separated the good arguments from the bad ones, placed trust where it was more deserved, and learned to stand on giant's shoulders so well we thought thoughts the giants never would have.
Anyway, I'm much more optimistic about the worlds with AIs that make the effort to properly explain themselves.
I'm convinced that piles of markdown and effective search (which may involve vectors) is all you need.