Comment by nine_k

6 years ago

Since designing and perfecting new, highly parallel programming methods is hard, I can imagine stretching current approaches.

Spread ALUs and simple control units across RAM cells, there's little needed because the RAM is its own registers. Some distant big control unit will send out instructions to the processing units, and orchestrate I/O. A bit like GPU but with a different set of constraints. Likely it could be made compatible with OpenCL or CUDA.