Comment by venusenvy47

6 months ago

The big players use parallel processing of multiple users to keep the GPUs and memory filled as much as possible during the inference they are providing to users. They can make use of the fact that they have a fairly steady stream of requests coming into their data centers at all times. This article describes some of how this is accomplished.