← Back to context

Comment by eru

6 days ago

Agreed.

If you wanted to and had enough engineering effort to spare, you could run an LLM deterministically at relatively small impacts to performance.

One approach is to make sure you run things in the same order. Another is to change your operations so that more of them become associative or even commutative.

See eg the paper 'A Lattice-Based Approach to Deterministic Parallelism' for some interesting ideas on the latter.