← Back to context

Comment by ttul

1 hour ago

Most people running local models would probably love to run larger models if only they had access to big enough hardware. I'm curious: to those of you running models locally, if there was a way to inference the model of your choice at a reasonable cost by effectively time-sharing a B300 rack through some privacy-protecting intermediary, would you consider that?

If there was a "Mullvad of GPU clouds", would that solve the privacy concerns?