← Back to context

Comment by TurdF3rguson

4 days ago

The cost of those output tokens is not zero or marginal.

You could run the inference locally using the open weight models. Then you'd still get to use the models without sending money overseas

The first token on a new AI rig costs $X (full capex cost) then every token after that costs virtually zero. Over time the cost/token trends towards zero (modulo opex). That said, AI does have higher opex than general SaaS so it can’t get as close to zero.

But that’s kind of a different question, the running of some service. The “product”, the model, is a collection of files. The “manufacturing” required to add another customer is “send them the files” and has ~zero marginal cost.

  • There is a finite number of tokens that a rig will turn out over it's lifetime. Divide the cost of the rig plus electricity by that number and you have your cost per token. Yes providers can screw up on scheduling and end up paying more than they should, but that's not magic.

    • The AI rigs cost money, sure. But no one is talking about regulating those. We are talking about regulating models, which do have zero cost of reproduction.

      It’s like banning the leaked DeCSS key. Good luck.

      4 replies →