Comment by Der_Einzige
17 hours ago
There is for $/creativity/token. LLM sampling settings are poorly supported even in open source serverless providers but are the single best lever you have for getting better outputs in regards to creativity (and quality for long context or highly quantized models).
I'm pretty sure you can adjust the creativity for many Chinese model inference providers.
Most of them don't expose more than top_p/top_k/temperature. Those are woefully inadequate compared to what open source inference engines support.