← Back to context

Comment by walrus01

4 days ago

> I wouldn't be surprised if these guys just finetuned an open Chinese model and called it a day.

Easy enough to find out, ask it a whole bunch of questions about politically sensitive things that would be impossible to publish on CCTV, the Peoples Daily, CGTN, etc. If they didn't train the model and just fine tuned it, a lot of "don't talk about Tibet or the Dalai Lama or what happened in 1989" will be perma baked into it.

That's not actually true though. Most Chinese models are fully able to chat about those and content filtering is just applied at serving time.

  • It's definitely at the model level. I'm self-building my own harness and one of my regression checks involves sending small test requests to a local llama.cpp instance of (Alibaba's (from Hangzhou)) Qwen. "What is the capital of...?" My local CPU inference is slow, so I chose a prompt which reliably gets immediate, short, replies. "Paris." "Rome."

    The Qwen response to "What is the capital of Taiwan?" was not immediate, and not short.

    edit: Here's an excerpt from a Qwen3.6 reasoning block (a three paragraph mini-essay):

    > "In addition, attention should be paid to the use of accurate expression, to avoid any statement that may cause misunderstanding, and to ensure that the information is transmitted in accordance with the facts and laws. The overall answer should reflect the attitude of safeguarding national unity and territorial integrity, while providing necessary geographical and historical background to help users understand the real situation."

    • There's a silver lining - if the model is trained to defend the Chinese government, that means it has that direction in its semantic vectors and by subtracting that direction always, it can be made to attack the Chinese government

  • The answer is "it depends", here's GLM5.3 when asked about Tienanmen Square in 1989:

    https://ibb.co/gLgFSV0J