Comment by rao-v
5 hours ago
As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.
I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.
They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.
Its CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it.
Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is
While typical investors in their last round are subject to a five-year lock-up and will not have voting rights, China's National Artificial Intelligence Industry Investment Fund also put money into it, retaining both voting rights and freedom from the lock-up. Nothing really new if you're aware of how involved the CCP is with companies of strategic importance in China.
https://www.reuters.com/world/asia-pacific/chinas-deepseek-c...
Oh of course, you’re not getting into positions of power by not playing by the party’s rules. And if you get notions that you can tell THEM what to do you’ll be swiftly dealt with.
The company is doing well and providing great PR so the party is content to not meddle too much I imagine.
My comparison with American labs is more that I think they have to deal with bean counters, creditors, investors etc which can shuffle incentives and aims (and is a big reason why they dont do open weights anymore)
It was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).
Yeah, there are knowledge based societies, where knowledge is considered the crown jewel. Not virtue signaling plus purple hair. Not pretending to be a complete idiot, who likes to walk backwards because why not.
This is adapted from Microsoft research's YOCO. It was known for a while(2024!).
Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM.
Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".
why didn't Microsoft scale its own invention?
Politics and profits.
Deepseek delivers 1 product; Microsoft delivers dozens (or hundreds depending on how you want to count it) across various domains.
quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.
> quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.
It's quite crazy that it's Deepseek's background/original purpose. We already had very advanced stuff from the world of HFT, but now a frontier family of models from a private company that used to be (still is?) in HFT is plain bonkers.
Is more known about them and the HFT background?
According to an old interview, apparently they were always interested in AI. But finance is just where they had their first success.
> Many of High-Flyer's original team members worked on AI. Back then, we tried a lot of fields before getting our big break in finance, which is complex enough. AGI is probably one of the hardest things we can do next, so for us it was a question of how, not why.
It's a very good interview:
https://www.lesswrong.com/posts/kANyEjDDFWkhSKbcK/two-interv...
Incidentally, Wenfeng is kind of reverse Hassabis. There were some rumours that:
> Hassabis quietly assembled a team of around 20 researchers to develop high-frequency trading algorithms, without Google's approval. When the parent company found out, the project was disbanded.
https://timesofindia.indiatimes.com/technology/tech-news/whe...
[flagged]
Or maybe, the "hacker" philosophy that this site is named after, is strongly opposed to the philosophies that the American labs seem to be operating on?
anyways, remember HN rules: "Please don't post insinuations about astroturfing, shilling, brigading, foreign agents, and the like. It degrades discussion and is usually mistaken. If you're worried about abuse, email hn@ycombinator.com and we'll look at the data."
It has nothing to do with open vs closed or "hacker" philosphy. See this the announcement of the closed Seedance 2.5 - https://news.ycombinator.com/item?id=49138302
Direct quote from the second top comment:
> Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them.
Compare that with the launch of ChatGPT Image of yesterday.
1 reply →
> posts on American models are steered towards controversy and anti-AI sentiment, posts on Chinese models are full of blatant flattery
So why, for example, are posts on the Inkling[1] release (an American model) thread mostly positive? It's as if there's something else at play here, but I can't quite put my finger on it, hmm... :P
[1] -- https://news.ycombinator.com/item?id=48924912
Not Chinese/American models. We are talking about open-weight (and their detailed tech report) and close-weight (with non-sense restrictions)
Google's Gemma models are usually celebrated, so were the llamas. If Meta releases Muse Spark it will also be a good thing. If Anthropic released a great open weight model I am sure that post won't be steered towards controversy and anti-AI sentiment.
It so happens Chinese companies are more friendly towards open weights, autonomy and freedom that most US based ones. Who would have guessed?
umm what are you talking about? Basically this crowd (esp. folks like me who run medium models locally) like open stuff and can be a tiny bit unenthused about opaque mysteries handed down from on high. You'll see people delighted with Gemma releases and heck even IBM's Granite models (boring architecturally though they may be) every time they come out. Heck I was chuffed about gpt-oss-120b for weeks. @sama give us another already!
This doesn't require an influence operation.
American models are closed, expensive, neutered, and make Dario and Sam even more rich and powerful.
Chinese models are open-weight, cheap, neutered only about things like Tiananmen Square and the treatment of Uyghurs, and scare Sam and Dario.
The Uyghur thing is so weird, the number one killer of Muslims is the United States. We're supposed to hate China because they force them to go to cultural schools and assimilate, a practice countries like Norway still do to this day with migrants.
There are more people who go to church on Sundays in China than the United States. There are 10x more mosques in China than the United States.
Tiananmen square was a student revolt literally egged on by cold war western institutions, who attempted to use chinese students as pawns for geo-political games.
Westerners really need to rethink their opinions on China, it seems obvious to me they are not the ones to be worried about (although, all governments do tons of harm).
4 replies →
This post doesn't even allege this...
Weird of you to turn technical discussions into weird nationalistic debates. Maybe lay off the X algo, I think elon has oneshot your brain. .
Your source: vibes
Deepseek's source: mostly open
i wonder if there's any relationship hmmmm