Rumor is that in the last few days Sama and Dario were calling for a slowdown in AI because there was some the parameter to tweak to scale AI models that scales similar to number of weights, amount of training data and compute. This would be huge, these "big 3" scaling parameters were untouched since basically the beginning of AI decades ago?
I also head this breakthrough was made at Google, which wousld explain their current silence. Another rumor is Google has achieved a new method of RSI (Recursive elf Improvement).
China can't do it because it' incompatible with distillation
Imagine your outcomes being bound to which passport you have obtained or on what land you were born. Wouldn't that be the overlord's dream? "The people I control are just just more intelligent"? lol!!!
Counterpoint 1: Deepseek R3 moment, Deepseek R4 moment, Kimi K3 moment. Maybe we'll have a Qwen4 moment next? GLM6 moment? Minimax M4? Hy4?
Counterpoint 2: I once worked in a company where a QA trainee had to go back to China for family trouble and then disappeared off the face of the earth. 4 months later there was a new system on the market, made in Shenzhen, with very particular design choices I recognized, because I spec'd them. Are we saying Google doesn't have any non-American employees?
I strongly recommend against nationalist arrogance in decision making around what your perceived adversary can and cannot do. Underestimating
humaningenuity is the number one cause of being caught with your pants down. And I say this as someone that is not at all charmed with the Chinese model.What does any of this have to do with the topic? They claim this new approach is unavailable to China. Why is it unavailable to China? I speculate it's because it's incompatible with distillation.
Oh. I skipped the step that the whole "China only has distillation" is an American hopium narrative. Similar to "China doesn't have the GPUs required" (beaten by Deepseek) or "China doesn't have the traces" (beaten by Kimi) or "China doesn't have the optimizations" (beaten by Deepseek 2nd time). Every time someone claims that China is inferior and cannot do something, then gives a reason for the why, it turns out to be an arrogant statement that doesn't age well.
Have you used Kimi K3? Is it really inferior for you? It was for me at least equal to Fable 5 in both depth and usability for the times that I have used it. Maybe I'm missing a point here and there are use cases where it is poor, but all these things Chinese labs couldn't do are narrated by reducing the humanity of the researchers there to "China", then making up reasons why Chinese tech is inferior, then claiming that this is why USA wins. But is that really true? What's left of the 10 year headstart? Seems like it is 2-3 months now.
Also, whenever OpenAI and Anthropic researchers proudly claim that their LLMs are doing most of the work for the next iteration, they in all their joy are forgetting that if Kimi K3 == Fable 5, then the next iteration of Kimi has no reason to be less well trained as the one that Fable 5 trained for Anthropic, or GPT 6 for OpenAI. Worse for everyone wanting to dominate the world by setting a research pace, Kimi K3 is open, so basically that training capability is now a commodity.
Thus, what I am warning about is that the nationalism of political leaders that need votes and the reality of what everyone's capabilities truly are, may either be already not be true, or otherwise turn out really easy to overcome.
You seem to be under the impression that distillation was a bad thing?
This is not the case. This is not a "hopium" narrative. Quite the opposite: it's a pro-china narrative that makes the US situation look bleak because of how much better this strategy is
The narrative I hear out there - maybe this isn't what you personally mean - and maybe I am interpreting this wrong, is:
All that I'm saying is that if the new tech is just some proprietary Google software then there is no moat, because software can be reverse engineered in a few days now, thanks to LLMs. See for example this repo that reverse engineered the unpublished Kimi K3 training pipeline
If there is some Nvidia or Goog hardware needed as well, then the moat will last for a few months at most; the Chinese hardware manufacturers can go from lab prototype to production at scale much faster than anyone else.
Therefore, I reject any claim of "China cannot do this". Not because I am pro-China, but because I am extremely skeptical of location/geopolitical based superiority of intellect and attitude. I think I know this to be a lie, as I've had the honor of working with people in almost every culture thinkable and the one thing we all excel at is overcoming barriers, if only we want to overcome it.
China can't do it because this "axis" doesn't exist
Trump says full-speed ahead. Pausers are losers.
Thanks for finding us something to read on what went down that isn't paywalled and anti-archive!
Altman, Dario et al:
"Chat, how do we make people not think we are bad guys before going pUbLiK???"
Chat:
"Fantastic idea. Right now, you are not too popular--it is good to acknowledge your shortcomings. I would play on base human emotions of crowds. Go mass hysteria. Tell them that you built a technology that can take over the world, and spin it in a way that makes you look like the good guys. RSI is your best angle here. You probably want to try some type of astro-turfing on obscure online forums--that should really stir up their insecurities. The key will be emphasizing the potential danger while asserting you in control.
"Do you want me to develop a 10-step astro-turfing campaign to target all the most popular internat discussion forums?"
I think its more "oh shit china models are closing gap....how do we keep pumping stock!??!"
Chat:
"Simple. tell them you are so close to ending world you decided to pause development. You. Are. Just. That. Good"
Fake. Double hyphen detected instead of em dash, a clear fingerprint of hooman writing.
Haha!
https://twiiit.com/teortaxesTex/status/2099055661937996180