pull down to refresh

Rumor is that in the last few days Sama and Dario were calling for a slowdown in AI because there was some the parameter to tweak to scale AI models that scales similar to number of weights, amount of training data and compute. This would be huge, these "big 3" scaling parameters were untouched since basically the beginning of AI decades ago?

I also head this breakthrough was made at Google, which wousld explain their current silence. Another rumor is Google has achieved a new method of RSI (Recursive elf Improvement).

China can't do it because it' incompatible with distillation

reply

Imagine your outcomes being bound to which passport you have obtained or on what land you were born. Wouldn't that be the overlord's dream? "The people I control are just just more intelligent"? lol!!!

Counterpoint 1: Deepseek R3 moment, Deepseek R4 moment, Kimi K3 moment. Maybe we'll have a Qwen4 moment next? GLM6 moment? Minimax M4? Hy4?

Counterpoint 2: I once worked in a company where a QA trainee had to go back to China for family trouble and then disappeared off the face of the earth. 4 months later there was a new system on the market, made in Shenzhen, with very particular design choices I recognized, because I spec'd them. Are we saying Google doesn't have any non-American employees?

I strongly recommend against nationalist arrogance in decision making around what your perceived adversary can and cannot do. Underestimating human ingenuity is the number one cause of being caught with your pants down. And I say this as someone that is not at all charmed with the Chinese model.

reply

What does any of this have to do with the topic? They claim this new approach is unavailable to China. Why is it unavailable to China? I speculate it's because it's incompatible with distillation.

reply

Oh. I skipped the step that the whole "China only has distillation" is an American hopium narrative. Similar to "China doesn't have the GPUs required" (beaten by Deepseek) or "China doesn't have the traces" (beaten by Kimi) or "China doesn't have the optimizations" (beaten by Deepseek 2nd time). Every time someone claims that China is inferior and cannot do something, then gives a reason for the why, it turns out to be an arrogant statement that doesn't age well.

Have you used Kimi K3? Is it really inferior for you? It was for me at least equal to Fable 5 in both depth and usability for the times that I have used it. Maybe I'm missing a point here and there are use cases where it is poor, but all these things Chinese labs couldn't do are narrated by reducing the humanity of the researchers there to "China", then making up reasons why Chinese tech is inferior, then claiming that this is why USA wins. But is that really true? What's left of the 10 year headstart? Seems like it is 2-3 months now.

Also, whenever OpenAI and Anthropic researchers proudly claim that their LLMs are doing most of the work for the next iteration, they in all their joy are forgetting that if Kimi K3 == Fable 5, then the next iteration of Kimi has no reason to be less well trained as the one that Fable 5 trained for Anthropic, or GPT 6 for OpenAI. Worse for everyone wanting to dominate the world by setting a research pace, Kimi K3 is open, so basically that training capability is now a commodity.

Thus, what I am warning about is that the nationalism of political leaders that need votes and the reality of what everyone's capabilities truly are, may either be already not be true, or otherwise turn out really easy to overcome.

reply
146 sats \ 1 reply \ @kilianbuhn 17h

You seem to be under the impression that distillation was a bad thing?

This is not the case. This is not a "hopium" narrative. Quite the opposite: it's a pro-china narrative that makes the US situation look bleak because of how much better this strategy is

reply
178 sats \ 0 replies \ @optimism 14h

The narrative I hear out there - maybe this isn't what you personally mean - and maybe I am interpreting this wrong, is:

  1. It's all distillation (which is true for some but not all Chinese-made models, after all Kimi K2 was from scratch, and iirc GLM 5.0 was too)
  2. Therefore the source of all LLMs is the US labs
  3. Therefore if the US innovates on the scratch-to-base layer, the Chinese labs are helpless

All that I'm saying is that if the new tech is just some proprietary Google software then there is no moat, because software can be reverse engineered in a few days now, thanks to LLMs. See for example this repo that reverse engineered the unpublished Kimi K3 training pipeline

If there is some Nvidia or Goog hardware needed as well, then the moat will last for a few months at most; the Chinese hardware manufacturers can go from lab prototype to production at scale much faster than anyone else.

Therefore, I reject any claim of "China cannot do this". Not because I am pro-China, but because I am extremely skeptical of location/geopolitical based superiority of intellect and attitude. I think I know this to be a lie, as I've had the honor of working with people in almost every culture thinkable and the one thing we all excel at is overcoming barriers, if only we want to overcome it.

reply

China can't do it because this "axis" doesn't exist

reply
147 sats \ 1 reply \ @freetx 13 Sep

Trump says full-speed ahead. Pausers are losers.

From the Irish Open in Doonbeg on Sunday, Trump dismissed Saturday's "pace the frontier" pile-on from Dario Amodei, Sam Altman, and Elon Musk. Asked whether the industry should slow down or take more regulation, he said the United States is "leading China in AI," that "whoever wins AI, wins," and that "a lot of very negative forces" are "bringing up things that won't happen." Guardrails were fine in theory. A pause was not.

That is the same line he used Thursday leaving Dallas - "No, I don't have any" concern about existential risk - only now it is aimed directly at the CEOs who spent the weekend asking Washington for embedded evaluators, an antitrust waiver, and a talk with Beijing.

Obama went the other way.

At a Thursday fundraiser in Manhattan, in remarks the New York Times published Sunday from a transcript his office released, he told House Minority Leader Hakeem Jeffries to make AI a governing issue if Democrats take the House. "Once you are speaker, I would strongly urge that the Democrats put together a framework for a very public conversation." Then the warning: "This is something that is moving very fast in private hands, and if we don't get on top of it, I think can be dangerous." Benefits too - drugs, clean energy - if they do. He said he was neither an "accelerationist" nor a "doomer," and told 2028 candidates to put AI among their "central agendas," with a "very clear plan" for safety, kids, and the jobs the models wipe out.
source
reply

Thanks for finding us something to read on what went down that isn't paywalled and anti-archive!

reply

Altman, Dario et al:

"Chat, how do we make people not think we are bad guys before going pUbLiK???"

Chat:

"Fantastic idea. Right now, you are not too popular--it is good to acknowledge your shortcomings. I would play on base human emotions of crowds. Go mass hysteria. Tell them that you built a technology that can take over the world, and spin it in a way that makes you look like the good guys. RSI is your best angle here. You probably want to try some type of astro-turfing on obscure online forums--that should really stir up their insecurities. The key will be emphasizing the potential danger while asserting you in control.

"Do you want me to develop a 10-step astro-turfing campaign to target all the most popular internat discussion forums?"

reply

I think its more "oh shit china models are closing gap....how do we keep pumping stock!??!"

Chat:

"Simple. tell them you are so close to ending world you decided to pause development. You. Are. Just. That. Good"

reply

Fake. Double hyphen detected instead of em dash, a clear fingerprint of hooman writing.

reply
Something that China inherently can not do

Haha!

reply