Everyone in politics wants what they want. And when Donald Trump doesn't personally share the same enthusiasm, they feal betrayed.
Take a step back and think.
The United States is the most powerful Empire ever to exist.
Trump is the rare example of someone who actually wants something for his country. He is a Great Man. You do not know what he knows.
Chimp all you want, but realize that you are dumb and have no idea what's happening.
And just know that it is the United States vs Evil.
@sdcat2@Foxfire1st@Sebastian__One@amitisinvesting China has released good foundational models. Didn't say they didn't. I use Chinese open source models daily.
What did I say wrong about transformers?
1. I have no idea what you think I said, but I will make it clear. The system prompt for LLMs that you interact with (instruct models) is written in such a way that works with the model's training. That's obvious. You wouldn't use Claude's system prompt on ChatGPT. If you use API (Open Router, etc) calls or if you load your own model, you get to define that system prompt.
2. I obviously wasn't clear enough in my initial reply, so I'll clear that up too. I did not say that the model was trained with the system prompt built in. The understanding of the system prompt is built in however. In that same way, a model may reject certain tasks because it was trained that way. It could also be trained in a subversive malicious way as well. Extremely simple to poison a model's ability to code.
3. A US company may download the weights for a model off of HF and do what they will with it, yes. But, if that model is an instruct model, that gets extremely difficult to change. You would really have to know what model weights had been shifted around by the instruct training. If the model is the base model, then yes, it's completely up to that US company.
@Foxfire1st@amitisinvesting Not sure what you are disagreeing with me on here, again.
If you believe I'm saying that the weights themselves are somehow going to send my data to the CCP, that isn't what I was trying to convey.
Not sure what you are disagreeing with here.
For models like Kimi releasing their 2.8T model, we haven't gotten their base model for that yet. I don't think we will either, because I don't think Moonshot has released many base models recently. (I haven't kept up here, doesn't matter too much either)
But, if you are talking about the weights from the pre-trained instruct models, then you aren't going to change them meaningfully because instruct is already built in.
And you said that system prompts have nothing to do with weights. That isnt true. A model that you talk to, ChatGPT, Claude, etc., is using an instruct layer that the system prompt is written to interact with.
@thetronchguy@amitisinvesting When a foreign country's lab blatantly rips off one of our frontier models, what should I say?
And if nothing is done about it, why do you think OpenAI or Anthropic would even bother making new models if they can't have protection from their government?
Don't know what you are getting at. You think companies make models for fun, regardless of anyone using them?
Usage stats is how they, and other companies, know that their product is good.
Also, when people start seeing Chinese models on the leader board of top AI benchmarks, what exactly do you think their reaction will be?
As for backdoors, it's extremely simple. Every model that you interact with is Question/Answer. The dataset is typically trained on or finetuned in such a way that it can answer your questions and speak in a way that you can understand. So, a dataset can be made in such a way that the AI *can* output code that is insecure. It's simple, the AI wouldn't know that it's doing it, since it's literally just it, and there's the backdoor.
@thetronchguy@amitisinvesting Usage stats, mainly.
Makes China look good, makes people say "wow! China is close!" As well.
But, to the point of the post, companies shouldn't be able to use Chinese models due to backdoors.
I'm not sure if I know what you are trying to say.
1. Models are trained in such a way that allows for backdoors. Think an LLM that follows system prompts. They are actually trained with that backdoor in them. That's how backdoors work. LLMs can be trained in such a way to make very, very small "mistakes" or flaws in code that allow for easy access (if someone knows what they are looking for) in, say, a web app.
2. Not sure what you are getting at with the human data. The same human data that the American models trained on is available to most everyone, aside from the stuff they paid for.
This guy does not get it.
1. The AI companies did illegally download a lot of text and books. However, a lot of it was in a legal grey area because of their publishing + most of that text is still downloadable by any 3rd party, such as yourself. However, you can not train your own model using that text and what makes a model usable (the question/answer portion) is largely from paid for sources such as Reddit and Qoura.
2. Since you, and obviously some Chinese companies, cannot train their own model given most of the same information (excluding what Anthropic, OpenAI, Google, etc. has purchased) since you do not have the same compute that these companies have paid billions nor do you have the talent these companies have created.
There is no question that these AI companies have illegally stolen text and data. However, the scale is simply not the same and by making it the same, you actually defraud more and tread on more companies and people.
Yes, they paid for those tokens. But when you use the API service (or any service) from Claude, it specifically says that you cannot use it to train or distill your own models.
Same for ChatGPT and their frontier models.
By using those services, you agree that you won't train or distill a model.
As for the other services, for the most part, sign away most of your data just by using that service.
I don't like it when my data is stolen and used for anything at all. But they were American companies that in many cases actually fairly bought your data and trained on it.
These Chinese companies are directly using our frontier model companies and training their own models to compete.
It's unlucky that US companies don't release huge frontier models, but how would they? It literally takes the smartest minds and millions of dollars to train these models. These Chinese companies skip that and just steal.
That's fair enough. But that doesn't go against my argument.
See, some of these Chinese companies distill off our models and actually steal from our companies.
If you trained a model and it was one of the best ever, and China uses your model to create their own, how would you like that?