“The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum.” - @benthompson
Who’s Afraid of Chinese Models?
Everyone is worried about Chinese models, but the frontier labs will be fine; we need to enable open U.S. alternatives.
https://t.co/Q5cbH229jj
Yet another lawsuit brought against US AI innovators for training AI. Chinese AI corps fall outside of US IP enforcement and enjoy zero lawsuits while US corps adopt their models, trained on US copyrighted data including from direct competitors of the US corps deploying them.
BREAKING LATE: @AnthropicAI's just been sued for direct infringement by publishers @SonyMusic and @WarnerChappell for torrenting + scraping. The suit also names D. Amodei and B. Mann for torrenting and contributory liability. A 4th DMCA count v @AnthropicAI pleads CMI removal:
two unanswered/unreported questions:
1. does the 40M or 450K reported by TR include the 650M paid for casetext?
2. what data was the model TR is using, Qwen, trained on?
Answers:
1. obviously not
2. unlicensed US copyrighted data, including from TR's direct competitors
Y'all shouldn't be surprised about this. Big part of dead internet theory(reality) is the action of nation states to coordinate these kinds attacks and it's not just on X (credit to them for publishing!)
Now that @GlobalAffairs has confirmed Chinese bot swarms have been pushing anti-US AI propaganda, and the below is true, your next question should be:
Why are Chinese propagandists so thrilled Thomson Reuters has decided to build on Chinese AI models and move off US AI models?
Chinese propaganda accounts on X are very excited that Thomson Reuters are using open source Chinese models trained on massive amounts of unlicensed US copyrighted material including from Thomson Reuters direct competitors. Says a lot.
look at this, all these accounts pushing the EXACT same messaging over thomson reuters choice to rebuild on chinese AI models. no paid advertisement disclosure either @elonmusk
generative AI was not invented recently for those interested, this was from 2016. we've all certainly come a long way but AI training remains AI training
“Distill OpenAI and Anthropic’s models, then give their competitors the ability to do something they could likely never attempt in the U.S. without facing legal consequences.”
@ReutersChina Distill OpenAI and Anthropic’s models, then give their competitors the ability to do something they could likely never attempt in the U.S. without facing legal consequences.
Chinese propaganda accounts on X are very excited that Thomson Reuters are using open source Chinese models trained on massive amounts of unlicensed US copyrighted material including from Thomson Reuters direct competitors. Says a lot.
chinese open models are built different*
* chinese open models are built outside the US IP regime and trained on vast amounts of unlicensed, unpaid US copyrighted data including from direct competitors of US companies who post train on them. no litigation/licensing = cheaper
Thomson-1.0 is built on Qwen under Apache 2.0: use, modification, and redistribution are all permitted. TR released its derivative under PolyForm Strict: none of the three.
TR's card says frontier AI shouldn't belong to a few funded players. TR's license choice says it should.
Maybe I'm missing something, but the Thomson-1.0 model license doesn't seem to permit any commercial use of any kind, including self-hosted use by a law firm to support its lawyers.
TR's decision to launch on Alibaba Qwen seems to directly contradict its published Data and AI Ethics Principles. Qwen is trained on US copyrighted data, including from direct TR competitors. Now TR is using Qwen despite its published data principles and its stance in TR v. ROSS.
From Alibaba Qwen's own training data summary: “Datasets may include copyrighted, trademarked, or patented content, as well as public domain content.”
Link: https://t.co/ZcoxPaRpG5
With one hand, Thomson Reuters drives up the cost of US AI tokens. While with the other hand, TR reaches for Chinese AI models not subject to the legal rule it's pushing in courts, trained on copyrighted work, including data from its direct competitors, because it's cheaper.
I'm sympathetic to what I understand of ROSS' position: the headnotes were only used internally for training, and this isn't fundamentally different from studying a competitors' product.
But I'd suggest some of the claims you make here are very broad and sort of non sequitur. For example:
> no training data market gets created in the US at all
All licensing stops because everyone has access to open source models?
What is it that ROSS wants here and how is it better?
You say it will make tokens more expensive. But what about higher order considerations?
For example, what good are cheap tokens if the whole market structure that supports the production of high-quality content upstream that those tokens are generated from collapses?
What if affording an inconsistency across jurisdictions means that:
1. Cos with proprietary data are able to capture more value via their own post-trained, self-hosted models
2. Frontier labs are not able to become monopolies
3. So tokens generally become cheaper
4. While also supporting a healthy market structure that incentivizes quality content creation the models need?