GPTs and Hallucination : "... the more obscure or controversial the subject matter, the more likely it was that the result would be incorrect..." https://t.co/aJNNfY2z8x
🚨 [AI RESEARCH] "AI models collapse when trained on recursively generated data" by @iliaishacked, @Zakobian, @aaronzhao123, @NicolasPapernot, @yaringal & Ross Anderson is a MUST-READ for everyone in AI. Quotes & comments:
"The development of LLMs is very involved and requires large quantities of training data. Yet, although current LLMs2,4–6, including GPT-3, were trained on predominantly human-generated text, this may change. If the training data of most future models are also scraped from the web, then they will inevitably train on data produced by their predecessors. In this paper, we investigate what happens when text produced by, for example, a version of GPT forms most of the training dataset of following models. What happens to GPT generations GPT-{n} as n increases? We discover that indiscriminately learning from data produced by other models causes ‘model collapse’—a degenerative process whereby, over time, models forget the true underlying data distribution, even in the absence of a shift in the distribution over time"
-
"Our evaluation suggests a ‘first mover advantage’ when it comes to training models such as LLMs. In our work, we demonstrate that training on samples from another generative model can induce a distribution shift, which—over time—causes model collapse. This in turn causes the model to misperceive the underlying learning task. To sustain learning over a long period of time, we need to make sure that access to the original data source is preserved and that further data not generated by LLMs remain available over time. The need to distinguish data generated by LLMs from other data raises questions about the provenance of content that is crawled from the Internet: it is unclear how content generated by LLMs can be tracked at scale. (...)"
-
➡ As I've been discussing in my newsletter, most general-purpose AI providers rely on legitimate interest as the lawful grounds to process data - including personal data - to train their AI models & systems. When attempting to justify their legitimate interest grounds, these providers often cite the collective & societal benefits to be reaped from their AI. I disagree with the legitimate interest argument from a data protection perspective (links to my newsletter article below). But leaving the data protection context aside for a moment, there is a general interest in preserving the quality of those models from ethical, environmental, and fairness perspectives (at least), not to mention the loss of human, technical & computational resources in case most existing models collapse. Perhaps regulation should also ensure that AI training, as a rule, preserves the sustainability of present and future models.
➡ Link to the paper below.
🔥 To stay up to date with the latest developments in AI policy, compliance & regulation, including excellent research, join 32,300 people who subscribe to my newsletter (link below).
Après ma première vidéo à 5,8 millions de vues, Bardella a répondu, et du coup je me suis intéressé à ce qu'il a dit. Merci de continuer à ne pas participer à sa diffusion ici, ça agace le petit Jordan.
Cette vidéo a fait 5,6 millions de vues entre tiktok et instagram, au point que Bardella s'est senti obligé de faire une réponse. Merci donc de ne pas participer à sa diffusion ici aussi.
C'était avant-hier.
L'incroyable collectif @nosservicespub publiait une somme de plus de 300 pages, retraçant l'évolution des services publics dans les 10 à 40 dernières années
Deux jours plus tard, petit défi : résumer ces 300 pages en moins de 12 tweets.
C'est parti ⤵️
Hedi, 22 ans, a été touché par un tir de LBD, roué de coups et laissé pour mort par la BAC de Marseille. Une partie de son crâne a dû être retiré pour le sauver. Il témoigne et raconte comment il envisage l’avenir.
There's a new programming language in town - it's Mojo! I'm more than a little excited about it. It's Python, but with none of Python's problems.
You can write code as fast as C, and deploy small standalone applications like C.
My post is below, and a 🧵
https://t.co/0IWGqcpEY7
On se demandait de quand pouvait dater l'expression "y'a moyen de moyenner", le @Wiktionnaire m'a indiqué... 1640. Visage choqué.
Et en suivant la référence, j'ai bien retrouvé ça dans "Curiositez françoises": https://t.co/ueeYU9lLNg
Relire le chapitre de Freakonomics qui attribue la chute de la criminalité à NY dans les 90's, non pas à la politique de Guiliani mais à Roe vs Wade, avec cette citation d'un haut gradé du bronx "the only effective crime-prevention device adopted in this nation since the 60's"
In 1980, the same year Dr. Slump debuted as a manga, Akira Toriyama designed an instruction manual for an Aiwa cassette recorder. It is as delightful as you can imagine. Photos courtesy of @CnZBYF1ctAj1R23