@antirez The core issue is the double standard: these companies built their AI on everyone’s shared data, including maybe some that was never meant to be scraped, yet they now behave as if that collective knowledge is their private property
Modern AI resulted from research made also by many non-US scientists (Hinton, the French folks, Linnainmaa, many others). The pre-training corpus was produced worldwide with massive code contribution from Europe OSS. What is happening with frontier LLMs is unacceptable.
# On the "hallucination problem"
I always struggle a bit with I'm asked about the "hallucination problem" in LLMs. Because, in some sense, hallucination is all LLMs do. They are dream machines.
We direct their dreams with prompts. The prompts start the dream, and based on the LLM's hazy recollection of its training documents, most of the time the result goes someplace useful.
It's only when the dreams go into deemed factually incorrect territory that we label it a "hallucination". It looks like a bug, but it's just the LLM doing what it always does.
At the other end of the extreme consider a search engine. It takes the prompt and just returns one of the most similar "training documents" it has in its database, verbatim. You could say that this search engine has a "creativity problem" - it will never respond with something new. An LLM is 100% dreaming and has the hallucination problem. A search engine is 0% dreaming and has the creativity problem.
All that said, I realize that what people *actually* mean is they don't want an LLM Assistant (a product like ChatGPT etc.) to hallucinate. An LLM Assistant is a lot more complex system than just the LLM itself, even if one is at the heart of it. There are many ways to mitigate hallcuinations in these systems - using Retrieval Augmented Generation (RAG) to more strongly anchor the dreams in real data through in-context learning is maybe the most common one. Disagreements between multiple samples, reflection, verification chains. Decoding uncertainty from activations. Tool use. All an active and very interesting areas of research.
TLDR I know I'm being super pedantic but the LLM has no "hallucination problem". Hallucination is not a bug, it is LLM's greatest feature. The LLM Assistant has a hallucination problem, and we should fix it.
</rant> Okay I feel much better now :)
You know how image generation went from blurry 32x32 texture patches to high-resolution images that are difficult to distinguish from real in roughly a snap of a finger? The same is now happening along the time axis (extending to video) and the repercussions boggle the mind just a bit. Every human becomes a director of multi-modal dreams, like the architect in Inception.
Coming back to Earth for a second, image/video generation is a perfect match for data-hungry neural nets because data is plentiful, and the pixels of each image or video are a huge source of bits (soft constraints) on the parameters of the network. When you're training giant neural nets in supervision-rich settings, your train loss = validation loss, and life is so good.
My favorite place to keep an eye on the AI video space unfold atm is probably https://t.co/l1xRaq71C4 , or the individual Discords.
AI video generation is getting shockingly good.
In a couple of years, anyone will be able to create a movie from a smartphone- we're entering a whole new era of film.
Here are some of the best examples I've seen:
Hopefully, it is now extremely obvious that Europe should restart dormant nuclear power stations and increase power output of existing ones.
This is *critical* to national and international security.
@TRENORD_miVC Incoveniente... abbiate almeno un minimo di rispetto! Già il servizio che offrite non è dei migliori, poi definite inconveniente tecnico una tragedia con delle vittime, pessimi!