Higher data quality is key to efficient AI training. Did you know you can use AI to improve your data quality and help train another AI? Here's a great example.
FineWeb released a special version of their general dataset focused on high educational content—FineWeb-Edu. This version is 10 times smaller but much more efficient for training your LLMs. How is that possible?
Despite the general dataset being cleaned and deduplicated English web data, it’s still quite large—15 trillion tokens. FineWeb managed to reduce it to 1.3 trillion tokens without losing training capability, and even improving it. They achieved this using AI.
Here is the recipe. First, FineWeb had an LLM AI annotate 500k samples, scoring each for educational quality on a scale from 0 to 5 (no Elon's tweets survived). Using these annotations, they trained a specialized classifier AI model. This classifier then processed the general dataset cooking it for 6,000 H100 GPU hours. Finally, with this filtered data they trained another LLM.
The results are impressive:
- FineWeb-Edu surpasses the general FineWeb and all other open web datasets, with remarkable improvements on educational benchmarks.
- It achieves the same performance with significantly less data, requiring 10x fewer tokens, which means less compute for training and much lower costs.
The key takeaway: quality data is crucial for AI education.
So yes, surprise-surprise, we need artificial teachers to train artificial intelligence models.
When internal combustion engine were created in early 19st century it was a complete black box. You just put fuel in and you get moving force produced out of it. Sounds ridiculous, right?
This is the situation with modern AI in the early 21st century. You send input, get output, and no one knows how it happens. Imagine looking inside an engine with 132 billion gears randomly moving. Would you understand it?
And that is why Anthropic's latest research is groundbreaking in AI. They found a system in the billions of neural connections in their LLM. They discovered that entities and concepts aren't tied to specific single neurons but are instead spread across whole network. This means no single neuron represents Scarlett Johansson; instead, groups of neurons do. And a neuron can be a part of multiple groups.
With this knowledge, you can find neuron groups and related concepts, observing how the model thinks and views the world.
Even more fascinating, you can manipulate the model's thinking process to focus on specific topics or answers, even breaking rules to produce harmful texts or reveal secrets.
Researchers say this requires vast compute power, but the possibilities are invaluable. I believe that entity mapping might become essential for universal industry-grade models.
The tools and knowledge acquired during this process might also shed light on how human brains work.
Amazing, right?
LLM infrastructure is enormously capital-intensive. All that compute required to run these models is quite expensive (hello Nvidia) and consumes tons of energy. Mark Zuckerberg is even talking about building 1GW data centers— that's a lot!
In response to this problem, there's a rising trend in improving the efficiency of language models.
Welcome SLMs!
Yes, SLM means Small Language Model. It's small, which means you need less compute and, therefore, less power to run it.
Apple’s OpenELM and Microsoft’s Phi-3 focus mainly on cost and size efficiency.
If we expect LLMs to become bigger, we end up with an AGI. SLMs are the opposite—they shrink, trying to stay as smart as possible. In the near future, we might be able to run SLMs on our phones.
Key takeaway: LLMs are big and expensive; SLMs try to deliver a similar quality level while being smaller and more cost-efficient.
#AI #LLM #SLM
AI can learn from examples even without changing the model itself.
Fine-tuning is the way to make an AI model produce better results in a specific knowledge area like legal texts or medical diagnosis. It is done by retraining it on an additional dataset, and therefore, may be costly.
How could we evade it?
A new study from Carnegie Mellon and Tel Aviv University shows that ICL can give similar results or even outperform fine-tuning.
ICL — In-Context Learning is the approach to raise the quality of model output by filling the context with high-quality examples related to a specific request. (How to do that is a different question; there are several approaches.)
So the thing to remember is simple: A bigger context means you can include more examples, and that will give you better results.
Or even simpler:
LLM context: the bigger — the better.
https://t.co/rqkhGQOt6S
We've hit the Inflated Expectations Peak on the LLMs/GenAI tech adoption curve. The AI market has blown up, flooded with hype typical of groundbreaking tech. Next up, we'll dip and rise again as most of the users are only about to be onboarded in during the next Enlightenment Scope stage.
Illustration from here: https://t.co/NgEeaXuZ5t
AGI is inevitable. But when?
To train better AI today, you need to feed it more data. Once you've given it all the world's literature, you're limited in options: generate artificial training data (which is usually lower quality) or improve the AI's architecture, making it more efficient. Better architecture means less data is needed for the same results, or with the same amount of data, it can produce even better results.
It's challenging to train good AI in areas with a small amount of data. But as AI becomes smarter with improved architecture and more compute power, it will be able to excel in areas with less data.
Continuing this trend could lead to AI that can be trained with just basic examples and explanations. At this point, AI will be as capable as a human, marking the achievement of AGI—Artificial General Intelligence.
It's not happening tomorrow, but with continuous architecture improvements and cross-domain connections, it's inevitable.
Highlighting the impact of cybercrime in 2023. The top 3 months recorded significant asset theft with almost a billion in total value:
1. July: ~$263M lost
2. September: $323M stolen
3. November: $344M
The graph reveals the broader landscape. #aml#hack#cryptocurrency
🚨 ledger library confirmed compromised and replaced with a drainer. wait out interacting with any dapps till things become clearer.
https://t.co/xapunW8zC3
A really serious issue is currently unfolding across most hosted crypto frontends.
There is a supply attack on a popular connector, the @Ledger connect-kit.
It has been infected with a drainer, which you can confirm by deobfuscating the code.
Be extra vigilant!
Cream Finance exchanged another 150,000 DAI from Ethereum to Bitcoin.
The hacker sends all funds to bc1q7zufuefkhtp48ef6q5krrj93kmwzh5damhf3e4
This is one of the branches on the Bitcoin side with deposits to Coinspot and Binance.
#CreamFinanceHack#Ethereum#Bitcoin
How about we stop googling at all (not a question). Asking questions to ChatGPT is way better. If you've got a tricky question and Stack Overflow doesn't jump at you from the first link - you're dead. Seriously, digging through tons of SEO garbage is exhausting. Hail AI!