Assistant Prof @ CU Anschutz Med School | TISLab | Biomedical Informatics | Data Engineering | ML | Bioinformatics | Curious Learner & Passionate Educator
We introduce LLM2Vec, a simple approach to transform any decoder-only LLM into a text encoder. We achieve SOTA performance on MTEB in the unsupervised and supervised category (among the models trained only on publicly available data). 🧵1/N
Paper: https://t.co/1ARXK1SWwR
Just got GPT-3.5-turbo-instruct and turbo function calling to create math animations (manim) just from text
Now you can easily ask to graph, answer questions, and you'll get a beautifully rendered animation explaining the concept
It works thanks to few-shot prompting 🎯
This is huge: Llama-v2 is open source, with a license that authorizes commercial use!
This is going to change the landscape of the LLM market.
Llama-v2 is available on Microsoft Azure and will be available on AWS, Hugging Face and other providers
Pretrained and fine-tuned models are available with 7B, 13B and 70B parameters.
Llama-2 website: https://t.co/PKrrXgHdem
Llama-2 paper: https://t.co/aINNrXNhMb
A number of personalities from industry and academia have endorsed our open source approach: https://t.co/N7HwgW9Suh
Church’s lambda calculus and the Turing machine are equally powerful but differ in the fact that Turing machines use mutable state. To this day, there is a rift between functional and imperative programming languages, because of the separation of Church and state.
@cto_junior@AravSrinivas@minimaxir@denisyarats Exactly this. Langchain is sucking the oxygen out of the room. It sounds compelling to higher-ups and gives the impression any complex project is just a few lines of code. I was excited too before I actually tried to use it.
We are releasing a whole-brain connectome of the fruit fly, including ~130k annotated neurons and tens of millions of typed synapses!
Explore the connectome: https://t.co/EWcwRiO0Oz
Reconstruction paper: https://t.co/wCI3hASUfD
Annotation paper: https://t.co/3bPTNK9hRk
1/6
“We are, of course, rapidly integrating chatbots into our email, calendar, and office suite, and fully intend to monetize that data stream. Don’t say we didn’t warn you.”
@nabeelqu@willdepue Yuuuup GPT-4 is fast and dumb for me this morning. I was trying a bit of code debugging and it suggested I add an import that was already there, in a fairly short prompt even. Same issue via the API.
@nabeelqu@willdepue I haven’t noticed a degredation, and it’s still pokey compared to 3.5. A few weeks ago I noticed it going much faster for a while, though I didn’t try to test the quality much at that time. I bet they’re A/B testing.
In this interview with Le Monde, Yoshua Bengio expresses his fears of some catastrophe scenarios that could be enabled by progress in AI.
One such scenario he is worried about is a flood of disinformation and political propaganda on social networks. He says that we have a "moral imperative" to act to prevent the erosion of democracy due to disinformation.
But here is the thing: that scenario has been happening for years and effective countermeasures have been implemented.
Social networks (Facebook and Instagram in particular) have implemented countermeasures against attempts to corrupt the democratic process since at least 2017. countermeasures against bots and fake accounts, misrepresentation of identity, spam, hate speech, bullying, child exploitation, terrorist propaganda, violent content, and misinformation that endangers public safety, have been in place for years.
And the most paradoxical thing: AI is part of the solution here, not part of the problem! It is because of progress in natural language understanding in multiple languages (due to self-supervised transformers) that, for example, hate speech can be detected and taken down in hundreds of languages.
Meta has a large organization called Integrity that is entirely devoted to security and content moderation. They make massive use of AI.
They publish a quarterly report on what they do:
https://t.co/C4Dyxi6SIh
For example, the number of fake accounts taken down every quarter oscillates between 400 million and 1.5 billion.
The prevalence of hate speech is somewhere between 0.01% and 0.02% of all content.
82% of hate speech is taken down automatically by AI before anyone sees it.
Text generation models, such auto-regressive LLMs may help with producing more "fake" content. But the bottleneck of disinformation is *not* in the production (e.g. QAnon is two guys). There are two bottlenecks: passing through the content moderation systems of dissemination media such as social networks and capturing the limited attention the public has.
LLMs do not help with any of that.
If bad actors attempt to use AI to flood Facebook with misinformation, it will be the bad guys' AI against the good guys' AI.
https://t.co/pLgBVlKKQ6