Become the best at what you do, not because of the legacy you’ll leave, but because of the lives you'll bless. - B P Hardy;👩🏫💁♀️🇵🇭 NLP| ML| Nature| Cats
Disappointed with your ICLR paper being rejected?
Ten years ago today, Sergey and I finished training some of the first end-to-end neutral nets for robot control 🤖
We submitted the paper to RSS on January 23, 2015.
It was rejected for being "incremental" and "unlikely to have much impact"
Our resubmission to NeurIPS was also rejected
It now has >4,000 citations (and more importantly, end-to-end training is widely accepted!)
It's also cool to think about what's changed and what's the same --
- The network was 92k parameters and trained on ~15 minutes of data
- The code was a combination of matlab, caffe, ROS, a custom CUDA kernel for speed, and a low-level 20 Hz controller in C++, all talking to each other. ROS+matlab was as bad as it sounds.
- We pre-trained the encoder and did inference off-board on a workstation with a larger GPU.
- We were paranoid about varying lighting messing up the network, so we did all the experiments after sunset (so long nights running experiments on the robot past 3 am)
Now, we have manipulation policies that are far more dextrous, far more generalizable, and maybe on the cusp of breaking into the real world. :)
(the paper: https://t.co/qDlGSdDExL)
🤔How do multilingual LLMs encode structural similarities across languages?
🌟We find that LLMs use identical circuits when languages share the same morphosyntactic processes. However, they involve specialized components to handle tasks if contain specific linguistic features⤵️
The Philippines braces for a triple threat as three tropical cyclones are expected to barrel through the Philippine area of responsibility this week, PAGASA said.
#NikaPH already made landfall in Aurora on Monday morning, November 11, forcing thousands to evacuate. #WeatherPatrol
READ: https://t.co/rs5Xsyz8cv
MIT's "Mathematics for Computer Science".
A 1048 page available for Free.
Focuses on explaining the use of mathematical models and methods to analyze problems in computer science.
True Zero-shot MT
Some thoughts on translating to truly unseen languages, Gemini 1.5's results on the MTOB long-context MT dataset, and similarities to L2 language acquisition.
https://t.co/6U0DmmozSq
Happy to announce this will be presented at #EACL2024 .
Special thanks to the anonymous reviewer who suggested measuring oscillatory hallucinations with TNG: our method reduces oscillatory hallucinations by 75-92% on average!
For this week’s NLP Seminar, we are thrilled to host @jacobeisenstein remotely to talk about "A Distributional View of Trustworthy NLP"!
When: 01/25 Thurs 11am PT
Non-Stanford affiliates registration form (closed at 9am PT on the talk day): https://t.co/x84lB5HGG1
You can build a full-stack application using Python alone.
You don't need JavaScript, CSS, or HTML.
If you are a data scientist or someone dealing with data, here is an open-source Python library that will let you build end-to-end production applications without worrying about learning web development:
https://t.co/0NQxKtJC6n
Star the repo!
Taipy works with Python. It has a library of pre-built components to interact with data pipelines, including visualization and management tools. It supports tools for versioning and pipeline orchestration.
It's open-source and comes with a Visual Studio Code extension that will get you started without writing any code.
Thanks to the team behind Taipy for collaborating with me on this post.
Adding this to your tool belt is one of the easiest ways to improve your Data Science career in 2024.
Exploring and distilling broader NLP research narratives was the goal of The Big Picture Workshop at EMNLP 2023.
Here is a look back at the excellent invited talks where authors discussed topics from in-context learning to attention from different perspectives.
Gemini has launched! We are starting a series of notebooks demonstrating Gemini in Colab, starting with using Sheets and Colab to facilitate creation of a personalized holiday note. Once you have a Gemini API key from Google AI Studio, try it out here: https://t.co/PMAriZOj7L
Introducing Gemini 1.0, our most capable and general AI model yet. Built natively to be multimodal, it’s the first step in our Gemini-era of models. Gemini is optimized in three sizes - Ultra, Pro, and Nano
Gemini Ultra’s performance exceeds current state-of-the-art results on 30 of the 32 widely-used academic benchmarks. With a score of 90.0%, Gemini Ultra is the first model to outperform human experts on MMLU.
https://t.co/yCsjfQKO9F
BPE is great for subword tokenization but it's not linguistically relevant, right?
Well... we show that typological knowledge can be induced from simply raw text and a compression algorithm
📃Check out our CL work at #EMNLP2023
https://t.co/PDzYd66Cbu
I’m very excited to share our work on Gemini today! Gemini is a family of multimodal models that demonstrate really strong capabilities across the image, audio, video, and text domains. Our most-capable model, Gemini Ultra, advances the state of the art in 30 of 32 benchmarks, including 10 of 12 popular text and reasoning benchmarks, 9 of 9 image understanding benchmarks, 6 of 6 video understanding benchmarks, and 5 of 5 speech recognition and speech translation benchmarks. Gemini Ultra is the first model to achieve human-expert performance on MMLU across 57 subjects with a score above 90%. It also achieves a new state-of-the-art score of 62.4% on the new MMMU multimodal reasoning benchmark, outperforming the previous best model by more than 5 percentage points.
Gemini was built by an awesome team of people from @GoogleDeepMind, @GoogleResearch, and elsewhere at @Google, and is one of the largest science and engineering efforts we’ve ever undertaken. As one of the two overall technical leads of the Gemini effort, along with my colleague @OriolVinyalsML, I am incredibly proud of the whole team, and we’re so excited to be sharing our work with you today!
There’s quite a lot of different material about Gemini available, starting with:
Main blog post: https://t.co/NzSycJl7aE
60-page technical report authored by th Gemini Team: https://t.co/CEdMRyYSLo
In this thread, I’ll walk you through some of the highlights.
Do multilingual language models (MultiLMs) have what it takes to reason across languages?
Our #EMNLP2023#NLProc paper proposes a new attention mechanism that considerably improves the cross-lingual generalization of MultiLMs!
We still see lots of links to old releases of CS224N. Make sure you're getting the latest goodness (RLHF, prompting, transformers) from the 2023 release!
YouTube: https://t.co/6hb0EIaR5Z
Website: https://t.co/5ReBrK9J4n
For fee cohort-based online class: https://t.co/bDh197g0X6