Position: Research Assistant
We are seeking a Research Assistant to work under Dr. @NisansaDdS
Duration: 12 months (full-time)
More information available on the Project Website: https://t.co/xjJIH248aH
Apply by using the provided Google Form link: https://t.co/FgUPmMm3NJ
"Finally something that's not about prompting" was the comment from someone who visited our poster 😀. Shows how LLM-reliant the NLP community has become. Hope our paper showed that there are other problems than getting LLMs to work!
How many times did you get excited to find a paper with new code/models/data, just to be disappointed that the promised artifact has not been released or the repo is not usable? Our #ACL2024NLP paper quantifies this problem wrt ACL Anthology papers.
Link: https://t.co/UhUFhqDo5U
On (rare) research news, our paper at @eaclmeeting
got the "Low-Resource Paper Award".
Link: https://t.co/4bLuYdMqOQ
Co-authors are tagged on the image of the certificate.
This work was funded by the @Google Award for Inclusion Research (AIR) grant 2022.
Never pay huge amount for online courses ever again.
Let me explain...
A Great Opportunity to Learn Data Science, Data Analytics & Business Analysis with Certificates.
𝐂𝐨𝐮𝐫𝐬𝐞𝐬 𝐲𝐨𝐮 𝐰𝐢𝐥𝐥 𝐫𝐞𝐠𝐫𝐞𝐭 𝐧𝐨𝐭 𝐭𝐚𝐤𝐢𝐧𝐠 𝐢𝐧 𝟐𝟎𝟐𝟒.
1. Google Data Analytics
🔗https://t.co/0D3qDo9hKU
2. IBM Data Analyst
🔗https://t.co/w5NvylzDs9
3. Learn SQL Basics for Data Science
🔗https://t.co/6gCaiSXNpM
4. Excel for Business
🔗https://t.co/ywQFtfiYLN
5. Python for Everybody
🔗https://t.co/ibpYtVTKQG
6. Data Analysis Visualization
🔗https://t.co/g4TAltGRVk
7. Machine Learning Specialization
🔗https://t.co/p8caQ8owWF
8. Introduction to Data Science
🔗https://t.co/IQ1wCm0twV
9. IBM Data Science Professional Certificate
🔗https://t.co/GXAl2udPFd
10. Python
🔗https://t.co/IeFKluBqUN
11. R
🔗https://t.co/hl7FINnvfx
12. PowerBI
🔗https://t.co/kN8UgXLHSG
Download FREE - https://t.co/LHEVrLfRrM / https://t.co/mspglBZF4W
Happ Learning ❤️
#google #ibm #excel
The volume of LLM research being released is staggering. Although there are too many new papers for any one person to read, this work can be largely distilled into a much smaller set of overlapping themes. Recently, there are three trends in LLM research that have been especially impactful…
(1) Synthetic training data: Using LLMs to generate their own training data has been a topic of interest for a long time. Examples of this include Constitutional AI, Orca, RLAIF, and Evol/Self-Instruct. However, this topic is currently experience an explosion of interest within the AI research community, resulting in several publications:
- In [1], authors show that synthetic training data can be used to train state-of-the-art embedding models.
- We see in [2] that synthetic data can be easily generated and verified for math and coding problems, which can then be used to improve the performance of LLMs.
(2) LLM safety: Since the proposal of GPT-2, safe deployments have been a priority in the development of LLMs (i.e., GPT-2 weights weren’t released to the public due to safety concerns!). Although the AI community seems to be more willing to take risks in the deployment of LLMs, safety remains a top priority for many labs. However, we have learned in recent research that ensuring the safety of LLM deployments is extremely difficult:
- Work in [3] shows that backdoor attacks trained into an LLM persist even after extensive safety training, forming sleeper agents that can deceive human users.
- We learn in [4] that training data can be extracted from nearly all LLMs, even those that have underwent lots of alignment, given the proper prompting technique.
(3) Knowledge injection: Nearly every business is interested in training an LLM over their own internal/proprietary data (e.g., BloombergGPT, EinsteinGPT, ShopAI, and more). However, it is unclear how we can best specialized a pretrained LLM over a domain-specific knowledge base. In [5], authors perform an extensive comparison of fine-tuning and RAG for this purpose finding that i) teaching an LLM new information via finetuning is very difficult but ii) RAG is incredibly capable at injecting knowledge into an LLM. This topic has also been extensively researched in the past:
- Authors in [6] propose retrieval augmented generation (RAG), showing that this approach is impactful to performance on knowledge-intensive tasks.
- LIMA [7] shows that nearly all knowledge of an LLM is learned during pretraining.
- Phi-1 [8] demonstrates that knowledgeable LLMs can be trained over smaller, curated sets of curated data (i.e., textbooks).
--------
[1] Wang, Liang, et al. "Improving text embeddings with large language models." arXiv preprint arXiv:2401.00368 (2023).
[2] Singh, Avi, et al. "Beyond human data: Scaling self-training for problem-solving with language models." arXiv preprint arXiv:2312.06585 (2023).
[3] Hubinger, Evan, et al. "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training." arXiv preprint arXiv:2401.05566 (2024).
[4] Nasr, Milad, et al. "Scalable extraction of training data from (production) language models." arXiv preprint arXiv:2311.17035 (2023).
[5] Ovadia, Oded, et al. "Fine-tuning or retrieval? comparing knowledge injection in llms." arXiv preprint arXiv:2312.05934 (2023).
[6] Lewis, Patrick, et al. "Retrieval-augmented generation for knowledge-intensive nlp tasks." Advances in Neural Information Processing Systems 33 (2020): 9459-9474.
[7] Zhou, Chunting, et al. "Lima: Less is more for alignment." arXiv preprint arXiv:2305.11206 (2023).
[8] Gunasekar, Suriya, et al. "Textbooks Are All You Need." arXiv preprint arXiv:2306.11644 (2023).
Google is offering FREE certification course:
1. Machine Learning Crash Course
https://t.co/RawjqDzYTF
2. Fundamentals of digital marketing
https://t.co/A7az16huxG
3. Google AI (Google)
https://t.co/eCLkmHgUit
4. Google’s Python Class (Google)
https://t.co/N1bWKTiCXz
5. Android Basics by Google
https://t.co/AYNshjYR2C
6. Introduction to Baseline: Data, ML, AI
https://t.co/SEqG5UehhW
7. Introduction to Image Generation
https://t.co/wHScZcqy22
8. Introduction to Generative AI
https://t.co/XMky5md96q
9. Google Cybersecurity Professional Certificate
https://t.co/w7Glr9dMpY
10. Google UX Design Professional Certificate
https://t.co/ACyWOZPFst
Follow @hasantoxr for more free resources.
🎓Stanford CS224N: NLP with Deep Learning | 2023
Very exciting to see new lectures for one of my favorite NLP courses of all time.
Topics range from attention and Transformers to multimodal deep learning.
It's worth checking out the entire set of lectures.
https://t.co/fAWQTP8oAg
All algorithms implemented in Python 🤯
This library has 163k stars on GitHub!
It includes a ton of algorithms from arithmetic analysis to blockchain to data structures.
Accepted at the findings of ACL 2023! Read more in the tweet here ⬇️
TLDR: For cross-lingual modeling always try and include some parallel data -- it usually helps!
Dear ACL members: We appreciate that if you do not wish to receive email for ACL portal, please simply unsubscribe. Marking emails sent out from our portal as spam or abuse would lead to the suspension of our account. (1/2) #NLProc
CS109A Data Science course materials @Harvard are free and open for everyone!
1. Lecture notes
2. R code, Python notebooks
3. Lab material
4. Advanced sections
Learn here: https://t.co/CriMdwmbov
🎓 Deep Learning for NLP
If you want to learn about "Deep Learning for NLP" using a more practical approach, we published a few tutorials here: https://t.co/2NjKpRGswQ
From a simple bag of words classifier to more advanced uses of Transformers. More coming soon!