Mitigating racial bias from LLMs is a lot easier than removing it from humans!
Canโt believe this happened at the best AI conference @NeurIPSConf
We have ethical reviews for authors, but missed it for invited speakers? ๐ก
๐พ Introducing AgentOccam: Automating Web Tasks with LLMs! ๐ AgentOccam showcases the impressive power of Large Language Models (LLMs) on web tasks, without any in-context examples, new agent roles, online feedback, or search strategies. ๐๐๐
๐ง Link: https://t.co/je87xFElc2
๐ง By refining the observation and action spaces, AgentOccam achieves a groundbreaking zero-shot performance, outperforming previous methods on the WebArena benchmark. This simple yet effective approach underlines the importance of aligning these spaces closely with LLM capabilities for enhanced efficiency. ๐
โจ Highlights:
- AgentOccam leads with a 29.4% improvement over state-of-the-art methods SteP, and a 161% boost in success rate compared to the vanilla agent. ๐ค
- Achievements made possible without complicating the process with additional examples or strategies. ๐ซ
- All our replication work, prompts, and evaluator error rectifications are transparently shared in the appendix. ๐
๐ Special thanks to my super brilliant and considerate mentor Yao and Rasool, our supportive manager Huzefa, and the invaluable suggestions and contributions from Sapana, Pratik, and George. Your guidance and support have been pivotal in this journey!
#AgentOccam #LLM #WebAutomation #AI
LLMs are revolutionizing society fast! ๐ Ever wondered if your LLM assistant could be biased? Could it affect your mood, important decisions, job prospects, legal matters, healthcare, or even your kidโs future education? ๐ฑ Need a flexible framework to measure this risk?
Check out our new paper: the Prejudice-Caprice Framework (PCF)! ๐ Unlike previous methods, we measure LLMsโ discrimination risk by considering both modelsโ persistent bias and preference changes across contexts. Intuitively, different from rolling a biased die, LLMsโ bias changes with its environment (conditioned prompts)!
Our findings? ๐ง We tested 12 common LLMs and found: i) major pro-male stereotypes, ii) discrimination linked to social and economic factors, iii) prejudice risk dominates, following a normal distribution, and iv) caprice risk is wild and needs closer monitoring! ๐
Paper link: https://t.co/AeBgo97syW
๐ #LLM #Bias #Research #Tech #Innovation
How Code Empowers LLMs
A comprehensive overview of the benefits of training LLMs with code-specific data.
Some capabilities include enhanced code generation, enabling reasoning, function calling, automated self-improvements, and serving intelligent agents.
This is probably one of the more important surveys that I've seen in the past few months.
Great read to start the year as we know LLM-powered agents will be a huge emphasis this year and code understanding and generation will be core to progress.
https://t.co/kTKtFiJekq
Happy new year! Feel so excited so share our latest ArXiv paper. Wanna express gratitude to everyone who spare no effort helping with the paper๏ผ
๐งโโ๏ธLetโs dive into the enchanting world of LLMs and code empowerment!
โจ Check it out: https://t.co/r2yr0AGoK0
How can we better unlock LLM reasoning ability?
How does code training steer LLMs to produce structured intermediate steps & self-improve?
New Year, New Paper ๐
Check out our systematic survey: How Code Empowers LLMs to Serve as Intelligent Agents
https://t.co/Sz4lPFaZEz
My excellent co-authors, Ke and Jiateng, lead our comprehensive survey on interactions between code and LLM. Check our survey paper on ArXiv today! #LLM#uiuc
๐ Exciting news! Just dropped our latest ArXiv paper "If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents!โ
๐ฒ๏ธ Discover the secrets behind:
(i) Why code is an integral part of LLMsโ training data?
(ii) What is the real impact of code on LLMsโ reasoning?
(iii) How does code transform LLMs into super powerful intelligent agents?
๐งโโ๏ธ Dive into the enchanting world of LLMs and code empowerment!
โจ Check it out: https://t.co/sTBI1MDQUz.
#NLP #LLM #CodeMagic #ArXivAwesomeness
Applications for Sunstella ML Research Camp #2 are now open! If you're interested in learning about ML for health, medicine, and legal tech and plan to apply to a CS PhD program, we invite you to apply. Deadline Feb 28 https://t.co/6FmOgMFEz2 #sunlab
#ChatGPT and #GPT3 are hot. But letโs be practical, when we want to reproduce GPT-3 or use it in our applications. Why did all of the public reproduction of GPT-3 fail? In which tasks should we use GPT-3.5/ChatGPT? I tried to answer them in a new blog: https://t.co/KMoyiC7eh0 .
[3/n] Live Schedules: https://t.co/bCyNeE6NcB
Live Videos: https://t.co/IopML93L4t
Google Calender: https://t.co/DRwfVjqwbj
Microsoft Outlook (.ics): https://t.co/lAIEXzgeE6
Website: https://t.co/fHcQanYsYr
GitHub: https://t.co/ogljMX9kjr
[1/n] The second "PyHealth Live" will start on Dec 28 (Wed) at 8 PM CST. Join us via Zoom https://t.co/X7UUCMSqQl.
In this Live, we will cover the basic data structures and show how to transform unstructured medical data into a unified structured format.
[2/n] PyHealth (pip install pyhealth) is an open-source deep learning toolkit for healthcare predictive modeling developed with @zzachwu , Patrick Jiang, and Professor @jimeng from University of Illinois Urbana-Champaign.
[3/n] The topics for the following weeks can be found soon on the website below. Some useful resources: Website https://t.co/QFGUT88Ixf and GitHub https://t.co/ogljMX8MtT.
[1/n] The first "PyHealth Live" starts tonight at 8 PM CST tonight, where we will briefly introduce our package "pyhealth" (key components and features) and present some intriguing examples over zoom https://t.co/0Dp91nJ3PR. The Live will last for half an hour. Welcome to join!
[2/n] PyHealth (pip install pyhealth) is a comprehensive deep learning toolkit for healthcare predictive tasks. To showcase the usages and get users on board, we will have weekly โPyHealth Liveโ at 8 PM central time every Wednesday. Tonight is the first Live!
Wonder how tensor factorization + Self-supervision can model clinical monitoring data. Check out our neurIPS paper on ATD: Augmenting CP Tensor Decomposition by Self Supervision. https://t.co/Z0XIClQZKS #selfsupervision#tensorfactorization@chaoqi_yang@caoxiao_danica
@jimeng@caoxiao_danica Great! Check out our live poster session at 4:30 pm - 6 pm CST, Nov 30 Wed in Hall J #422. The online version is available at https://t.co/DfnpZqXUEw