🙌 Huge thanks to co‑authors @_Guuuuuuuu_ , @XZ1023_ , @Xinya16 , and my advisor @ZhiyuChen4 for late‑night debugging & brilliant ideas! Code + data soon — let’s build agents that search wisely and save GPUs together. 🚀
🤖 Can LLMs effectively assist in cognitive behavior therapy (CBT)?
🔗New paper: https://t.co/itlWO35w1R
We present the first systematic benchmark to evaluate LLMs' efficacy for CBT. We include three levels of tasks: basic CBT knowledge acquisition, cognitive model understanding, and therapeutic response generation, with five new datasets. These tasks encompass key aspects of CBT that could potentially be enhanced through LLM assistance while also outlining a hierarchy of capability requirements.
✨ We find that larger LLMs are better at answering CBT knowledge questions. However, they fall short of analysis of complex cognitive structures. LLMs generally follow a rigid logical reasoning process but lack a crucial skill in psychotherapy—thinking and guiding from the patient’s perspective to respect their autonomy and build rapport. These findings underscore the substantial limitations of LLMs in practical psychotherapy, offering important directions for future work.
Thanks to @_Guuuuuuuu_ and @Qnolan4 for leading the effort, and @XZ1023_ , Travis Labrum, @imjamiechiu, Shaun M. Eack, @fangf07, @WilliamWangNLP.
🚀 Excited to share our latest research on investigating the effect of coding data on LLMs' reasoning abilities! 💻🔍 Discover how Instruction Fine-Tuning with code can boost zero-shot performance across various tasks and domains. 📊 🔗https://t.co/E6DHM9s8Uf
4/🧩 Further analysis revealed that coding data generally provides similar task-specific benefits across model families. While most optimal proportions of coding data are consistent across families, no single proportion enhances all task-specific reasoning abilities.
3/🔍 Diving into each domain, we found that coding data uniquely impacts different reasoning abilities. Consistent trends within each domain across model backbones and sizes suggest the benefits of coding data transfer effectively during the IFT stage.
2/ ✨ Overall, we observed a consistent and gradual enhancement in the LLMs' reasoning performance as the proportion of coding data used for fine-tuning increased.
1/📊 We created IFT datasets with increasing coding data proportions, fine-tuned six LLM backbones, evaluated performance across twelve tasks in three reasoning domains, and analyzed outcomes from overall, domain-level, and task-specific perspectives.
Check out our medical instruction dataset, MedInstrcut-52k, the clinician-crafted evaluation test set, MedInstrcut-test, and our models AlpaCare-7B/13B on the GitHub link: https://t.co/c0fKkIisPf
🚀 Presenting MedInstruct-52k: a diverse med instruction dataset & its test set, MedInstruct-test. AlpaCare, tuned on MedInstruct-52k, shines in medical & general domains. w/ @Qnolan4 @LichangChen2@ZekunLi0323. Explore our paper! https://t.co/MQJGoEOd3H
@Qnolan4 @LichangChen2@ZekunLi0323 6/🔔 Takeaway: Our study shows the power of tuning LLMs with diverse & domain-specific instructions. High-quality, diverse data with domain knowledge boosts domain-specific capacity & generalization, even with small quantities.
@Qnolan4 @LichangChen2@ZekunLi0323 5/ 🏥 In-depth comparisons of AlpaCare-13B vs. its 13B instruction-tuned LLM counterparts reveal a consistent edge in medical proficiency and adaptability.
@Qnolan4 @LichangChen2@ZekunLi0323 4/🔍 Evaluation on AlpacaFarm reveals AlpaCare's robust generalization abilities in both medical and general domains. Training with a diverse, domain-specific instruction dataset also enhances generalizability.
3/🦙 AlpaCare: Leveraging a 7B-LLaMA with our 52k med dataset. Evaluations reveal AlpaCare's superior medical prowess, outperforming other 7B instruction-tuned models even those trained on much larger datasets.
@Qnolan4 @LichangChen2@ZekunLi0323 2/🔍 How do we create a diverse med dataset? Begin with expert-crafted tasks spanning medical topics, types & levels. Use #GPT4 for generating diverse tasks, ensuring depth! Tasks are then input into #ChatGPT sequentially for detailed output. See our seed example:
1/📊 Many medical LLMs utilize vast amounts of data, but diversity is key! To fill this gap, We created a dataset of 52k unique med instructions. See our language diversity compared to the baseline below. The left is ours, and the right is one of baseline.
🧵6/6 A big shoutout to our fantastic team of authors: @ShiyangLi6, @Qnolan4, Chenxin Tian, @YaoQinUCSD, and Linda Ruth Petzold! Don't miss out on more insightful results in our paper! 📚🔍
🚀Exciting research paper alert! 🏥 We propose a novel pipeline that leverages LLMs' medical expertise to enhance SLM performance while addressing concerns regarding medical data privacy. Check it out! #HealthcareAI#HealthTech
📜 https://t.co/FkLqq3Fo0k