The new era of LLM assisted coding actually means those who have good communication skills will excel, as working effectively with an LLM at scale for coding means you have to be really good at communicating and giving feedback, just like working with an intern or co-worker.
🦜🛠️ Introducing LangSmith 🦜🔗
A unified platform to help developers debug, test, evaluate, and monitor their LLM applications.
Integrates seamlessly with LangChain, but doesn't require it.
This is a pretty cool paper.
It first introduces a new multimodal biomedical benchmark with 14 different tasks.
It then presents a proof of concept for a generalist biomedical AI system called Med-PaLM Multimodal.
It supports different types of biomedical data like clinical text, imaging, and genomics.
"In a side-by-side ranking on 246 retrospective chest X-rays, clinicians express a pairwise preference for Med-PaLM M reports over those produced by radiologists in up to 40.50% of cases, suggesting potential clinical utility."
Multimodal systems keep making progress. I believe it's possible that we might end up with all these general-purpose AI systems that cater to different domains. It would be cool to incorporate these as components to more general AI systems based on needs and problems. As an example, curious how this could be combined with Galactica for instance for more research-related tasks. There is a ton of research that could be done on the interplay of LLMs - not only for evaluation but also for increasing capability.
There is also a lot of work to be done to better evaluate these systems too. The proposed benchmark is a good start but for a domain like health, we need a pretty robust set of tools and benchmarks for assessing the robustness, safety, and quality of models.
(paper in the replies)
ChatGPT Code interpreter is doing it's own error correction loops, so I guess that's progress? Still tedious as a collaborator, but things are moving forward.
I've been adding in an "about me" section to my training scripts for some time now, good to see this getting baked in. Though I'd really rather see code interpreter actually working well.
Starting today, you can set custom instructions in ChatGPT that will persist from conversation to conversation. 👀 📌
You can enable custom instructions in the beta panel from the settings.
Thanks for taking the time to do this research! The team is aware of the reported regressions and looking into it.
Side note: it would be cool for research like this to have a public OpenAI eval set. That way, as new models come online, we can test against these known regression cases.
Let me know if I can be of any help on that.
Many of us practitioners have felt that GPT-4 degrades over time. It's now corroborated by a recent study. But why does GPT-4 degrade, and what can we learn from it?
Here're my thoughts:
▸ Safety vs helpfulness tradeoff: the paper shows that GPT-4 Jun version is "safer" than Mar version, as it's much more likely to refuse sensitive questions (answer rate drops from 21% -> 5%).
Unfortunately, more safety typically comes at the cost of less usefulness, leading to a possible degrade in cognitive skills. My guess (no evidence, just speculation) is that OpenAI spent the majority of efforts doing lobotomy from Mar to Jun, and didn't have time to fully recover the other capabilities that matter.
▸ Safety alignment makes coding unnecessarily verbose: the paper shows that GPT-4-Jun tends to mix in useless text even though the prompt explicitly says "Generate the code only without any other text". This means practitioners now need to manually post-process the output to be executable - a big annoyance in an LLM software stack.
I believe this is a side effect of safety alignment. We've all seen GPTs add warnings, disclaimers (I'm not a <domain> expert, so please consult ...), and back-pedaling (that being said, it's important to be respectful ...), usually to an otherwise very straightforward answer. If the whole brain is tuned to behave like this, coding would suffer as well.
▸ Cost cutting: no one knows for sure if GPT-4-Jun is the exact same mixture-of-expert configuration as GPT-4-Mar. It's possible that (1) parameter count drops, (2) number of experts is reduced, and/or (3) simpler queries are routed to smaller experts, and only complex ones maintain the original computation cost.
▸ Continuous integration will be a crucial LLM R&D topic: the AI world is barely catching up on things that the general software world takes for granted. Even this study paper doesn't do a comprehensive regression testing on benchmarks like MMLU, Math, and HumanEval. It only studies a particular prime number detection problem.
Does GPT-4 regress on trigonometry? What about other reasoning tasks? What about quality of code in different programming languages, and the ability of self-debugging?
▸ Open-source for the win: it's funny that this paper comes out at the same time as Llama-2. OSS LLMs don't have such mysteries. We can rigorously version and trace regressions, diagnose and fix all of them together as a community.
@GuillemBraso@OrcunCetintas@lealtaixe This is exiting work. I’m curious how it might apply to agricultural scenarios, where it would be valuable to both decompose plants hierarchically, and track these features that change over long time spans.
0/ Descriptions of finding product market fit, and category creation often miss the hard part of getting to $10m in ARR ...
... which is the crazy effort required to tug, pull and hammer the shit out of the product *and* the market to make them hold together 👇
Earthformer - a space-time Transformer for Earth system forecasting. Earthformer is based on a generic, flexible and efficient space-time attention block called Cuboid Attention
GitHub: https://t.co/gtoZ5gykbv
Try it out with Colab: https://t.co/4gLDv3QMhn
#DeepLearning#python
Important validation for almond growers that winter flooding for groundwater recharge doesn't impact the root zone or yield. #groundwater#almonds#ag
https://t.co/o7v60dedgO
The #PathToAutonomy is your journey. How do you want to make your operation more efficient with technology? Visit https://t.co/makf3zwrRE and chart out your next investments in ag tech.