100% agree with Andrew on the unreasonable effectiveness of error analysis, and that it’s a bit different for genAI and agents. Turns out there is a structured & well established framework to help with error analysis in gen AI—grounded theory! Hamel and I go into detail and do live error analysis with a real agent in an episode on Lenny’s podcast that we recorded over the summer (link below)
$14.5M Series A raised to scale AI evaluation & governance! Led by Race Capital & joined by NXTP, Mindset, YC, Quiet Capital, Telefonica, KPN Ventures, and Mento VC.
Full details here:
https://t.co/9wgApVGOmi
✨ Our tests page got a refresh!
- better categorization
- easier to scan
- your custom tests are available to create from the UI
Try it out and let us know your thoughts 💭
just one of the many AI bugs I see every day... this one's in @googledocs' autocomplete. exactly the kind of thing @openlayerco can easily help catch and fix.
⚙️👷⚒️✨ We’ve been building something HUGE over the past couple weeks that we’re itching to share with you soon. Stay tuned next Tuesday.
Subscribe to our Changelog to be the first to hear: https://t.co/C2lCAguzau
🎉 Openlayer now supports Claude 3!
This is Anthropic’s most powerful model yet with significant improvements in intelligence, latency, accuracy. It features stronger vision capabilities and longer context windows.
Create prompts on Openlayer that directly call the Claude 3 API and see how it stacks up.
Using LLMs to generate code? With just a few clicks, set up an automated test to ensure all your outputs are in Python. Get Slack or email notifications whenever they’re not. Try it out on Openlayer.
Thrilled to share I've joined @Felicis as a GP ✨ I've long admired their commitment to founders, new categories & high-conviction bets.
I'll continue to invest in my favorite categories of databases, dev tools, infra & AI from inception to Series A.
https://t.co/pjxwFMFl62
AI model performance and reliabliity
There has been an explosion in the number of LLM models and apps over the past year. While it's important to move fast, it's equally important to think about the performance of these models and apps in the following ways.
First, it's critical to know how your users interface with your model and how your model responds. Logging your production requests is the cornerstone of any LLMOps setup.
Second, beyond measuring your usage, it's important to run stress tests to understand your model's weaknesses and get alerts if those failures come up in production. For example: if PII is leaked, or if the model starts giving irrelevant responses or hallucinating, you need to know ASAP.
Finally, it's important to track your experiments and version your prompts, models, and datasets. At each iteration, you must run tests to guarantee you're always moving the needle forward, versus falling back.
@openlayerco is a next-gen monitoring and eval platform that allows developers to keep track of how their AI is performing, beyond token cost and latency. Their comprehensive toolkit for AI reliability gives developers unique insights into the performance and reliability of their apps and models.
We're live on Product Hunt! 🤩
We'd really appreciate your support as early followers to help us get to the top.
Our launch shows off the newest features we've shipped over the past few months, from LLM evals to monitoring mode and so much more.
https://t.co/abROvNZDLz
Probably the most underrated tool for debugging and improving machine learning models:
Error analysis
Instead of looking at the evaluation metric for all data points, compute it for different groups. Identification of these groups can even be automated.
Our new product page at https://t.co/y2AHQXyP3f is now live 🚀
It shows off some of the new features we’ve built over the past few months, and includes resources on how to get started.
Kudos to @ericaoutput for the designs and animations. If you like what you see, follow her for more design content!