✨ just shipped: Appwright
we're building AI agents for software QA (more on that soon!) – and we saw how painfully e2e testing for mobile apps can be
to solve this, we built a new test framework that combines the best of @AppiumDevs and @playwrightweb — and it's called Appwright
https://t.co/5bb9uyxvp6
the "hype" is coming in the way of a foundational understanding of what LLMs are, and what a developer needs to know about them
to fix this, I've put together a pragmatic intro guide—something I wish I had a year ago
no serious developer would build an app that only runs on AWS, but chances are that your LLM app only runs with OpenAI models
the new smaller Llama 3 model is insanely good: model accuracy is as good as GPT 3.5, while being cheaper and faster. but chances are that your prompts are tightly coupled with your LLM, making it harder to switch
@empiricalrun enables you to track and compare model performance, and iterate towards the best answer for your use-case—so that you can ride the tailwinds of LLM improvements
🧪rapid prompt-engineering + intuitive evals 🩺
[prompt-learner 🤝empirical]
🟢tailor-made prompting, specific to task & model, aids smaller models significantly more than larger models
🟢2x ⏫ on sql execution with haiku
🟢sql syntax correctness 🔼99%🤯from 56% on llama3
every LLM launch uses benchmarks like HumanEval to showcase performance
these benchmarks are everywhere—but they have little utility for app developers. they are mere numbers on a blog post. what's missing is a "vibe check"
what if we could play with these benchmarks, tweak them for our use-cases, and do this quickly with fast iteration cycles?
🦉 new on @empiricalrun
we've built the best playground for developers to try the new Gemma models
compare models and prompts side-by-side, and across multiple scenarios at the same time
link in thread ⬇️
🦉 new on @empiricalrun
prompt engineering can be frustrating, especially when you need to support multiple scenarios
we shipped a better playground that lets you "prompt in bulk". demo ⤵️
introducing @empiricalrun 🦉
new LLMs are launching daily, and showcasing their capabilities with theoretical benchmarks
BUT app devs need more: they want to play with and observe model behavior on real-world scenarios
we're solving that with empirical. launch post ⬇️
India in last 3 decades at Asian Games
From that 1 Gold to 25 Gold,
From those 23 Medals to 100 Medals.
India is becoming that sports powerhouse, we dreamed of back then.
We just launched our Interactive Live Stream SDK on Product Hunt 🐱
Devs can now build Instagram Live-like experiences right inside their apps
100% customisable. Scale to millions. Go live in hours!
Show us some love on PH
https://t.co/2UgE7sNOUi
🚀 Super excited to launch something we've been working on for a bit with @vercel!!!
🗣️ The all-new virtual events starter kit with live video!
We're bringing the power of custom live video to all developers with a self-hosted starter kit:
Demo- https://t.co/GOWbwb8y7d
(1/4)