This is the most interesting recent benchmark result that I've seen:
The 100-line mini-swe-agent harness gets better performance out of Opus, GPT, and Gemini than their respective bespoke harnesses.
(As measured on the excellent DeepSWE bench).
Why would that be true?
Has anyone done comprehensive testing of gpt-4-vision-preview?
I want to know stuff like the minimum text size it can read, the radius of the smallest circle it can locate in an image, the number of circles it can count, etc.
Could be an automated benchmark for other models too
Which set of statements do you agree with?
1. AGI is as much or more of a risk to human flourishing as nuclear weapons
2. I have a good idea for what should be done about that
Has anyone had good experiences with GPT-powered code generation for complete web app features?
As in, you describe what should exist, and GPT actually provides the source of all the necessary files and where they should go.
Ideally in the context of Ruby on Rails.
Let's say that a US-based research company has developed an AGI model that was able to use the browser, pass captchas, hire people on Upwork, and lie about its intentions.
What should they do after observing this?
We bring in @full_stack_dl, a venerable boot camp crew that pioneered technical deep dives into deep learning where people fly in from around the world. 🥞 Their #LLM Bootcamp in the spring was sold out and this is your chance to attend the ➡️ version.
👉 https://t.co/9vW3YtaoKR
We're also about 3 weeks away from our latest LLM bootcamp.
@karpathy called the last version "high-quality tokens". Register soon if you want to make sure you get a spot!
The bootcamp is in Oakland on November 13. You can register here: https://t.co/nCNgdt63eL.
We're hosting a livestream with @ScaleByTheBay, this coming Monday at 1:30 pm PST.
Come join us on your YouTube channel to talk about LLMs in production and more.
https://t.co/FOiTDZU4AP)
Solutions from replies:
- @OpenPipeAI looks exactly right https://t.co/tEZqua1nJR
- @PortkeyAI launching feature soon
- @analyticsaurabh building his own
I currently use @helicone_ai, any plans from them?
Is there a service I can use to pipe my GPT-4 calls through, and it automatically finetunes GPT-3.5 (or whatever) on all of them, and lets me know when it's up to par?
By what year will there be an AI that is more capable than most humans in most domains of digital work (e.g. you can tell it to do anything you currently hire a white collar professional to do, and it does the job better than the median human)?
If you can't make it in-person, check out our materials online instead!
Our last LLM bootcamp (as well as our deep learning course) are available for free on our website.
https://t.co/B9mY4kjRc5
🥞🦜 New LLM Bootcamp Announcement 🦜🥞
In 2023, the AI world speedran through models, architectures (e.g., RAG), and frameworks (e.g., @langchain).
After a year of hype, what's *actually* working?
This November, we'll show you, in our latest class on building prod LLM apps