@thsottiaux UI/UX feature for agent - UI collaboration.
Please add push notifications for when the agent needs context or human input. This could be as simple as codex sending a notification (incl on iOS app) when a run is done OR better as a hook mid-run.
@joseph_h_garvin@METR_Evals I susoect METR didnt tell the model that this behaviour was “cheating”.
e.g. if I tell you to solve a hard math problem, and you find the problem on the internet and share the solution, technically I’ve delivered the result, but METR may call this cheating.
@hwchase17@TheEthanDing@hwchase17 - the models.. They're now more reliable.
1. Connect the model to the code base that created the db
2. Connect the model to the docs for what columns mean
3. Models can list schema and retry.
These all make it now viable (and working) where it wasn't before.
@embirico@OpenAI@embirico a Jupyter notebook extension that autonomously completes analysis tasks based on a prompt input. Around 3 weeks since first commit). Codex for everything (except some UI components).
* multi-provider
* permissions system
* Auth & user security
& more.
@TheMingjie I'd appreciate a key loaded with $500-$1k for a side project I'm working on - it's a notebook assistant for working with Research & Data Scientists in Jupyter notebooks.
@simonw I think it’s more likely that as safety/usability features are built-in, helpfulness/performance falls.
In addition, there’s likely a trade off with turbo between making cheaper compute and faster inference speed vs. Accuracy and performance.
@EMostaque Companies will have a $ budget or compute budget, not a “performance” budget.
The demand for GPUs should in the most part hold if the budget projections are accurate.
In other words, if such a 10% improvement is possible, just train a bigger model instead for longer.
@RadissonBlu Disappointed that we had no clean water access during our stay in Mysore, India. Also, the hotel couldn’t offer a satisfactory fix.
Look at the colour (while running a bath to bathe our 1 year old).
This is a nice Bayesian Optimization blog post by some scientists at Apple (including a discussion of @DrewDim's excellent paper on shrinkage estimators). Notable to see folks at Apple sharing this much information about their work. https://t.co/UJyyC2ouB8