@nileshgiri34 When it reaches ASI ,it doesn't only manifest in llm. It exceeds all human creativity and imagination on how intelligent it could be. If it ASI - it will be in hardware and lot of other integration. It will be exponential.
Agreed. We need to ask “Did the agent behave intelligently and safely while doing it?”.
Instead of “Did the agent finish the task successfully?”.
Now is the time i'm almost tired of some of the benchmarks - because they don't make sense.
So it should include Trajectory of actions AI models took before succeeding or failing. So that we can differentiate whether it got lucky or it is actually smart.
I don't like essays because my handwriting is bad and i'm bad at expressing sometimes but AI shouldn't have MCQ test but an essay test - because they don't care about hand writing -so judge them on best Evals. This was just and example.
Tomorrow we will bring back the 5h limit for Plus accounts across ChatGPT Work and Codex. I had mentioned this a while ago, but then postponed it.
This is necessary as (a) the 5h limit allows us to smoothen the load on our compute, allowing to keep the plan generous in terms of weekly usage and (b) users on the Plus plan are relatively casual and new users, but then also just accidentally eat through their whole weeks usage and then are confused, making it not a great experience.
We are for the upcoming months keeping the 5h limit not enabled for Pro $100 and Pro $200 subscriptions.
Agreed. We need to ask “Did the agent behave intelligently and safely while doing it?”.
Instead of “Did the agent finish the task successfully?”.
Now is the time i'm almost tired of some of the benchmarks - because they don't make sense.
So it should include Trajectory of actions AI models took before succeeding or failing. So that we can differentiate whether it got lucky or it is actually smart.
I don't like essays because my handwriting is bad and i'm bad at expressing sometimes but AI shouldn't have MCQ test but an essay test - because they don't care about hand writing -so judge them on best Evals. This was just and example.
I have 1 year of gemini ai pro sub, and i don't even use it - i hate google for that.
Still hoping for their comeback. Love-hate kinda situation w google.