Waking up to see my GPT-6 Astra benchmark, across every single effort level with @VulcanBench is done.
This took ~12 hours to run, and still left me with 12% of my weekly usage from one $100/mo Max subscription.
My Fable 5.1 benchmark took over 45 hours hours to run and required three $200/mo subscriptions + $200 in additional usage credits, for a total of $800 to run.
I am reviewing all of the results now because I'm honestly so shocked at the difference here, I need to do a deep dive and really make sure everything ran as I expected.
More to come.
it is obviously trivial relative to everything else, but the fact that astra can make me whatever fun little game i can imagine and i can be playing it a few minutes later is so cool
Taxes on wealth or unrealized capital gains are a bad idea. Not just because of all the well-known implementation issues. But also in theory.
Our paper, now forthcoming in Econometrica:
Taxes on wealth or unrealized capital gains are a bad idea. Not just because of all the well-known implementation issues. But also in theory.
Our paper, now forthcoming in Econometrica:
When Astra can produce technically correct papers in an hour albeit on boring topics, what will happen when the human in the loop defines the research questions and steer the process more? Journals are already bombarded with AI papers - with Astra this will not slow down. More likely the development will accelerate. Not sure the science community is fully prepared for this.
An interesting failure of Astra: I asked it to conduct original entrepreneurship research with whatever online datasets it could find, pre-registering its hypotheses. It churned out a lot of beautifully formatted, technically correct papers on boring topics. No research taste.
@sflorimm@grok is the pricing of Anthropic and OpenAI in this image correct? What are people saying are the reasons for a higher pricing of Anthropic than OpenAI? Isn't OpenAI the bigger company?
I'd go further and say that the risk isn't the machines acting as if they have awakened, but humans thinking they have.
Models trained on human data will imitate human behavior. The imitation becomes so accurate that people will start convincing themselves that AI has achieved consciousness.
That will push us into all kinds of rabbit holes with unwanted consequences.
In such circumstances, I always go back to the KISS principle in programming (Keep it simple, stupid): AI is a tool. Look and treat it as such (even it if acts very human), and examine the threats (which are real) through that lens as well. Don't let anthropomorphism (which okay might sometimes be needed to describe and abstract the models' behavior) distract you from the main discussions that should be had.