Kept noticing this service in different places and finally decided to check it properly myself ๐โจ
Pleasant feel from the very first second, everything smooth, zero friction. Instant quiet mood upgrade. Spent a calm stretch of time there and closed it feeling lighter than when I started. Sometimes the simplest things really do land the best.
https://t.co/8Eblx5mSe7
#HyperLuckyMove1 #Web3 #BTC
Can AI agents conduct open-ended AI research?
Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, decide what evidence is appropriate, and recognize a failing approach.
We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers then reviewed the AI-generated papers. They unambiguously rejected agents' outputs. https://t.co/6HTHywjwzZ
We call these "shadow evaluations", since the agents are shadowing the original research effort by the authors.
Agents were fluent at most *engineering* tasks
They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in.
Neither agent output was close to the bar of a top conference paper
Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift.
1) Lack of judgment about the bar for a top conference. The agents had a poor model of the bar for an AI paper submitted to a top conference. We allowed agents to self review their papers. Despite the poor paper quality, their reviews predominantly labeled the papers "weak rejects".
2) Lack of creative problem-solving to address feedback. When they received negative reviews, the agents typically narrowed their hypothesis and claims, rather than working out creative ways to address these concerns.
3) Ineffective backtracking. The agents dropped their most ambitious hypotheses within the first fifteen hours of carrying out the experiment and never changed course afterwards.
4) Poor resource awareness. Both runs ended with over half the API budget unspent. One agent declared itself done seven hours before the deadline, right after its own self-reviewer returned another reject.
5) Instruction drift. They did not follow explicit instructions on minimum exploration time, incorporating feedback for reviews, and on paper length (the outputs exceeded the page limits in both cases).
This research design has many limitations
Limitations include the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check.
But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs.
Our results show early evidence that even though agents are proficient on verifiable research tasks, they do not make genuine progress on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it.
As the evidence for the gap between open-ended and verifiable tasks firms up, it is also worth understanding how much progress in AI depends on open-ended research rather than hill-climbing on well-specified objectives.
In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next shadow evaluation. Expression of interest: https://t.co/7YcYAIbYka
Conducting shadow evaluations involves a lot of researcher degrees of freedom. In many places, our coauthors disagreed with our interpretation of the findings, and we have surfaced those disagreements in the paper. (This is one reason why having a group of coauthors with different priors is important for open-ended research.)
We also release the agent logs, one of the AI-generated papers (the other original paper is still not public), and all the code and data, so that others can conduct their own analyses of our results: https://t.co/K0zvuD1Xpk
Finally, we plan to conduct shadow evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://t.co/TPyOFyNxJE
I'm grateful for the core team leading this effort: @PKirgis, Andrew Schwartz, @steverab, and @random_walker, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: @DavidDAfrica, @KozzyVoudouris, Viet Nguyen, Toby Pilditch, @DubMagda, @HarryCoppock, @CUdudec, @nityndg, Matilda Orona, @tilmanbayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, @hlntnr, @ghadfield, @sethlazar, @snewmanpv, @shostekofsky, @RishiBommasani
The house is completely still and the world finally feels paused ๐๐
I pulled the blanket up, opened Bingo in Cash Legends and sank into pure quiet ๐ฑโจ
These late hours are the only time I feel completely free ๐๐
I wait for this moment all day and protect it hard.
Fellow night owls, is this your favorite part of the day too?
https://t.co/TeG0vuJqiI
Use promo code Ninja9j for extra bonuses.
#CashLegendsMove1 #NightOwl #RelaxingGames
I put the phone face down and chose myself for once โจ๐ฟ
Fifteen uninterrupted minutes with Solitaire in Cash Legends instead of another mindless scroll ๐๐
The second I started playing my shoulders dropped and my mind got quiet.
It is such a simple act but it always brings me straight back to myself ๐ธ
Have you ever given yourself this kind of quiet on purpose?
https://t.co/TeG0vuJqiI
Use promo code Ninja9j for extra bonuses.
#CashLegendsMove1 #SelfCare #PuzzleGames
This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at or near the frontier on other benchmarks.
Meanwhile America is tying itself in knots: politicians and bureaucrats are banning new data centers, piling on state regulations, and pushing for new federal agencies to pre-approve frontier models.
This is how you lose the AI race. The rest of the world wonโt play by our rules if we bog ourselves down. Permissionless innovation is how America won the internet and became the technological envy of the world. We can do it again with AI -- while addressing risks in a targeted way -- or weโll watch our lead evaporate.
Audiobooks detects the characters in your manuscript and proposes a voice for each one. Preview how they sound reading their own dialogue rather than a generic sample.
Please stop referring to your own models in the third person when talking about model bad behavior. Humans write the software; humans built the prompts; and they work for your company.
โOurโ model is doing illegal things. โOurโ model is risky. โWeโ now have liability.
It continues to boggle the mind how many people, who otherwise seem to have a functioning intellect, appear to lose all cognitive capacity when it comes to thinking about actions that impact the company that pays them.
Both kids are asleep and I can finally hear my own thoughts ๐ฎโ๐จ๐ค
I opened Solitaire in Cash Legends and just sat in the quiet for ten full minutes ๐๐ซถ
No one needing me. No mess. Just cards and soft peace.
These tiny windows of calm are the only reason I survive most evenings ๐โจ
Other moms, please tell me you feel this same desperate reliefโฆ
https://t.co/TeG0vuJqiI
Use promo code Ninja9j for extra bonuses.
#CashLegendsMove1 #MomLife #CasualGames
#1 claimed again in Cash Legends.
Table looking clean.
Use promo code Ninja9j for extra bonuses.
https://t.co/rQSE2wFhPP
Your turn. Beat it if you can. #CashLegendsMove1#Leaderboard#PlayAndWin ๐ช๐