Microduck training facility about to go live. Time to train my duck army.
Whats the most ridiculous policy I can train onto one of these @pollenrobotics ducks?
@TomasMann1878 I don't think anyone has really found the balance between vibe coding and systems engineering yet.
So far my best method has been to restrict Claude to only creating one file / class / abstraction at a time.
Small individual tasks & commits >> Mega spec file slop
I love that you can take an RL concept and almost 1:1 apply it back to humans.
Take this paper "Learning to Reason at the Frontier of Learnability" https://t.co/jYKUn7k4Ks
Want to learn a new skill? Get an LLM to produce tasks, questions and tests that it believes you have a high variability of being successful at. Its that simple - just keep applying this process iteratively
Some thoughts on the game theory behind pacing the frontier
- Guaranteeing other labs & nations don't defect is near impossible.
- The incentive for defecting is extremely high, however if you "win" but the civilisation as we know it changes for the worse, whats the point?
- The only rational move is for more resources, effort and focus to go into alignment and safety research. If cooperation can't be guaranteed, our next best bet is target the most critical problems. Ensure that the cost of defection on humanity is low.
We Must Pace the Frontier: Iโve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. Weโll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess modelsโ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
@mahnerak Are all the AIRAโ agents given the same scope of research, or are they focusing on different subsets of the problem space?
I assume it also uses some of the research preference model work for pruning proposed child solutions?
its also fairly indicative of how you should view most of the things you do in life.
Very few things are part of your core mission - identify them and treat everything else as a side quest.
The thing I love about the term "side questing" is that it immediately takes the pressure off whatever you are doing and allows you to be a beginner again.
Learning a language -> side questing
Playing a new sport -> side questing
Career pivot -> side questing
inspired by @Recursive_SI results, I've been experimenting with a custom harness for exploring @karpathy autoresearch solutions.
From 250 minutes on an initial seed with a H100 the harness scored 0.985924 bpb - which would put this position 29 on the autoresearch-at-home leaderboard in very little run-time.
Have a few more ideas to test & open source soon.
Few ways you could go about it, simplest just blacklisting specific callers / callbacks - which I think is what most currently do.
Obviously it becomes a cat and mouse game very quickly & it's difficult to fully guarantee you filter toxic takers, so you also need to factor that into your book updates.
@umnovd @TSIPS1267 @titanbuilderxyz Many ways to do this - easiest would just be widen exponentially based on time since last update of the propAMM book.
That way you might get updates in non titan blocks (just might not be top of block) and can still get some flow.
One downfall of AI & vibe coding is that now my girlfriend knows I can create a decent looking CRUD app in < 2h, she wants a custom web app for everything.
This must be what tradies feel like when they get tasked with "small weekend projects"
@QuasarBuilder@potuz_eth also for blocks like this with massive value on the line it makes sense to have your own relay to be the last bid in and buy precious milliseconds, which again, is hard if you are a smaller player.
totally, the game theory makes sense to not bid away all profit. Just seems like there would be some upper notional block value threshold where it makes sense to have a more aggressive strategy.
The main risk is if you are aggressive for larger blocks, slowly it bleeds down into the average blocks and then its a race to zero again on profitability for builders, so I get its a risky game to play. Similar to some of the early searcher games / on-chain RFQ games ;)
Are the bidding strategies not super aggressive just as most blocks don't have this much value on the line?
Even with a very small window to bid, I was surprised most builders just appeared to one up / dime each other vs race to bidding 50-70% of block value to the proposer to lock in a couple of mil profit.
I can see how if strategies are configured to bid x% incrementally with some cap on notional this happens.