Feel like more and more people are coming to the conclusion that clarity of communication is incredibly important while orchestrating teams of agents across a large number of projects - I predict that in not long this will become an official benchmark
btw Claude Code lets you configure your own output style!
drop instructions in ~/.claude/output-styles, then run /config → Output style
this is the one I like to use after a long day lol. honestly helps so much
Roon, I expected more of you. Can someone help me understand what I’m getting wrong with my extrapolation? Based on the current game on the field, open source models will be coming soon at fable level and that will likely continue with the gap closing from ~3-0 months soon. Open source models won’t refuse or be constrained, so how is this not counterproductive to the goal of helping us compete with them and ensuring our models are competitively priced? Aside from the fact that models you train to limit their capabilities will in all likelihood not be as good at finding and fixing exploits.
Also, let me use another analogy that’s going to trigger some people: say you could buy a Glock from the USA which determines when you could and could not shoot it using some unknown decision process, and you went up against a fully automatic M16 rifle that always shoots whenever you pull the trigger. Who in their right mind would ever buy the Glock, and why would the Glock’s manufacturer exist in 2 years?
@shadcn I have my own prompt that
improves my first draft prompts, being able to use that would be great, however if there’s a way to construct this to use the existing subscription for codex or ChatGPT itself that would be awesome
@0xSero Almost perfect - except for xai since 4.5 is a good model, wonder if that’s all cursors team and their RL data/training knowledge and experience or if there’s something about the base model but it’s definitely top 4
@signulll Honestly the better business, why open up another legendary restaurant if you can own the building, restaurants are competitive, NYC real estate is scarce
True but once the data center is built and ready for use (lease commencement): The contract activates. Meta recognizes a “Right of Use” asset and Lease liability (present value of future lease payments) under ASC 842.
It’s at this point that this lease liability functions as debt, it’s a legally binding obligation to make future cash payments and most importantly it appears on the balance sheet as a liability.
It’s essentially debt, but structured to be less risky with a delayed financing cost, since their obligations only kick in once the data centers are fully operational and ready for use, so a good deal for meta (and potentially less expensive due to lower interest / delayed start time of their payments, unless this has already been factored into higher lease costs)
Does no one remember Enron?
Off balance sheet financial obligations worked out so well last time
On the other hand it’s a great move for Facebook, after 2008 established the playbook on how force the governments hand on a bail out , mainly by tying the financial health of the core banking institutions to your own, and they’ve done that, so kinda gotta respect the move - which is why Howard Lutnicks winning this round against sacks and everyone else and we might actually see a ban on open source AI
@arvidkahl Are you selling this to hedge funds as a way to track sentiment and what stocks/financial assets are being talked about the most and in what light?
And if not would you?
The best approach is how Codex is rumored to do this: train on the KV cache itself and use a custom model that compacts the input into an output of max size (~25-50k tokens). This is optimized on a synthetically generated dataset composed of tasks utilizing a variety of unique categories/cases, specializing in task completion and coherence by rewarding the model based on its ability to use that compacted KV cache and complete the task fully without losing any specifics or nuances.
I hope your saying this sarcastically
Testing sol in grok build and Claude codes harnesses showed that codex is not great, there’s a surprising amount of quality difference just based on how they define the tools and how they construct the subagent orchestration and communication patterns and tools
- codex is such a mess if you read how it does this, and how it does tool calling, and tool discovery is horrific and just print giant globs of json which eats up usage at a scary clip, for a single tool selection there’s no reason for this
And don’t get me started on the system prompt, people should be fired over that - it’s crazy how much more you can get out of sol in a better harness even though it wasnt designed for it or trained in it, has no one read it for themselves?
Such a perfect depiction of the subtle nuances of how fable just has design taste and sol is not quite there yet - even though it has made massive progress compared with 5.5
These two generations are the best visual depiction of Fable vs 5.6-sol I've seen
It's also the best way I can explain "big model smell". Big models don't vomit all of this extra llm-y garbage into the UI
This same difference happens in the code the models write
These two generations are the best visual depiction of Fable vs 5.6-sol I've seen
It's also the best way I can explain "big model smell". Big models don't vomit all of this extra llm-y garbage into the UI
This same difference happens in the code the models write