@jdchawla29@dylanbowmanSF@BarrAlexandra fundamentally its all about the ability to collect that data. some of this is a privacy challenge (what cant be shared?) some is a collection challenge (how do you capture/unify data across multiple comm streams?) and some of it is a pure tech issue. i'm working on the tech part.
Partnerships is the difference between companies that can scale and companies that collapse trying to do everything themselves. It's so underrated and such a cheat code that it's strange that startups have a weird aversion to it
I used to think partnerships were nonsense in startups. Now I think they're everything.
Before you start a company, your impression of who's winning is shaped by their social media. Partnership announcements that look cool. From the outside, they seem like a big deal.
A year in, you realize those announcements mean nothing. Seeing behind the curtain convinces you partnerships are mostly noise.
Then, years later, you understand they're actually everything.
Just not the kind with the side by side logos. The real partnerships are your long-term customers, your trusted vendors, your brokers, your lenders, the people you call at 11pm. Built over years of relationship building, flying out to meet each other, late calls.
Startups can't do it alone. It takes your whole team, but it also takes a handful of other companies you trust enough to build something bigger than either of you could alone.
Relationships run the world. The trust I have with a few partners means we can schedule an ad-hoc call and ship a product that's never been built before.
When a @Meow customer needs something, we try to show up the same way.
"Partnerships" are nothing in startups. Partnerships are everything in startups.
"Transformers" by Daniel Jurafsky and James H. Martin is one of the clearest and most mathematically grounded introductions to the Transformer architecture I have ever read.
Chapter 8 introduces the Transformer as the standard architecture behind modern large language models. What makes this chapter particularly interesting is its step-by-step presentation of the underlying mechanisms: contextual embeddings, self-attention, query, key and value vectors, scaled dot-product attention, multi-head attention, residual streams, feedforward layers, layer normalization, masking, and the parallel matrix formulation of attention.
In particular, the treatment of attention as a weighted sum of contextual representations is especially valuable. The chapter first develops an intuitive, simplified view of attention and then gradually derives the full formulation using the Q, K, and V matrices. This approach makes it easier to understand what is actually happening inside the architecture from an algebraic and matrix-based perspective, rather than simply viewing the usual block diagrams.
I think it is an excellent resource for anyone interested in understanding how Transformers work from linguistic, mathematical, and computational perspectives.
https://t.co/3fitdPy6Fv
Introducing autoresearch for arXiv papers
Change 'arxiv' to 'autoarxiv' in any paper URL
An agent deploys to resolve setup issues on the codebase, run a minimal reproduction, and estimate full replication cost. Read more below
cheap detection of reward hacking behavior: our humble contribution to the reward hacking space is available on preprint https://t.co/qyCS0cFEZB we recover frontier model performance (LLMAAJ) behavior with small encoders.
special thanks to @neversupervised and team for Terminal Wrench!
> In their blog post, Anthropic defended its decision by saying the jailbreak isn’t serious. That is not what the trusted partner and the USG believe;
I’ve had a number of conversations with folks inside and outside government about the current situation with Anthropic, and here is what I believe to be true:
— As we know, Anthropic publicly released its Mythos class models earlier this week under the commercial name Fable.
— Fable is Mythos with guardrails. But if those guardrails fail, then you’ve exposed Mythos and its advanced cyber capabilities to people who shouldn’t have them. (Keep in mind that Anthropic itself widely promoted the idea that Mythos was a cyberweapon and needed to be regulated as such. They asked for government regulation of Mythos and championed the guardrails on Fable. If there is a vulnerability — big or small — it is Anthropic’s responsibility to patch.)
— A highly credible trusted partner of both Anthropic and the USG who was testing Fable came forward with a jailbreak of those guardrails. The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused.
— In their blog post, Anthropic defended its decision by saying the jailbreak isn’t serious. That is not what the trusted partner and the USG believe; nor is that kind of minimizing language consistent with Anthropic’s brand as the AI safety company. It’s difficult to fathom how they could claim a jailbreak allowing operability of a cyber weapon could be defined as not “serious.”
— In the past, Anthropic has always said that safety must be top priority and taken super seriously. In this case, Anthropic prioritized the continued offering of the consumer model over safety.
— In reaction, the Admin issued the export control. The Admin did this reluctantly. It’s been very surprised that Anthropic hasn’t wanted to cooperate with a reasonable safety request (ie fixing the jailbreak issue). Anthropic’s reaction is very much at odds with their branding and ethos as a safe AI research community.
— The Admin’s hope now is that Anthropic remediates the safety issue, the export control is lifted, and Fable goes back into general release. The Admin wants all of this to happen as soon as possible. It is frankly bewildered that Anthropic hasn’t wanted to comply with safety requests that it previously said were its highest priority.
— Those trying to misdirect and tie this action to the prior DoW/Anthropic issues are wrong. The Admin values Anthropic’s technical capabilities and feels that this issue, while serious, should be easily resolved. The ball is in Anthropic’s court.