Everyone (myself included) assumes AI is swallowing Search and killing organic discovery. Surprisingly, Shopify's numbers say otherwise: AI-referred sessions are up 3x YoY, but organic still grew 12% on a much bigger base. The pie itself got bigger.
https://t.co/XAmLIFyn5c
@gavinmckew Great question! You actually don't need to - gisting takes care of it automatically. The blog post on gisting is coming soon, you'll see what I mean.
Asked the team to do a post on how we run "The Flywheel": compressing production failures into model weights every day, beating frontier-model quality and cutting serving costs 96%. Shopify FTW! And yes, we did have a banger earnings release today :-)
https://t.co/QoaH5JlFF5
@gavinmckew Simple answer: always both! You want to run ACE constantly https://t.co/fcEFwa7sxD (compressing the prompt, of course - gisting blog is coming up!) and have the Flywheel add it to the training data on the next iteration.
And the next iteration of Gaussianization regularizer can be grabbed here: https://t.co/sJcnKSDZxw
Faster, high-dimensionality robust, necessary and sufficient, with linear-in-batch-size approximation :-)
My first NeurIPS was 2001, first ICML 2005. This is the first year I feel something has changed and we are facing a real crisis. ICML paper quality was clearly lower than before, this time I saw NeurIPS reviews stating “this paper is not for the general public” - ?!!
NeurIPS should not let authors reply to reviewers’ responses to their own rebuttal if they haven’t responded to the rebuttals of papers they reviewed. The authors of every paper I reviewed replied within an hour after I responded to their rebuttal, yet not a single reviewer has responded to mine.
@marcel_hussing Why don't you just message the authors? Not using LLMs for rebuttal is unrealistic nowadays, but using it to generate AI slop is not OK. Messaging the author, saying: "Hey, this is LLM-generated, I need you folks to actually answer the following:..." should do the trick.
OpenAI API has much stricter (read: slopier, lazier) content controls vs ChatGPT/Codex itself. Helped my hedge funds buddies to debug a weird intermittent bug in their API usage - turns out asking about corporate actions of biology-related companies makes it barf randomly...
I finally found one thing that's better on Mac than Windows: when I RDP into a Windows machine, I have an option of using the full alternative desktop for it and an easy swipe gesture to switch between them. @pavandavuluri, you should consider adding it into Windows itself.
@Yingzhe0301 This is true. This is the source of all these problems: reviewing is a thankless chore. We should include quality reviews into H-index, add reviewers as co-authors.
A very standard experience. We have LLMs, why can't we look at reviewers' publishing history and rout to the those who actually understand the topic? 50% of the last ICML was "engineering slop", because reviewers can understand it.
Got NeurIPS reviews. Every concern raised by the reviewers is either already addressed in the appendix or stems from unfamiliarity with the relevant literature. Of course, we’ll cover everything in the rebuttal, but I don’t think it’s realistic to turn a 3/3/2 into positive scores
@npew I'm constantly hearing these discussions: "Oh, it's not really better at coding", "it thinks for 30 min", the future of the API is uncertain, there is no ability to use it as an advisor model in Codex... But for the class of actually hard problems, it is an unbelievable unlock.
It's that time of year: iterating on a rebuttal for NeurIPS review, have 5.6 Pro collaborating with Fable 5 - Fable is not even close to Pro in these math-heavy applications. I am very worried OpenAI is underestimating the power of Pro...
Best-of-n rules. Best for math and everything, really. Just today pitted it head-to-head with Fable - a very clear winner. If only it were available in Codex…