BREAKING: Missourians voted to end the state's abortion ban, the very first abortion ban enforced after Roe was overturned.
This is a huge win for reproductive freedom in Missouri.
Don’t ask whose fault this was. Plenty of time for second-guessing and recriminations. Ask instead, what can you do? For my part, that means telling the truth as I see it, as long as I can. The media will be under a lot of pressure to toe the line; don’t capitulate in advance
@MarriottBonvoy Are you able to help? We evacuated to inland away from Tampa. Our next evacuation hotel is by TPA. The hotel phone line is busy (not surprising due to what happened in Tampa) Is there a way to find out whether the hotel is operational before we head back west? 🙏
In the latest article about Kamala Harris’ staff turnover in the Washington Post, former staffers complain that she expects her team to be just as prepared as she is, and to be able to explain everything they put in her briefings and schedule.
OK, here is my best guess on the state of LLMs:
- The scale increase between gpt-3 and gpt-4 was 100x
- Doing that for the next model is going to be very hard
- We're nearly out of general language tokens. So let's say we can 2x that. And perhaps get more proprietary tokens and get to 3-4x. And do a lot of data cleaning and get to 6-7x.
- A 100x training run also requires a Gigawatt datacenter which we don't have yet
- Synthetic data is great, but it's not clear how that can be used for general language. I suspect this is why both OAI and Anthropic are focusing on math and code which can be improved via various "synthetic" compute methods (simulated data, or recursive self improvement of some sort)
- In the meantime, there is focus on getting more learnings from the same data. Perhaps there is a breakthrough there but I've not heard of it
- Planning can be pushed to inference in some domains (e.g. coding) which we're starting to hear about. But again, not clear how much this buys.
- Moronic policies like SB 1047 are threatening to slow all this down.
So tl;dr I don't see where the 100x jump will come from for general language reasoning. This is why we're seeing a focus on math and code. I'm glad teams are working hard at new algorithmic unlocks.
(btw, this is pure speculation, would love to know where I'm wrong!)