๐ HPC-AI is heading to #CVPR2026! Find us at Booth 237 from June 3โ7! Whether you're building AI Agents or hunting for top-tier compute, weโve got you covered.
๐ The Ultimate CVPR Swag & Perks:
1๏ธโฃ Limited Swag: Grab custom caps, canvas bags ๐งข ๐๏ธ
2๏ธโฃ Top-up Reward: ๐ฐ$15 = GPU Hour + Model APIs + Summer mini-fans / Bluetooth speaker (Only 20 set)๐ฅ
3๏ธโฃ Accepted Author Special: Top up $40 Get $20 voucher ๐ฐ
4๏ธโฃ Survey Gift: Just do a quick survey, choose your refrigerator magnets ๐งฒ
Stop by, talk tech, and claim your rewards! ๐ค
๐ Booth 237 | June 3โ7 | Colorado Convention Center
#CVPR2026 #HPCAI #AI #MachineLearning #GPU #ModelAPIs
Pressure Testing GPT-4-128K With Long Context Recall
128K tokens of context is awesome - but what's performance like?
I wanted to find out so I did a โneedle in a haystackโ analysis
Some expected (and unexpected) results
Here's what I found:
Findings:
* GPT-4โs recall performance started to degrade above 73K tokens
* Low recall performance was correlated when the fact to be recalled was placed between at 7%-50% document depth
* If the fact was at the beginning of the document, it was recalled regardless of context length
So what:
* No Guarantees - Your facts are not guaranteed to be retrieved. Donโt bake the assumption they will into your applications
* Less context = more accuracy - This is well know, but when possible reduce the amount of context you send to GPT-4 to increase its ability to recall
* Position matters - Also well know, but facts placed at the very beginning and 2nd half of the document seem to be recalled better
Overview of the process:
* Use Paul Graham essays as โbackgroundโ tokens. With 218 essays itโs easy to get up to 128K tokens
* Place a random statement within the document at various depths. Fact used: โThe best thing to do in San Francisco is eat a sandwich and sit in Dolores Park on a sunny day.โ
* Ask GPT-4 to answer this question only using the context provided
* Evaluate GPT-4s answer with another model (gpt-4 again) using @langchain evals
* Rinse and repeat for 15x document depths between 0% (top of document) and 100% (bottom of document) and 15x context lengths (1K Tokens > 128K Tokens)
Next Steps To Take This Further:
* Iterations of this analysis were evenly distributed, itโs been suggested that doing a sigmoid distribution would be better (it would tease out more nuanced at the start and end of the document)
* For rigor, one should do a key:value retrieval step. However for relatability I did a San Francisco line within PGs essays.
Notes:
* While I think this will be directionally correct, more testing is needed to get a firmer grip on GPT4s abilities
* Switching up prompt with vary results
* 2x tests were run at large context lengths to tease out more performance
* This test cost ~$200 for API calls (a single call at 128K input tokens costs $1.28)
* Thank you to @charles_irl for being a sounding board and providing great next steps
Dishonest CBC headline:
"Canada's AI pioneer Geoffrey Hinton says AI could wipe out humans. In the meantime, there's money to be made".
The second sentence was said by a journalist, not me, but you wouldn't know that.
Even if you're in last place. ๐
Even if the weather is terrible. ๐ง๏ธ
Even if it feels like you can't do it. ๐ซ
๐๐๐ซ๐๐ง ๐๐๐ซ๐ ๐ช๐ฅ ๐ช
Nothing was going to stop Cambodia's Bou Samnang ๐ฐ๐ญ from finishing the women's 5,000 metre race at the #SEAGames.
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
Turns LDM Stable Diffusion into an efficient and expressive text-to-video model with resolution up to 1280 x 2048.
proj: https://t.co/BS3fFbLCSs
abs: https://t.co/oICRrLdSjE
JUST IN: NVIDIA dropped new text-to-video research.
While still far from Hollywood quality, it's pretty damn impressive how fast this is moving.
"a storm trooper vacuuming a beach"