@jackclarkSF Approximately what %age salt are you adding? Most sausage recipes use somewhere around 1.7-2%.
So for 100g pork/main product, 1.7-2g salt (if you’re sensitive to salt, maybe try 1.5?). Might help with the texture too if that’s been an issue.
Anything special for the pastry?
@MichaelTontchev@TheZvi I also wonder if it might just be a property of larger models, reasoning not necessary? Like maybe we could just ask opus 3 something like “{prompt} + roughly how many tokens do you think you’ve used in this answer”? Idk, just spitballing
Really cool work as always!
Would be curious to see some latency matched figures. To a certain point, adding more width is ~free wall clock time wise, and adding depth is serial so much slower.
This also reminds me of something I was curious about with Gemma 4: the 26b MoE (30 layers) uses ~2x as many tokens on the AA index as the 31b dense model (60 layers), so the actual number of layers that were used is ~the same. The 31b scores higher (though it’s confounded with 31b active vs 4b), but is consistent with your point that CoT is basically using more layers but lossy.
@labenz Still have a hard time with generated voices so haven’t listened a ton honestly, but the segments have been good. Would rather have more regular full cogrev eps but understand it’s not really a 1:1 trade :)
over the course of this 4 hour flight the rather trim person next to me has consumed:
- a cinnamon bun
- two cans of ginger ale
- a bag of doritos
- a bag of peanut M&Ms
- two entire kiwis
- a bag of gardettos
- a protein bar