interpretability work is super interesting. ran an open-model (laguna XS 2.1) and read through the J-lens of the model to try and analyze the failures in coding problems. in a few cases where the model messes up, the word success shows up prematurely, leading the model to think it was correct before it actually was.
to combat, injecting the word wait and/or urging the model to check its work more frequently and in depth leads to better results (on a small sample). early success tokens hurt results, neutral tokens didn't help when injected, and the wait token improved results where early success was common. i think the implications are likely more relevant for non-code problems (as code and math can be verified deterministically). by pushing LLMs to cook for longer at a token-level, seems like it performs better (and the activation space changes to urge the model to do more work). would be curious if anyone's done anymore work looking into the J-space paper from ant (https://t.co/RciwoDXZuq).
initial experiment inspiration via: @venkat9165
Today, HUD is excited to share our Series A funding!
We are the platform for building high quality post training datasets. Over 50 businesses use HUD to build RL environments, sell them to AI labs, or train their own models from them.
Our mission is to enable a generation of data entrepreneurs. The previous generation built apps to impact the world. We believe people will build infrastructure around data, both digital and physical, that align AIs to their specific goals.
Weโve raised $16M in total funding from @Standard_Cap, @ycombinator, @ExceptionalCap, @Liquid2V, @twentytwovc , amazing angels such as @dylan522p from @SemiAnalysis_, @tszzl, @ivanburazin, @theo, and more!