@kalomaze What's the intuition behind the binary classification ^4 scaling? If model gets 4/8 correct, it's guessing, don't want to reward this behaviour, if it's 7/8, it's picking up smth?
@kalomaze Which hyperparams? I keep getting junk text after 100 steps (500 samples) and I can't tell if it's a sampling issue. Lr 5e-7. Happens even before the Lr decay. generally happens when the format rewards saturate
@alkimiadev@UnslothAI What kind of rewards are you using for text classification? Straight up 1/0 reward? Any intermediate reward steps?
Thanks for the info, it's very interesting.
@macrocephalopod Curious to hear more on that, what does integrating quant techniques mean there? You plug in your earnings estimates and model spits out expected return?
@OkayEstimator Agree that it should never be in the backtesting engine but you often find in the signal itself, most likely created during feature engineering.