I will introduce theoretical considerations such as how to assess the quality of probabilistic classifiers, and conversely, how the choice of the performance metric impacts the ability of the hyper-parameter tuning procedure to reliably find good models.
@jn2clark How many epochs? Maybe ADOPT works best (as an optimizer) but makes it easier to overfit as a result? Strong regularization might be needed (e.g. dropout or SAM-style updates?)
@nworbmot Indeed, I read some of the details in the links of the thread after asking :) I suppose this kind of energy storage can help justify the increase of demand side flexibility from 5% to 10%.
I am looking forward to giving a keynote presentation at @PyDataParis 2024, September 25-26, Cité des Sciences. This will be an opportunity to reflect on the meaning of "noise" and predictive uncertainty in machine learning applications.
@benjaminwalker@srush_nlp But one could argue this is just a UI problem: the LLM could be prompted to output an inner voice "train of thoughts" first followed by a final answer and the UI hides the transient inner voice output to the user. I am sure that some 'chat' platforms already implement this.
@karpathy > No production-grade *actual* RL on an LLM has so far been convincingly achieved and demonstrated in an open domain, at scale.
The solver network of AlphaProof seems like a good state in that direction:
https://t.co/eiSGRwW49x
although it might be a stretch to call it an LLM.
The schedule of PyData Paris 2024 is LIVE! 🚀
Have you secured your tickets yet? 🎟️
Join us for two inspiring days of open-source data science, insightful talks, sprints, and countless networking opportunities, on September 25-26! 🌐
https://t.co/bOZDXHFmfJ