@yoavartzi@dlwh You can always use things as functions (params, inputs, state) -> (inputs, state) so it's generally easy to interact across framework boundaries.
I used a mix of both haiku and flax models for a couple project and I didn't run into any problems.
@suchenzang I tried asking it to tell me what happened to Trump in July 2024 without searching, and it did list the assassination attempt with really low latency
@pli_cachete@himbodhisattva If you use the filter at inference time, maybe you generate 1000 candidates, and one passes (or none). If you train the model to pass the filter, then you can just run it once at inference time.
@srush_nlp We had an ACL paper arguing that searching for high-likelihood outputs _should_ give you garbage, under mild assumptions: https://t.co/In0GcGWitE
IMO this is a big reason why sampling beat out search over the last 5-6 years.
@srush_nlp 2. The response to "but this task looks hard even for humans" was "okay but people say this is AGI so it should be better than humans".
I didn't see anything which would reject LLMs being (e.g.) 25% as good at reasoning as humans.
@srush_nlp Not to agree too closely with the aggressive question/comment which was made, but I felt like it was arguing against a bit of a strawman. I consider it an open question still to what extent LLMs are reasoning, since:
1. "Cannot do X reasoning task" != "Cannot reason in general"