@suchenzang for some reason this makes me think that @thinkymachines will really make it in long term, they are doing( and releasing) good research work despite all these setbacks and that goes long way!
@osanseviero 1. pls obv maitain the token efficiency (unlike qwen) its the best thing
2. a ~100B class MoE model with excellent knowledge work and agentic/coding capabilities (eg: laguna ?)
3. in house trained dspark drafters?
4. long running tasks, eg: 100-200 coherent tool calls?