I wonder if they know how many works and project are conducted by these international interns.
For me,
GEAR(KV compression, Neurips2024) and TurboAttention(attention acceleration, Mlsys2025) are used in @MSFTResearch coding model service.
ThunderAgent(Agentic Infra ICML2026) is widely used in @togethercompute and @nvidia Dynamo.
Most of them are the works during part time cpt internship.
This is pretty hilarious. Three things make it practically unusable:
1. Cost: $0.7/M output for a “dLLM” is absurdly expensive.
2. Accuracy: scoring 12 on the Artificial Analysis Intelligence Index means it’s basically gibberish.
3. Closed source: no way to finetune or tweak it.
The fastest AI model on the market. Independently verified.
No enterprise contract needed. $1m+ tokens per minute with pay as you go.
$0.2/M input, $0.7/M output.