Many people think any given ML project is 99% training.
In reality, it’s 50% evaluation, 40% data cleaning, 8% integration, and 2% training.
The first two set the noise floor for learning. No ML magic matters; the model cannot lower the noise floor, as that’s the optimal bound of Shannon encoding of your data.
Thus, not a single day goes by without me thinking about ontology. Even the old labels have to be constantly reviewed.
今还看到了一个非常激进的观点,来自 Yuchen Jin。他说,任何没有尝试过使用人工智能进行编程的科技公司 CEO 都错过了机会。谷歌的谢尔盖在编程,Meta 的扎克伯格也在编程,Shopify 的托比在编程。如果你没有亲身感受到人工智能发展的速度,你就无法预见未来。你很可能会被那些预见到未来的人颠覆。
对于最近沉迷 vibe 无法自拔的我来说,可太喜欢这句话了。AI 的变化现在是指数级的,人的认知大部分时候是线性的。不亲自下场,就会本能低估这个速度。让 AI 给你改一个老项目、重构代码、写个工具,眼看着从“完全不行”到“不太行”再到干得漂亮,你大概就知道 AI 进步的速度了,你会在方方面面做出完全不一样的判断。