A mutualistic collection of cells. Also a @HeriotWattUni academic doing research into machine learning, bio-inspired computing and other interesting things.
Great to have been involved in this initiative led by @sayashk and @random_walker to (hopefully!) improve the use of machine learning in science. Further thoughts in my Substack post: https://t.co/pgCDOqWWdC
Excited to share that our paper introducing the REFORMS checklist is now out @ScienceAdvances!
In it, we:
- review common errors in ML for science
- create a checklist of 32 items applicable across disciplines
- provide in-depth guidelines for each item
https://t.co/M21VWGIZhC
I’ve been meaning to write a book about computers for years. However, this requires time that I probably don’t have. Instead, I’ve decided to set up a substack: https://t.co/z0Sdjnnrz9
📄New Paper Alert! 🌐 "IoTGeM: Generalizable Models for Behaviour-Based IoT Attack Detection" with @justmikejust @michael_lones.🚀This paper tackles the challenge of adapting IoT attack detection models to new data with an innovative approach.
https://t.co/CPDT97CFQJ
New year, new pitfalls! I’ve just updated “How to avoid machine learning pitfalls: a guide for academic researchers” to v4, with new material on LLMs, ensembles, checklists, and data usage, and numerous tweaks elsewhere. https://t.co/CKjUgyE9A9
ML-based science is facing a reproducibility crisis. We think clear reporting standards for researchers can help.
Today, we're introducing REFORMS, a consensus-based checklist authored by 19 researchers across many disciplines.
https://t.co/p9ISyyRnV0
Would be nice if Google Scholar could sort by interestingness, perhaps as decided by a LLM. My most interesting papers (as decided by me) are buried deep in my cited-by list.
Interesting to hear political bias in “AI chatbots” discussed on #PoliticsLive today, though the implication this is due to left-wing programmers suggests a lack of understanding in political circles about how these things work.
Women in Data Science Edinburgh will be held at Heriot-Watt University (Hybrid) on 15th May 2023 and invites contributions from women who use data science: https://t.co/ASiPxKyrZv
@RestIsPolitics@LondonPalladium@RoryStewartUK@campbellclaret Media coverage of ChatGPT etc has focused on plagiarism. A greater concern is automation of highly-skilled jobs. This paper, just released, lists many of the professions which are in danger: https://t.co/ZfTV7gnUZp. Any thoughts on the social, economic and political consequences?
Updated “How to avoid machine learning pitfalls”. New sections on handling spurious correlations, temporal dependencies in time series data, and keeping up with deep learning, plus lots more references: https://t.co/CKjUgyE9A9
in what order did you learn your languages?
1. python
2. R
3. haskell
4. C
5. rust
6. scheme
7. {java,type}script
8. clojure
9. scala
10. java
11. idris
12. go
An interesting paper on why tree-based ensemble models, such as random forest, tend to outperform deep learning models on tabular (i.e. not image or text) data. https://t.co/d3oWzpHDV7
Just updated "How to avoid machine learning pitfalls” to version 2. New sections on data augmentation pitfalls and when not to use deep learning, plus more on leakage, stats and transparency, new references and various tweaks. https://t.co/CKjUgyE9A9