@NikHarmon@aaroe_lene 3) We reran a number of experiments to double-check some results for the sake of reproducibility - scientific rigour adds to project costs!; 4) Some quick back-of-a-napkin cost-benefit calculations suggests that human annotator hours are still the best use of resources
@NikHarmon@aaroe_lene Hi Nikolaj! We didn't expressly measure costs but a couple of anecdotal observations. 1) Setting up robust LLM pipelines is non-trivial in terms of labour time; 2) Running LLMs at scale is compute intensive; [cont]
@ML_Burn@aaroe_lene Hi Mike! The supervised classification was done using 5-fold cross-validation (i.e. 80/20 splits) on 2,900 human annotates tweets. You can check out the code on Github (under supervised_classification.py): https://t.co/NOV3LZOaYz
To get a feel for how we ran the LLM classification experiments, check out this blog post by my junior colleague and co-author, Márton Kardos. He developed a Python package for doing exactly this:
https://t.co/MlfVEmO17k
Some interesting results here, I think! We originally set out on with this report to evaluate the performance of open-source LLMs relative to the major closed-source alternative.
🚨Preprint🚨 We provide a systematic evaluation of the performance of ChatGPT, Open Source LLMs & supervised machine learning classifiers for text annotation. We find that "Chatbots Are Not Reliable Text Annotators" https://t.co/xv0Z3C5K9y. We discuss results given #openscience
Applying LLMs in humanities and social science contexts is fashionable right now - and for good reason. But with this report we try to urge some caution and reflection on what this means for open science principles.
🚨Preprint🚨 We provide a systematic evaluation of the performance of ChatGPT, Open Source LLMs & supervised machine learning classifiers for text annotation. We find that "Chatbots Are Not Reliable Text Annotators" https://t.co/xv0Z3C5K9y. We discuss results given #openscience
For the Friday evening crowd still mulling over Chomsky’s op-ed, here’s a key passage from our recent letter in @cogsci_soc. There are absolutely some things LLMs cannot do. But what they can do, they do very well - and that demands attention. https://t.co/TFWExGgHVj
If you are unsatisfied by Chomsky et al.'s painfully petrified op-ed, boy do I have a read for you! It wasn't planned, but CogSci just published a letter by @MH_Christiansen , @ross_dkm and I arguing quite literally the opposite, at least re:language. https://t.co/fOW788uCyL 🧵
If you are unsatisfied by Chomsky et al.'s painfully petrified op-ed, boy do I have a read for you! It wasn't planned, but CogSci just published a letter by @MH_Christiansen , @ross_dkm and I arguing quite literally the opposite, at least re:language. https://t.co/fOW788uCyL 🧵
In which we argue that cognitive science should take seriously LLMs as working models of how human-like language can be acquired without the need for a built-in grammar. Really excited to see this one published. It’s short and it’s free - give it a read! https://t.co/ySdoCQajh9
In a letter in Cognitive Science, @pcontrerask,
@ross_dkm and I argue that Large Language Models - despite their many shortcomings - demonstrate that human-like grammatical language production can be learned from experience alone.
Read it for free here: https://t.co/t7Fd4FJn87
Interested in joining @uantwerp as a postdoc in computational humanities or digital text analysis? There is a nice call from the YUFE network with as focus areas: Sustainability; Digital Society (!), Citizens Wellbeing; and European Identity. PMs welcome.
https://t.co/eWAKA8V3oL
@KCEnevoldsen I haven't done it since I began to consistently use type hints, seems pointless. Google's Python Style Guide advises against doubling up, too https://t.co/NXuQsBpBES