Assistant Prof at SMU (Singapore) working on AI/ML related accounting/finance research with a bent towards NLP. Always searching for good coffee and food.
A strong and secure open ecosystem is important for the world to benefit from AI. We’ve always supported and contributed heavily to open source and science from Jax to Transformers to AlphaFold to Gemma open models which have now been downloaded 300M+ times. And the standards framework we’ve proposed supports responsible deployment of both open and proprietary models.
1/ 🚨 New in Review of Accounting Studies!
We study how misinformation laws affect corporate social media.
Firms tweet 29% less after such laws are passed—especially fewer posts refuting false info.
📄 w/ Yun Lou, Samuel Tan & Liandong Zhang
🔗 https://t.co/kA8i81E22p
2/ Using millions of tweets from 2014–2021, we find the drop is sharpest in countries with:
✅ High social media usage
✅ Strong investor protection
Firms seem to view these laws as credible—reducing the need to defend themselves online.
2/ Using millions of tweets from 2014–2021, we find the drop is sharpest in countries with:
✅ High social media usage
✅ Strong investor protection
Firms seem to view these laws as credible—reducing the need to defend themselves online.
🚀 Just published! Excited to share my newly published paper with Hai Lu and Wenli Huang: "Discretionary Dissemination on Twitter." We explore how companies use Twitter for financial disclosures and its impact on transparency. Check it out in open access! https://t.co/DsqBspVgyC
We've added a new system prompts release notes section to our docs. We're going to log changes we make to the default system prompts on Claude dot ai and our mobile apps. (The system prompt does not affect the API.)
If you are a student or academic researcher and want to make progress towards human-level AI:
>>>DO NOT WORK ON LLMs<<<
LLMs are an off ramp.
Thousands of engineers are working on LLMs with enormous computing resources.
The only way you could possibly contribute is by analyzing existing LLMs and showing their power and limitations.
But it's more fun and impactful to come up with new ideas and new architectures and show that they might work, even on small problems.
This is a nice summary of what to expect GPT3 and LLMs to do, currently. GPT3 is getting pretty oversold currently -- it's an interesting and sometimes useful model, but it is only a *language* model and can't be expected to do everything people are claiming it will do.
My unwavering opinion on current (auto-regressive) LLMs
1. They are useful as writing aids.
2. They are "reactive" & don't plan nor reason.
3. They make stuff up or retrieve stuff approximately.
4. That can be mitigated but not fixed by human feedback.
5. Better systems will come
Before we reach Human-Level AI (HLAI), we will have to reach Cat-Level & Dog-Level AI.
We are nowhere near that.
We are still missing something big.
LLM's linguistic abilities notwithstanding.
A house cat has way more common sense and understanding of the world than any LLM.
Excited by a panel of Gavin Foo (IPOS), professors @prof_rmc@jerroldsoh & moderated by librarian @bellaratmelia discussing implications of Singapore's new Computational Data Analysis provision that some have called a "game changer" for text mining https://t.co/VjJ3Mt3Amv
Being a self taught developer is pretty exciting because you're always being told that the code you wrote is actually using the reverse factory service discovery injection visitor pattern, and you're all like, cool, I didn't know it was called that.
Shoutout to the lawyers that were like, "no, the tweets aren't ALL 280 characters, they're a MAXIMUM of 280 characters, so we need to make that clear."
Wachtell getting thousands an hour for this level of attention to detail while the rest of us chumps see our families and kids.
Another interesting data dump from Pushshift. It could certainly be useful in a lot of research fronts, especially those tracking social media phenomenon.
I gave an talk a couple weeks ago on how various #misreporting detection algorithms performed over the years: a severe drop in performance in more recent years, unless you use a tree-based #MachineLearning approach. Slides and code are on my website at https://t.co/u3Vl34XN20
It took me about 30 years to figure out I should always include the date in the form YYYYMMDD at the end of descriptive file names. How many more years will it take me to convince my students to do the same? 😬
I'm more likely to receive something called draft.pdf. 😏