Just a irrelevant point: LLMs know a LOT, did tons of testing on politics related questions. And the more I know China & politics, the more I realize how good LLMs are. You will also be amazed how much nuanced stuff are in internet corpus, despite heavy censorship.
A Princeton PoliSci China expert publicly admitting that Claude, an LLM that doesn't actually "know" anything, knows more about China than he does is basically telling on the entire discipline of political science.
1/
A top quant lost his job due to an AI reasoning model replacing him at his company. He kept applying to different companies and tried his hand at macro writing but to no avail.
Eventually he swallows his pride and talks to his school friend who is now a plumber.
"I understand your old position was a finance maths guy. Why don't you come to our company and apply for a plumber position? You will earn half your old salary but the overtime and the union benefits make up for the rest. But remember, when you apply, tell them that you completed only seven elementary classes. They don't like educated people."
So it happened. The quant got a job as a plumber and his life significantly improved. He just had to seal a screw or two occasionally, and his salary was good enough and he had zero stress.
One day, the board of the plumbing company decided that every plumber had to go to evening classes to complete "basic financial literacy" certification. So, our quant had to go there too. It just happened that the first class was retirement planning. The evening teacher, to check students' knowledge, asked
โIf you invest half your money in stocks and half in bonds, how do you calculate the portfolio return?โ
The person asked was the quant. He jumped to the board, and then he realized that he had forgotten the formula. He started to reason it, and he filled the white board with He defined a filtered probability space (ฮฉ, โฑ, {โฑ_t}, โ) and posited two correlated geometric Brownian motions for the risky and โrisk-freeโ assets. He invoked the Radon-Nikodym derivative to switch to the risk-neutral measure โ, then switched back because the question was about realized returns, not prices.
He filled the whiteboard with stochastic discount factors, covariance matrices, CRRA utility functions, and pages of Itรด calculus. He derived the wealth process under a self-financing strategy:
dW_t = W_t[(w R_s + (1โw) R_b) dt + w ฯ_s dB_t^s + (1โw) ฯ_b dB_t^b]
Then he solved the HJB equation for the optimal allocation, noted that Mertonโs solution collapses to a constant weight under log utility, and circled w = ยฝ as the given constraint.
He invoked the linearity of expectation. He cited Markowitz (1952). He drew a small efficient frontier in the corner for context.
Finally, exhausted, chalk-dusted, eyes wild, he arrived at:
R_p = ยฝ R_s + ยฝ R_b
Then forty plumbers, in perfect unison, slammed their wrenches on the desks and roared:
โYOU FORGOT THE VOLATILITY DRAG, YOU FUCKING TOURIST!!โ
Based on a survey on coding usage among social scientists fielded in March 2026, only 20% of social scientists regularly use coding agents, with large differences across sub-field. 2/6
I feel there is a trend of "enboringization": we distribute more energy to boring tasks due to failure of proper detection and evaluation. More closed book test and auditing. I also spend much more time on examining referencing and sentence format.
This post is hype and nonsense.
We scientists have not "lost our jobs" to AI.
As someone who has been working on AI for scientific automation, for the past few months, and extensively tracked studies like these - I am of the firm belief that LLM-based multi-agent systems are great for automating tedious or incremental work in:
- Analysing a research landscape.
- Posing new hypotheses.
- Designing and orchestrating experiments, with integrated equipment and routine approaches.
- Analysing results.
- Designing and carrying out follow-up studies etc.
But what they are NOT good for is thinking too far outside of the box of their training data, RAG-based vector databases etc.
i.e. You can make an AI system emulate a boring, routine, and incremental scientist - which is extremely useful - but you cannot yet make AI automate Nobel Prize winning tail-end creatively ingenious ideas.
We are nowhere close to an AI Einstein, Newton, Darwin, Turing, Gรถdel, or Conway.
Anyone who claims otherwise is pumping hype rather than reality.
It reminds me of two types of evaluations for LLMs:
Static benchmarks for reasoning, like QA tasks.
Dynamic benchmarks that are constantly refreshed, like predicting market.
It would be interesting to design a mechanism for students that leverages type 2 evaluations.