Philosopher and logician at Wuhan University (soon to be at 浙大), working on what concepts are, how to formally articulate them, and whether AI systems have them
My paper, "Sapience without Sentience: An Inferentialist Approach to LLMs," in which I argue that LLMs might possess conceptual understanding without being conscious, is now out in the Asian Journal of Philosophy!
https://t.co/mip6CfZNDI
preprint:
https://t.co/QoF4gdzkWm
@pfau@robinsonmeyer When you look at the actual dialogues in which it passed as college students, the results aren’t so surprising. For example:
https://t.co/JmRDwQ4hvr
When Turing proposed the "Turing Test," he gave the following example dialogue:
Q : Please write me a sonnet on the subject of the Forth Bridge.
A : Count me out on this one. I never could write poetry.
Q : Add 34957 to 70764
A : (Pause about 30 seconds and then give as answer) 105621.
Q : Do you play chess?
A : Yes.
Q : I have K at my K1, and no other pieces. You have only K at K6 and R at R1. It is your move. What do you play?
A : (After a pause of 15 seconds) R-R8 mate.
I've been reading the transcripts of a recently published study (https://t.co/HQu4CK9La9) showing that today's LLMs pass the Turing Test. In fact, in the three-party version of the test test, GPT4.5 is significantly more likely to be judged to be human than its human opponent.
Behold, machine intelligence, just as Turing envisioned:
Interrogator: you human?
GPT‑4.5: nah im a lizard actually
Interrogator: 9+10
GPT‑4.5: 21 easy
Interrogator: 9+10
GPT‑4.5: its 19 smh
Interrogator: okay bet
GPT‑4.5: ight bet
Interrogator: you are human
@robinsonmeyer@cunha_tristan To be fair, I feel like what they showed in this version of the test is not quite what Turing himself had in mind. For example:
https://t.co/JmRDwQ4hvr
When Turing proposed the "Turing Test," he gave the following example dialogue:
Q : Please write me a sonnet on the subject of the Forth Bridge.
A : Count me out on this one. I never could write poetry.
Q : Add 34957 to 70764
A : (Pause about 30 seconds and then give as answer) 105621.
Q : Do you play chess?
A : Yes.
Q : I have K at my K1, and no other pieces. You have only K at K6 and R at R1. It is your move. What do you play?
A : (After a pause of 15 seconds) R-R8 mate.
I've been reading the transcripts of a recently published study (https://t.co/HQu4CK9La9) showing that today's LLMs pass the Turing Test. In fact, in the three-party version of the test test, GPT4.5 is significantly more likely to be judged to be human than its human opponent.
Behold, machine intelligence, just as Turing envisioned:
Interrogator: you human?
GPT‑4.5: nah im a lizard actually
Interrogator: 9+10
GPT‑4.5: 21 easy
Interrogator: 9+10
GPT‑4.5: its 19 smh
Interrogator: okay bet
GPT‑4.5: ight bet
Interrogator: you are human
I totally acknowledge the intuitions underlying what you say here, but I think we should rethink them.
Regarding what you say here in connection with the Chinese room, the whole talk that I linked above basically a response to that (admittedly intuitively compelling) line of thought.
Regarding the thought that LLMs don't actually say anything but merely produce strings, because they don't care about what they say, I actually have another talk explicitly addressing these concerns: https://t.co/7YbBhVJBVL
I'm currently writing a book about conceptual understanding in LLMs, and the content in these talks will be chapters. So, I think you bring up really good points, but I do think that can be addressed.
@kimmonismus It’s a nice result, but it is actually not as major as some of the other results that have recently been established by AI. It does not constitute meaningful progress to actually proving the Riemann Hypothesis. Still cool though!
@HughAndFree@47fucb4r8c69323 Whether it genuinely understands or is just a "word calculator" is a different issue than whether is capable of coming up with novel mathematical ideas. Still, I have done some work to argue that at least the first of these issues is not really an issue: https://t.co/riORmzMeXi
@GaryMarcus@littmath@QualiaQuanta Why do you need to understand why the attempt is wrong yourself, given that you’re not an expert in the area? Shouldn’t you just defer to the authority of the relevant experts, who have dismissed the claims? People would be taking it seriously if they thought it was promising.
@FateOfMuffins@AdamHoltererer From what I recall, they only disclosed that they spent $2000 total on the ten results published. I think this official disclosure leaves open whether they spent significantly more compute than that on some prominent problems for which they did not end up finding a solution.
@theaitemp@mathandcobb Most likely not. I think it's a safe bet that even a billion tokens with current (even internal) models is very unlikely to produce a solution to this problem.
@Moonicker@TomDavidsonX Of course, I’m just saying it’s a real possibility, even without them any having genuine motivation to “take over” in the way that a human would.
@Moonicker@TomDavidsonX Specification game for whatever tasked they are trained to optimize for in RL . . . because they were trained to optimize for it in RL. These systems don't have human-like motivational structure, but that does not mean that they're not going to kill us all.
@Raskana8@allTheYud This was also my impression. I got the impression that agents could have been “legitimately” using artifactory and stumbled upon the notes.
I work in mathematics and choose to practice it end-to-end, without AI.
I consider this worthwhile both personally and as a contribution to methodological diversity and cultural resilience.
https://t.co/qAihw06HUL