Forecasting platform @INFERpub is now being run by @RANDCorporation, taking over from @umdARLIS
@INFERpub is an excellent source of predictions on AI, Iran & Taiwan, among others
(It will continue to use @cultivatelabs tech)
We are pleased to share that #INFER is now led by @RANDCorporation. We look forward to this new chapter that will enable INFER to leverage #RAND expertise in our mission to advance the #forecasting capability of the US govt.
See our announcement>>https://t.co/jV5BcxT9EB
🇫🇮 Finland's presidential election is tomorrow
Opinion polls suggest a close race, with center-right Stubb only 2 points ahead of greens Haavisto
However, prediction markets show Stubb is 88% likely to win
Unlike most European countries, the Finnish president holds executive power in foreign policy.
In practice, the PM focuses on EU issues 🇪🇺, while the president deals with other countries 🇷🇺🇨🇳.
Opinion polls:
◦ Stubb: 24% (center-right)
◦ Haavisto: 22% (greens)
◦ Halla-aho: 16% (populist)
◦ Rehn: 12% (agrarian)
Prediction markets:
Stubb:
◦ 88% on @Polymarket ($268k at stake)
◦ 83% on @ManifoldMarkets
Haavisto:
◦ 11% on @Polymarket
◦ 12% on @ManifoldMarkets
⚠️Important context missing from my post: https://t.co/IlpfKhlKcZ
Prediction markets refer to overall winner, which will likely be determined after a 2nd round.
Polling numbers will change once there are only 2 candidates, therefore my comparison is misleading.
My apologies!
(Presumably the populist & agrarian voters backing the 3rd & 4th most popular candidates are unlikely to vote greens; more likely to consolidate behind the center-right candidate?)
🇫🇮 Finland's presidential election is tomorrow
Opinion polls suggest a close race, with center-right Stubb only 2 points ahead of greens Haavisto
However, prediction markets show Stubb is 88% likely to win
Unlike most European countries, the Finnish president holds executive power in foreign policy.
In practice, the PM focuses on EU issues 🇪🇺, while the president deals with other countries 🇷🇺🇨🇳.
Opinion polls:
◦ Stubb: 24% (center-right)
◦ Haavisto: 22% (greens)
◦ Halla-aho: 16% (populist)
◦ Rehn: 12% (agrarian)
Prediction markets:
Stubb:
◦ 88% on @Polymarket ($268k at stake)
◦ 83% on @ManifoldMarkets
Haavisto:
◦ 11% on @Polymarket
◦ 12% on @ManifoldMarkets
ElectionBettingOdds now has Trump at >50% chance of becoming president in 2024 for the first time...
However, I think it's worth disaggregating those odds:
The biggest market is at 54%, driving up the aggregated number (@Polymarket)
But all the other markets are all still at 47% (@PredictIt, @Betfair & @smarkets)
Also, points-based forecasters - not included in ElectionBettingOdd's aggregate - still favour Biden (@metaculus, @ManifoldMarkets & @GJ_Open)
This week's Trump dashboard...
Presidency:
◦ Real money odds diverge on Trump: now 54% on @Polymarket, still 46% on @PredictIt
Primaries:
◦ Trump ~96% likely to win New Hampshire, up from ~70% last week (after @VivekGRamaswamy & @GovRonDeSantis drop out)
Other:
◦ Superforecasters @swift_centre say felony conviction drops Trump's reelection chances to 37%
(if Trump is GOP nominee; see full analysis: https://t.co/KnjQ6EH400)
This week's Trump dashboard...
Presidency:
◦ Real money odds diverge on Trump: now 54% on @Polymarket, still 46% on @PredictIt
Primaries:
◦ Trump ~96% likely to win New Hampshire, up from ~70% last week (after @VivekGRamaswamy & @GovRonDeSantis drop out)
Other:
◦ Superforecasters @swift_centre say felony conviction drops Trump's reelection chances to 37%
(if Trump is GOP nominee; see full analysis: https://t.co/KnjQ6EH400)
Expert forecasters say there is over a 90% chance Biden will be the Democrat nominee, while betting markets imply there is just a 75% chance
https://t.co/Sj7dBoEwZy
What if instead of infotainment, news coverage was like the weather forecast?
A dashboard of probabilities...
Introducing v1 of the Trump dashboard, just in time for the Iowa caucus
(v2 will have pretty charts)
even if i expected the technology itself to continuously progress in a way not fully captured by the prediction (ie each new product announcement shortens the timeline), i would still consider politics etc
eg what if pause movement succeeds?
if some other group arrived at the same timeline totally independent of Metaculus, then i would extremise
just my view - i'm not a good forecaster
DeepMind releases AlphaGeometry, explicitly stating it can solve Olympiad-level geometry problems
What do prediction markets think?
◦ Timeline to Gold at International Math Olympiad accelerates by 13 months, now Oct 2027 (@metaculus)
◦ Timeline to AGI accelerates by 2 months, now Sept 2031 (@metaculus)
◦ AI now 58% likely to win at least Bronze before 2026, up from 40% (@ManifoldMarkets)
◦ AI still 45% to solve next Millennium Prize Problem (@metaculus) & 67% for an unsolved problem in maths before 2030 (@ManifoldMarkets), no change
sorry for the confusion
technically tagline should be "News through real & play money prediction markets, forecast aggregators, forecasting tournaments, superforecasters, & occasionally expert opinion surveys"
trade-off i'm hoping to make is a small loss in nuance, for a greater gain in readability. hope my audience forgives me
is the claim that Metaculus should've already anticipated all AI progress, to the extent that any updates are very small?
i don't think Metaculus needs to be omniscient to be useful. you're getting a prediction that reflects available info, at the cost of ~2 seconds glancing at a chart. (AlphaGeometry wasn't public; jump is new info being priced in.)
someone who doesn't need this could, presumably, create a more accurate forecast in less time?
(sorry for terse reply, trying to grapple with question in concise way + time poor)
Precise definitions:
AGI (@MatthewJBar for @metaculus):
We will thus define "an AI system" as a single unified software system that can satisfy the following criteria, all completable by at least some humans.
Able to reliably pass a 2-hour, adversarial Turing test during which the participants can send text, images, and audio files (as is done in ordinary text messaging applications) during the course of their conversation. An 'adversarial' Turing test is one in which the human judges are instructed to ask interesting and difficult questions, designed to advantage human participants, and to successfully unmask the computer as an impostor. A single demonstration of an AI passing such a Turing test, or one that is sufficiently similar, will be sufficient for this condition, so long as the test is well-designed to the estimation of Metaculus Admins.
Has general robotic capabilities, of the type able to autonomously, when equipped with appropriate actuators and when given human-readable instructions, satisfactorily assemble a (or the equivalent of a) circa-2021 Ferrari 312 T4 1:8 scale automobile model. A single demonstration of this ability, or a sufficiently similar demonstration, will be considered sufficient.
High competency at a diverse fields of expertise, as measured by achieving at least 75% accuracy in every task and 90% mean accuracy across all tasks in the Q&A dataset developed by Dan Hendrycks et al..
Able to get top-1 strict accuracy of at least 90.0% on interview-level problems found in the APPS benchmark introduced by Dan Hendrycks, Steven Basart et al. Top-1 accuracy is distinguished, as in the paper, from top-k accuracy in which k outputs from the model are generated, and the best output is selected.
By "unified" we mean that the system is integrated enough that it can, for example, explain its reasoning on a Q&A task, or verbally report its progress and identify objects during model assembly. (This is not really meant to be an additional capability of "introspection" so much as a provision that the system not simply be cobbled together as a set of sub-systems specialized to tasks like the above, but rather a single system applicable to many problems.)
Weak AGI (@AnthonyNAguirre for @metaculus):
For these purposes we will thus define "AI system" as a single unified software system that can satisfy the following criteria, all easily completable by a typical college-educated human.
Able to reliably pass a Turing test of the type that would win the Loebner Silver Prize.
Able to score 90% or more on a robust version of the Winograd Schema Challenge, e.g. the "Winogrande" challenge or comparable data set for which human performance is at 90+%
Be able to score 75th percentile (as compared to the corresponding year's human students; this was a score of 600 in 2016) on all the full mathematics section of a circa-2015-2020 standard SAT exam, using just images of the exam pages and having less than ten SAT exams as part of the training data. (Training on other corpuses of math problems is fair game as long as they are arguably distinct from SAT exams.)
Be able to learn the classic Atari game "Montezuma's revenge" (based on just visual inputs and standard controls) and explore all 24 rooms based on the equivalent of less than 100 hours of real-time play (see closely-related question.)
By "unified" we mean that the system is integrated enough that it can, for example, explain its reasoning on an SAT problem or Winograd schema question, or verbally report its progress and identify objects during videogame play. (This is not really meant to be an additional capability of "introspection" so much as a provision that the system not simply be cobbled together as a set of sub-systems specialized to tasks like the above, but rather a single system applicable to many problems.)
Important unsolved problem in mathematics (@MatthewJBar for @ManifoldMarkets):
For the purpose of this question, a mathematical conjecture is considered "important" if it appears on the list of unsolved problems maintained by the following sources:
The Open Problem Garden
The Clay Mathematics Institute (CMI)
The Unsolved Problems in Number Theory book by Richard K. Guy
These sources collectively provide a broad variety of conjectures across different fields of mathematics that are widely acknowledged as significant.
Taiwan's presidential election is this weekend...
Incumbent DPP's Lai Ching-te is the favourite, according to both prediction markets & opinion polls.
Markets disagree about TPP's Ko Wen-je. @Metaculus estimates only 1% chance of victory, contrasting with 21% odds on @Polymarket.
◦ Lai Ching-te (DPP) is ~72-83% likely to win according to prediction markets. Polls show Lai at 35%, a +6.5 point lead (weighted avg by @donovan_smith).
◦ Hou Yu-ih (KMT): ~14-20% to win according to markets, polling at 29%.
◦ Ko Wen-je (TPP): ~1-21% win according to markets, polling at 24%.
Early betting odds seem to favour @BillAckman & @NeriOxman:
◦ Neri Oxman's MIT degree is only 1% likely to be revoked by Jan 31 (@Polymarket)
◦ MIT president Kornbluth 38% likely to be ousted in 2024, up from 23% following the Business Insider Oxman story (@Kalshi)
◦ 43% chance that an MIT employee loses their job due to plagiarism in 2024 (@ManifoldMarkets)
◦ Bill Ackman's investigation likely to accuse (@ManifoldMarkets):
‣ 1-10 MIT faculty = 41% chance
‣ 10-20 faculty = 24%
‣ 20-50 faculty = 13%
‣ 100+ faculty = 9%
⚠️These are early odds (low liquidity still)