Estamos no ponto em que usando IA eu me sinto um capataz tentando ter controle e garantir que meus funcionários estejam realmente trabalhando e não só me enrolando
Juntou muita gente, não esperava, tanto que nem coloquei a legenda do gráfico. Tem estabelecimentos faltantes, por ele só é possível ter noção da escala.
Se quiserem ver algo legal: Criei uma ferramenta que compara preços de m2 entre Rio, Nit, SP e BH.
https://t.co/Zq58EbEtyZ
2006: esse ano eu vejo a seleção ser campeã
2010: esse ano eu vejo a seleção ser campeã
2014: esse ano eu vejo a seleção ser campeã
2018: esse ano eu vejo a seleção ser campeã
2022: esse ano eu vejo a seleção ser campeã
2026: …
As new models emerge and benchmark scores keep improving, there's a temptation to reach for the flashy new model. But performance may degrade on bespoke tasks, and some models might surprise you if you give them a chance. Private benchmarks are key to these discoveries.
I prompted a range of models to produce research designs sitting at the frontier of questions in political psychology and behavior. Through blind ratings of outputs, I selected those that mirrored the design choices I would personally make and could serve as research design co-partners. Contrary to my expectations, models like Opus 4.8 that I routinely use for coding produced overly elaborate designs. Smaller proprietary models, and the much-hyped GLM-5.2, generated viable designs.
On tasks where I have ground-truth values, smaller open-source models like Qwen3.6 outperformed some of the proprietary models we might reach for when classifying text or generating stimuli.
Cutting through the sea of AI hype we encounter on this platform and elsewhere can be exhausting. Building private benchmarks can ground our judgments.
There are many ways to do this. @HannoHilbig and @dasanaike have been testing model performance on tasks like text classification and OCR, where you might already have gold-standard classifications or ratings. For taste-based judgments, you can ask a model to generate a research design, measurement approach, or recruitment strategy; select the outputs that meet your standards; and build a rubric that can be automatically applied to future models. @ahall_research argues that we should all be developing private benchmarks, and he's right. The field moves too fast for any single leaderboard to tell you what works for your problem. Only you can do that.
I'm proud to finally announce the first release of the Small-Area Global Elections (SAGE) dataset, encompassing global, granular, standardized, geocoded, polling station or equivalent-level election results for 110 countries, conditionally accepted at Nature Scientific Data. (1/n)
To keep pace with AI progress, we're advancing how we study Claude's economic impact.
Hourly sampling and survey data show us how the cadences of life shape usage, what people produce with Claude, and how perceptions of AI's impact may be changing. https://t.co/Waov1B6iG1
We cannot consider #AI to be morally neutral. In reality, every technical tool embodies choices and priorities through what it measures, ignores, and optimizes, and how it classifies people and situations. Ethical discernment cannot be limited to asking whether we are using a system for good or bad purposes. It must also examine how that system is designed and what vision of the human person and society is embedded in the data and models that guide it. #MagnificaHumanitas
URGENTE:
Ancelotti disse a Endrick que ele será TITULAR contra o Haiti.
Com isso, o trio de ataque do Brasil deve ser formado por Vini Jr, Raphinha e Ancelotti.
Alibaba Qwen3.7 slowly fading into irrelevance at the frontier due to proprietary stance.
In it's place we have Minimax M3 and... *checks notes* Rio 3.5 397b, made by the municipal IT company of Rio de Janeiro's city government.
https://t.co/JgIJYVhoEi