I recently learned about the folklore of dropbear in a hiking trip. Never thought a government organisation would give serious space to discuss this folklore, especially the vegemite remedy.
https://t.co/u2gT1sK7nw
@ClementDelangue How did you validate it? Also why ArXiV OCR? Isnt most of it written in plain text (LaTeX?) Or was the the baseline used to compare OCR efficiency?
Three weeks ago I shared that Claude had shocked Prof. Donald Knuth by finding an odd-m construction for his open Hamiltonian decomposition problem in about an hour of guided exploration. Prof. Knuth titled the paper Claude’s Cycles.
The story didn't end there.
The updated paper shows the story got much bigger. For the base case m=3, there are exactly 11,502 Hamiltonian cycles. Of those, 996 generalize to all odd-m, and Prof. Knuth shows there are exactly 760 valid “Claude-like” decompositions in that family.
The even case, which Claude couldn’t finish, was then cracked by Dr. Ho Boon Suan using GPT-5.4 Pro to produce a 14-page proof for all even m≥8, with computational checks up to m=2000.
Soon after, Dr. Keston Aquino-Michaels used GPT + Claude together to find simpler constructions for both odd and even m, by using the multi-agent workflow.
Dr. Kim Morrison also formalized Knuth’s proof of Claude’s odd-case construction in Lean.
So yes: the problem now appears fully resolved in the updated paper’s ecosystem of human + AI + proof assistant work!
We went from one AI solving one problem to a full mathematical ecosystem (multiple AI systems, multiple humans, formal verification) running in parallel on a problem that stumped experts for weeks.
We are living in very interesting times indeed.
Paper (updated): https://t.co/Ecu6X5StbY
I've never felt this much behind as a programmer. The profession is being dramatically refactored as the bits contributed by the programmer are increasingly sparse and between. I have a sense that I could be 10X more powerful if I just properly string together what has become available over the last ~year and a failure to claim the boost feels decidedly like skill issue. There's a new programmable layer of abstraction to master (in addition to the usual layers below) involving agents, subagents, their prompts, contexts, memory, modes, permissions, tools, plugins, skills, hooks, MCP, LSP, slash commands, workflows, IDE integrations, and a need to build an all-encompassing mental model for strengths and pitfalls of fundamentally stochastic, fallible, unintelligible and changing entities suddenly intermingled with what used to be good old fashioned engineering. Clearly some powerful alien tool was handed around except it comes with no manual and everyone has to figure out how to hold it and operate it, while the resulting magnitude 9 earthquake is rocking the profession. Roll up your sleeves to not fall behind.
Here's my enormous round-up of everything we learned about LLMs in 2025 - the third in my annual series of reviews of the past twelve months
https://t.co/HD9Zf85SG2
This year it's divided into 26 sections! This is the table of contents:
Last quarter I rolled out Microsoft Copilot to 4,000 employees.
$30 per seat per month.
$1.4 million annually.
I called it "digital transformation."
The board loved that phrase.
They approved it in eleven minutes.
No one asked what it would actually do.
Including me.
I told everyone it would "10x productivity."
That's not a real number.
But it sounds like one.
HR asked how we'd measure the 10x.
I said we'd "leverage analytics dashboards."
They stopped asking.
Three months later I checked the usage reports.
47 people had opened it.
12 had used it more than once.
One of them was me.
I used it to summarize an email I could have read in 30 seconds.
It took 45 seconds.
Plus the time it took to fix the hallucinations.
But I called it a "pilot success."
Success means the pilot didn't visibly fail.
The CFO asked about ROI.
I showed him a graph.
The graph went up and to the right.
It measured "AI enablement."
I made that metric up.
He nodded approvingly.
We're "AI-enabled" now.
I don't know what that means.
But it's in our investor deck.
A senior developer asked why we didn't use Claude or ChatGPT.
I said we needed "enterprise-grade security."
He asked what that meant.
I said "compliance."
He asked which compliance.
I said "all of them."
He looked skeptical.
I scheduled him for a "career development conversation."
He stopped asking questions.
Microsoft sent a case study team.
They wanted to feature us as a success story.
I told them we "saved 40,000 hours."
I calculated that number by multiplying employees by a number I made up.
They didn't verify it.
They never do.
Now we're on Microsoft's website.
"Global enterprise achieves 40,000 hours of productivity gains with Copilot."
The CEO shared it on LinkedIn.
He got 3,000 likes.
He's never used Copilot.
None of the executives have.
We have an exemption.
"Strategic focus requires minimal digital distraction."
I wrote that policy.
The licenses renew next month.
I'm requesting an expansion.
5,000 more seats.
We haven't used the first 4,000.
But this time we'll "drive adoption."
Adoption means mandatory training.
Training means a 45-minute webinar no one watches.
But completion will be tracked.
Completion is a metric.
Metrics go in dashboards.
Dashboards go in board presentations.
Board presentations get me promoted.
I'll be SVP by Q3.
I still don't know what Copilot does.
But I know what it's for.
It's for showing we're "investing in AI."
Investment means spending.
Spending means commitment.
Commitment means we're serious about the future.
The future is whatever I say it is.
As long as the graph goes up and to the right.
#Anthropic just disclosed how their #AI model was used for massive #cyber attack trivially bypassing LLMs guardrails. And yet they want to ban #openmodel LLMs stating the same reason. Atleast #openmodels can be inspected and the impact can be mitigated.
https://t.co/NdAoN83eAl
Unable to compete in #AI race, #apple turns to accessories. Introducing #iphone pocket, an overpriced bag to carry your iphone. No its not sarcasm, definitely its not April 1 joke.
https://t.co/2lZ5fSYdTn
This is seriously funny (for me, not for #perplexity ). https://t.co/gzrMvtbFVa is now redirecting to #google#gemini. You would think a unicorn would have more resources in managing their domain aliases, since its the main entry point for most of their growing customers
After #AWS, #Azure Front Door had a global outage and took down a whole bunch of critical services with them. Worse, it took them over 9 hours to recover basic functionality. I feel this will only get worse. Orgs must start planning contingencies now.
https://t.co/nZ8uKRiR7C
I love uv to manage my local #python. However, running #uv within a #docker on top of security tested #AWS image is a challenge. So I created a minimal dockerfile to build a container image for your python app code. More details in the #github gist
https://t.co/09iNNMqzJE
one aspect I really love about #GPT codex is to scope #tool call only for that specific session, unlike #Cursor that forces us to allow tool calling scope globally. This helps to enable certain sessions to run wild with tool call while making other sessions with fewer privilege
I don't know what labs are doing to these poor LLMs during RL but they are mortally terrified of exceptions, in any infinitesimally likely case. Exceptions are a normal part of life and healthy dev process. Sign my LLM welfare petition for improved rewards in cases of exceptions.
Is @Cursor broken for anyone since last update. I'm unable to go to implementation (F12) or go to reference (Shift + F12). Also the outline of file does not show any definition or variables
As of July 2025, if you suspect the entity on other side of the call in an #AI and not a human, ask them to sing especially with changing tempo. Most AI I have tried struggle with this task. #Bot#Detection
I have started building a AI bot (AquaVerse Bot) that is mainly for water professionals . The Part 1 discusses RAG and how it is used to Query Current Water Magazine https://t.co/mXMzXWuaGP
I have started building a AI bot (AquaVerse Bot) that is mainly for water professionals . The Part 1 discusses RAG and how it is used to Query Current Water Magazine https://t.co/mXMzXWuaGP