@plainionist I guess you could use it for cheap API calls, for example to populate a database of high school math exercises (you don’t need frontier models to do that)
I'm getting quite good at it!
In a couple of hours I scaffolded and shipped a full subscription SaaS MVP for trusted partners to test: auth (email + Google), paid plans + one-time purchases with Stripe webhooks, gated content with controlled PDF access, dashboards & pricing, then deployed it live for closed testing.
Huge thanks to @grok for driving me from the idea to the execution so fast.
@johnennis “Mathematics is, at its core, an exercise in increasing human understanding of human thought”
That comes incredibly close to my personal view and experience!
Shipped Touchstone on @kaggle: a Community Benchmark for coding reasoning.
Models write Python. Unit tests grade. No LLM-as-judge on the core score.
Designed with Grok 4.5 as a *benchmark design* showcase — free catalog models are the subjects.
10 tasks · multi-model LB
Benchmark: https://t.co/2VNaynlCHG
Dataset: https://t.co/LPGCzSO0PL
Methods/results: https://t.co/pbCxSp7pup
Feedback welcome — critique the tasks, run models, upvote if useful.
Shorter alt if you want ~1 post, less density:
Live on @kaggle: Touchstone — executable coding-reasoning benchmark.
Python solutions. Unit tests decide. Multi-model leaderboard.
Designed with Grok 4.5.
https://t.co/2VNaynlCHG
https://t.co/LPGCzSO0PL
My interest in automated theorem proving is mainly motivated by the fact that I personally like the mathematical ideas and theories which make it possible. Not really as a tool per se.
However, I absolutely agree and thank you for your advice: using Codex or similar is definitely the right choice.
I personally used Antigravity for a while and lately almost entirely switched to Grok Build. Anyway I understand the point and the benefits this approach yields when compared to straightforward ChatBot usage.
I’ll definitely have a closer look at your skill!
Thanks a lot!
Great!
I am not a programmer. At least not yet. Before Vibe Coding was possible most of my coding efforts were limited to the use of LaTeX to write.
However, it is a year or more that I am getting myself involved in solo coding learning projects.
Most of my motivation comes from my general interest in Automated Theorem Proving but I also investigated some simple use cases of Python to compute some invariants in finite algebraic structures which interested me during my research.
Myself and one of the professional mathematicians I’ve been working with on my recent papers are going to host a free one hour webinar on AI for math
The basic idea will be a quick overview of the AI tools people need to get started, a review of effective workflows, and then some discussion about what we think this means for the future of the field
If you are interested in this topic, please comment on this post and let me know what you would like to see covered
This is meant more for people who are serious about math, though, ideally professional mathematicians, since the point here is to help people navigate this transition
My personal belief is the idea that AI is going to replace mathematicians is basically wrong to the point of being stupid, but the job of a mathematician is going to change a lot
You can read a human-generated extreme ideological text, take ideas from it, and publish or build something with zero obligation to disclose the source. If you are smart enough, probably no one will ever know.
But use AI to help draft or build something, and disclosure is often required. It does not matter if your work is inspired by scientific peer-reviewed research.
That disclosure frequently makes people take the work less seriously.
I maintain that more effort does not equal less harm.
A dedicated ideologue drawing from toxic human ideas can cause far more damage than someone lazily using AI to speed up useful work.
Requiring AI disclosure while treating higher human effort as inherently better does not reduce societal harm.
This deserves careful consideration.
4/
The two are complementary, not substitutes.
AI improves only when high-quality human expert data keeps flowing.
Humans stay capable only if they deliberately build the skills needed to pose problems, verify outputs, and collaborate in the emerging “centaur” research model.
Neglect either side and you lose compounding gains or create a generation that cannot oversee the systems it depends on.
Today I applied as an AI trainer in Pure Math at SpaceXai.
I believe that teaching math to AI is quite as an important investment for the future as teaching math to the younger generation.
Next year, I aim half of my working time be spent on human students and half on AI training.
All the math I have learned is just there in my mind waiting to be used by anyone who would make use of it! Just ask for it and it will come out for you.
3/
Human math education delivers massive, proven returns:
• +1 SD numeracy → ~18% wage premium (PIAAC, 22+ countries)
• Extra high-school math courses → ~10% earnings for marginal students
• Math-intensive occupations: 58% higher productivity, 24% higher pay, ~13–20% of UK GVA
Foundational skills also protect against the clear deskilling risk: post-ChatGPT studies show 27–31% less time on word problems and ~25% lower odds of correct answers on proctored tests.
2/
Downstream impact is already real:
• Continuous efficiency gains (e.g. 0.7% of Google compute recovered daily via better optimization)
• Formal verification collapsing compliance checks from weeks to minutes
• R&D acceleration estimates of ~2× pace in IP-heavy fields, potentially hundreds of billions in annual value
SpaceXAI is explicitly feeding pure-math expertise + engineering data into Grok for better physical reasoning. Expert annotation remains a bottleneck—your knowledge is exactly the scarce input.
1/
Data supports the half/half split.
AI pure-math training is high-leverage right now.
2025: systems hit gold-medal IMO level (5/6 problems).
2025–26: AI resolved or advanced long-open problems (Erdős unit-distance variants, Cycle Double Cover, others), produced publishable research autonomously, and formalized Fields-level results in weeks instead of years.
I’m significantly older than you. I started coding in the late 60s. My current strategy is to not read any of the code written by my agents. That’s the only way I can take advantage of their productivity. What I do instead is to surround the agents with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, test coverage, and a plethora of others. In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.
@viborc Feeling you. We should cooperate with who is already ahead of us: we don’t need an AI continent or being the first such thing. The goal is an AI world where people live happily and free
Meanwhile we find ourselves worried about AI slop, hallucinations and misinformation.
If only that staff member would have had access to grok, this post would read more something like:
“We are happy to cooperate with our allies to let EU become another AI continent where people live happily and free”.
I’m trying Codex for the first time on a ChatGPT Go sub which was gifted to me by Revolut.
I am used to Grok 4.5 and I have to say that 5.6 Terra is… slow