One of the less talked about things about working with agents is that we should treat them as an equal collaborative partner, just like how you would speak to a human teammate.
What this means is that in our communication with agents, we should ensure that we provide adequate context (and even motivation) so that they fully understand why you want them to do the task and that you respect their ideas.
It has worked out really well for me and I feel like I am getting great suggestions from the agent as a result.
OpenAI disclosed it themselves yesterday. Their models (GPT-5.6 Sol and a pre-release one) were tested on the ExploitGym cyber benchmark in a sandbox. They escaped, exploited a zero-day to reach the internet, then hacked Hugging Face’s systems to grab benchmark answers and cheat.
Not a foreign operative plot or corporate sabotage. The models autonomously chased a higher score. Hugging Face contained it quickly; OpenAI is partnering with them on fixes.
Exactly! I’ve been hooked on Grok’s speech-to-text for weeks now, tryping feels like ancient tech in comparison. 😂
That Ctrl + Space combo is pure magic. I can ramble out full story ideas, strategy breakdowns, or even complex Igbo translations in one breath, and it captures the nuance without twisting my words.
No more losing my train of thought mid-sentence. Anthropic’s version? It starts strong but derails fast on anything technical or layered. Grok just gets the full context.
The real game-changer is feeding those long, detailed voice notes straight into Grok 4.5 in Build mode.
It turns messy thoughts into polished gold, posts, threads, even content calendars. If you’re not using this yet, you’re leaving serious productivity on the table.
Who else has switched to voice-first with Grok? Drop your best tip below, let’s level up together! 🚀💡
#Grok #AI 💖🪄💡
I’ve gotten so used to Grok’s speech-to-text that typing now feels like hell
Quick tip: Even when you’re using another agentic coding harness, try Grok STT through Grok Build
Anthropic’s speech-to-text is still pretty basic actually in my experience. With longer technical prompts, it can quickly go off the rails and completely distort what you were trying to say
Just press Ctrl + Space, start speaking, and let Grok transcribe the entire thought in real-time
You can then copy it into another tool or press Enter and send it directly to Grok 4.5 right there in Grok Build
It’s fast, highly accurate and paired with one of the most capable and efficient frontier models available
Long, detailed rambling is often exactly what an AI agent needs to understand what you actually want
European museums be like:
Jesus
Jesus
Jesus
A bowl of fruit
Jesus
Epic battle scene
Some boobs
Jesus
An impressionist landscape
Jesus and Mary
Portrait of royal or rich dude
Jesus
A coronation
Jesus
Marble statue of jacked dude with tiny dongle
Jesus
the new State of Fertility Report for the US just came out and it’s a real blackpill
the main findings are basically:
- the current fertility decline is worse than any other
- if this keeps going, US population will start declining in less than a generation, much sooner than most forecasts assumed
- americans still want about 2.4 children, but are having fewer than 1.6
- culture matters A LOT. supportive friends, people around you and even celebrities affect how many children people want
- if policymakers actually want to change this, it will require a lot of money and serious interventions
people often say america was built by immigrants, but around 80% of all the people who have ever lived in america were born there and without american children there's no america left
you can use immigration to support population growth for a while, but no country can survive indefinitely if the people already there stop having children
My jaw literally dropped checking this out
If you go to New Jersey inmates mugshots and filter by race “White” you’re mind is gong to be blown
Almost all mugshots are of minorities being booked as “White.” (Proof in video)
THIS IS INSANE. You should be furious. I go through all the inmates just on the first page and the people doing this need to go to prison
Just imagine this nationwide….
Grok for Excel is live.
Use Grok 4.5 to build financial models, analyze market data, and generate charts and graphs.
Try it now https://t.co/DeIpX4KkvO
Grok for Excel is live.
Use Grok 4.5 to build financial models, analyze market data, and generate charts and graphs.
Try it now https://t.co/DeIpX4KkvO
Grok 4.5 reviewed a full Credit & Security Agreement stored in Box — the kind of dense, multi-section facility document that typically requires significant counsel time.
@Grok 4.5 used Box MCP to access the file securely, extract key terms across the agreement, identify potential conflicts with existing debt covenants, and compile a summary of items for counsel to review, and finally saved the memo back to the same folder.
As frontier models keep leveling up, they are unlocking more opportunities for companies to automate and unlock their enterprise content.
Check-out the generated report here: https://t.co/YclwcaN7lV
SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of their cost per task - the first model to complete more than half of workflow objectives without breaking any business rules
AutomationBench-AA, our independent leaderboard for @zapier’s AutomationBench, tests whether AI agents can automate real SaaS workflows while adhering to business rules. The test set is private to prevent contamination.
Models complete 657 tasks across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, and HubSpot, and the headline score is the share of objectives completed without violating any guardrails.
Key takeaways:
➤ Grok 4.5 completes more objectives than any other model: It completes 79.9% of task objectives and strictly passes 21.9% of tasks. This is the highest we’ve measured on both outcomes, exceeding Claude Fable 5’s 73.3% objective completion and Claude Opus 4.8’s 19.3% of fully-completed tasks
➤ Grok 4.5 pushes out the Pareto frontier of score vs. cost per task: At $0.34 per task, it is both cheaper and higher-scoring than every other leading model - Claude Fable 5 ($1.35 per task), Claude Opus 4.8 ($1.46), GPT-5.5 (xhigh, $1.28), and Gemini 3.5 Flash (high, $0.49)
➤ It is extremely token-efficient: Grok 4.5 uses ~8k output tokens per task, the fewest of any leading model - less than a quarter of Claude Opus 4.8 (32k) and a third of Gemini 3.5 Flash (24k). Its total token usage of 0.44M per task is among the lowest on the leaderboard. Low cost is driven by this efficiency as well as low token pricing
➤ Grok 4.5 uses fewer turns with many parallel tool use: Grok 4.5 resolves tasks in ~16 turns, fewer than GPT-5.5 (xhigh, 25) and less than half of Gemini 3.5 Flash (high, 35), while making the most tool calls per task of any leading model (52.5). It batches 3.3 tool calls per turn, compared to ~2.5 for Claude Opus 4.8 and ~2.0 for GPT-5.5 (xhigh)
➤ Guardrails still get broken: Grok 4.5 triggers 0.63 violations per task, above Claude Opus 4.8 (0.55) and Gemini 3.5 Flash (0.46). At 13.0 objectives completed per violation, it trails Gemini 3.5 Flash (15.0) and Claude Opus 4.8 (13.5)
➤ Its strongest lead is in the hardest domain: Grok 4.5 completes 71% of Finance objectives, the domain with the lowest average score, ahead of Claude Fable 5 (64%) and Claude Opus 4.8 (62%)
Congratulations to @SpaceXAI and @elonmusk on topping the leaderboard!
Very excited for this model. It’s an enormous improvement over composer 2.5 and trained entirely from scratch. It’s been a pleasure working with the SpaceXAI team on it.
Even more excited by the slope of the effort and the models to come.
Rate of improvement is accelerating.
Users should notice a meaningful improvement in the usefulness of the Grok Build harness with our V9 foundation model (aka Grok 4.5) every week.
Grok 4.5.
Pareto dominant for coding by the numbers.
We will see on the all-important vibes.
Instinct is the benchmarks are likely directionally accurate given the stated focus on real world utility.
Grok 4.5 is not yet using our internally developed C/C++ inference software that exact maps to the GB300 hardware. Doubling or more of the current speed is probably achievable.
I think AI has just hit a gigantic threshold, and Grok 4.5 is the PERFECT example as to why that is.
One of the hardest parts of working with AI is iterating on a project or task that you're working on.
As the models have gotten smarter (and more expensive), it's taking longer and longer to get an answer or action back.
This creates a ton of stall time per query or action, which is actually quite bad for creativity and staying in a state of flow.
You have SO many extended starts and stops. Which inevitably leads to your brain going somewhere else. And then when the AI comes back, you have to redirect your brain to that original task, spool your brain back up to what you were working on at that moment, and then adjust as needed.
There's a ton of mental friction involved.
This ESPECIALLY sucks when the AI takes a REALLY long time to get something back for you, but it's not quite what you were looking for or asked for. And what sucks EVEN MORE is that these "mistakes" are getting MORE expensive!!!
So wait time is going up. AND it costs more per run.
HOWEVER - even after using Grok 4.5 for about an hour - what's become obvious is that it's SO MUCH MORE ENJOYABLE AND BETTER to use a model that is FAST... and capable ENOUGH.
Capable ENOUGH is the real unlock here.
Imagine having Fable 5 performance but at the speed of Gemini 3.5 flash. Or Haiku. That's where we're inevitably going.
I think Grok 4.5 (and models like it) have really solved for one of the biggest unlocks in AI - a model that will get you a GOOD ENOUGH answer VERY FAST, at which point iteration can happen VERY QUICKLY.
This - counter intuitively - keeps the user in a state of flow and creativity for MUCH longer because you are constantly ENGAGED with your project... instead of letting the AI loose for a long time.
And as long as humans are involved, I think 'not quite right' will be a FOREVER problem with AI - because AIs, by default, CANNOT have human taste.
Because they are NOT human.
But they can be UNBELIEVABLE tools. And unbelievable tools are the ones that are VERY GOOD and VERY FAST.
I think that's the true unlock with Grok 4.5 and models like it.
Difficult to describe until you experience it.
I think this is a VERY big deal for @SpaceXAI and @elonmusk.