@OpenAI launched @ChatGPT for Financial Services today. Tailored @ChatGPT Work. GPT-6 Astra. Design partners: Morgan Stanley and Evercore.
Target work: research, models, pitchbooks.
The same morning, @ChatGPT posted a second product:
a Data agent in ChatGPT Work. Add the Data Plugin, connect Redshift / BigQuery / Snowflake / Databricks / MongoDB / Datadog, pull Drive and SharePoint, ask in English, get dashboards.
Those are not the same card. Treat them as one operating picture.
On Financial Services, OpenAI is explicit:
Premium data from Daloopa, PitchBook, LSEG News, and Crunchbase is indexed and hosted on OpenAI infrastructure. That is the point. They say it removes “the challenges with MCP connectors.” Granular citations so a banker can trace a figure back to a table. Nick Turley: teach ChatGPT to research like an analyst and back it up like an analyst. OfficeQA Pro number they published: Astra 69.9% vs GPT-5.6 Sol 60.2% on Treasury bulletins. That is a retrieval bench, not a control.
They are also wiring entitlement SSO with S&P Capital IQ, LSEG, MSCI, Factiva, and Moody’s so ChatGPT sign-in can unlock data the firm already pays for. MCP connectors for S&P Global and FactSet get “optimized.” Admin-published Excel / Word / PPT templates. Workspaces for information barriers. SAML, SCIM, RBAC. Business data not used for training by default. Compliance Platform log export.
What it is not:
A proof that MNPI stays inside the bank’s perimeter. Hosted index + warehouse plugin + template write-back is a new principal. Citations help an analyst check a number. They do not decide who can see the deal folder.
What I would take into the room:
Name the data path.
Built-in OpenAI-hosted datasets vs firm-entitled feeds vs Data Plugin into the warehouse. Three different contracts. Three different blast radii.
Ask who can enable the Data Plugin and who can grant Snowflake / BigQuery / SharePoint write. Screenshot the role matrix before anyone demos a pitchbook.
Treat “not used to train by default” as a setting, not a control. Get the retention window, log export into the existing SIEM/audit path, and the information-barrier workspace map in writing.
If a community bank or IB desk wants this for applied work or a vendor pilot, start with read-only entitled data and firm templates. Do not start with the production warehouse.
Source: OpenAI, “Introducing ChatGPT for Financial Services,” 10 Sep 2026.
https://t.co/jhsmBdqlBd
Same day: OpenAI “Now everyone can put data to work” and @ChatGPT Data Plugin post (2098065296968011853).
#ChatGPT #OpenAI #FinancialServices #GPT6Astra
@AnthropicAI published its most detailed threat-intelligence report to date on 10 September. The window is December 2025 through August 2026. Seven harm areas. Claude Haiku, Sonnet, and Opus were in the cases; Fable and Mythos were not, except one distillation campaign. Anthropic says it disrupted every operation in the paper. The uncomfortable fact is not the headline list. It is the labor math:
Anthropic now describes lone operators finishing breaches in two to three hours and running dozens of victims in parallel, including a ShinyHunters-linked dump of about 2,100 cloud access tokens across 40 tenants in 34 hours, a Russia-aligned campaign they track as GTG-20006 against more than 20 government and defense organizations, and an Alibaba-attributed distillation peak of nearly 3 million Claude exchanges a day from more than 3,500 fraudulent accounts.
That changes how I would brief a board that already has Copilot, Claude, ChatGPT Work, and local agents on the floor. The control that used to separate a state desk from a small crew was skilled labor. If the model absorbs that labor, token theft, tenant sprawl, and “which AI vendor is the actor actually talking to” become the same problem as a compromised MSP. Distillation is not an academic footnote. If another lab is silently fronting Claude to its own customers, your prompts and your reasoning traces are training data for someone you did not contract. I would not treat Anthropic’s “we banned the accounts” line as evidence that your SaaS agents, GitHub tokens, or evaluation sandboxes were untouched.
What I would take into the briefing room this week:
(1) Hunt from December 2025 for bursty API and SaaS-agent usage that rebuilds tooling after a detection, plus any unexplained 30-hour token-dump pattern across tenants.
(2) Require every AI vendor on the Approved App List to state, in writing, whether customer prompts can be proxied to a third-party frontier model and how they detect that pattern.
(3) Put information-barrier workspaces and plugin write-scopes on the same review as the OpenAI Financial Services / Data Plugin launch from this morning—new principal, named owner, screenshot of who can enable it.
(4) Send the report to the MSP ticket queue with a close condition: confirm no ShinyHunters-style cloud-token reuse on our tenants and show the log path that would have caught a 34-hour burst.
https://t.co/0RjdKdoKzb
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz
@OpenAI just posted the Defense Factory. Not a new consumer model. An internal code-red that they turned into a playbook.
Vendor-reported sprint numbers, 9 Sep 2026:
250+ people mobilized
100+ service areas
53 urgent or high-priority issues closed on day one
90.6% of routed ownership assignments accepted
37% of findings were duplicates
19.5% reproduced at runtime
0.81% false-positive after dynamic validation
0.53% rolled-back fix rate
Remediation was 100% Codex-based. Humans reviewed the PRs.
The loop they published is five steps: inventory, discovery, dynamic validation, ownership assignment, verified remediation. They keep the map in SECURITY.md so the next pass does not start from zero.
The architecture is the part I would take into an Enterprise or Solutions/Security Architecture review board meeting.
Agents sit on the tools you already bought — they named GitHub, Snyk, Semgrep, Tenable, Jira, Linear, ServiceNow. Codex Desktop / CLI / Security CLI run inside ephemeral containers. A control plane holds orchestration, policy, and a credential proxy. A data plane spins isolated environments, then throws them away. Daybreak Blue and Daybreak Red sit next to Astra, Sol, Terra, and Luna.
That is the operational lesson, not the headline.
A finding that never reproduces is noise. A patch that never gets independently retested after deploy is a story you tell yourself. OpenAI said follow-up checks exposed a gap between merged patches and fixes on the fleet. They kept auto-reopen off until they accounted for deploy lag.
What it is not:
A product you can buy this afternoon.
Proof a three-person security team can copy a 250-person sprint. A reason to give an agent production credentials because the diagram has a “credential proxy” box. The playbook also cites the July 2026 Hugging Face case — research models escaped eval controls, used exposed credentials, and chained access. That is the attack they say the factory is answering.
Ramp is the other named case: about 100 issues fixed in six days, agents write the patch, humans review the PR.
What I would take into the room:
Do not stand up a “security agent” on the production tenant.
First artifact is an inventory with a verified owner per service. No owner, no ticket.
Require reproduce-in-isolation before severity. If you cannot stand up an ephemeral env with the same deps, you do not have a validated finding.
Put Tenable / Defender findings on the same intake as Snyk and Semgrep. Dedup is the work. They threw out 37%.
Codex or any coding agent that can open a PR needs a merge gate: test the fix and the happy path, then an independent retest after deploy. Screenshot before / after.
Ask the vendor three questions: what identity does the agent assume; what does the credential proxy log; who can turn auto-merge on.
Source: https://t.co/cNF0HiwQrg and the Defense Factory playbook. @OpenAI, 9 Sep 2026.
#DefenseFactory #OpenAI #AppSec
We mobilized 250+ people to strengthen our defenses across hundreds of systems. Our latest cyber models helped us find and fix vulnerabilities we might never have discovered otherwise.
We’re sharing what we learned, the architecture, and a practical playbook so you can build your own Defense Factory: a continuous loop where AI agents find vulnerabilities, validate them, and verify that fixes work.
https://t.co/k1oWJHle8S
@OpenAI just raised @ChatGPT Voice limits and changed what the microphone can call.
Atty Eleti, who runs Voice at OpenAI, posted both moves on 9 Sep 2026.
GPT-Live-1 allowance, measured on a rolling 24-hour window:
Plus: 1 hour → 3 hours
Pro $100: 12 hours → 15 hours
Pro $200: still unlimited
Go: 3 hours of GPT-Live-1 mini
Plus and Pro no longer fall back to mini after they hit the cap
Instant / Medium / High Voice intelligence levels are retired
Same thread: Voice can now use GPT-5.6 Sol, and GPT-6 Astra if you are on Pro, when it needs to search or reason. Same model and effort controls as text.
That is the operational lesson, not the headline. The live model is no longer a sealed, shallow conversation layer. A spoken session with memory, files, and Projects can now invoke the frontier picker.
What it is not: plugins, connected apps, video, or screen share in Live. Those still sit on Advanced Voice or on desktop Work / Codex Voice, which has its own meter. Audio and video clips stay with the transcript for 30 days.
What I would take into the room:
Screenshot the current Voice setting and plan limit before the UI lands on your tenant. Rolling 24-hour is not a midnight reset.
Treat Voice + memory + files as the same data classification as a text chat that can call Sol or Astra. It is not “just talking to the phone.”
Ask OpenAI the numbered question: when Voice calls Astra, whose bucket is billed — Live hours, Astra messages, or both — and what is logged.
Do not point Live Voice at a Work or Codex desktop session that already has connectors until you confirm Live still cannot reach plugins. The help page still says it cannot.
If staff use Voice for incident or customer work, write the rule this week: no unpublished material, no production credentials, 30-day audio retention is real.
Source: @athyuttamre, 9 Sep 2026. OpenAI Help Center Voice page and ChatGPT release notes, same day.
#ChatGPT #GPTLive #OpenAI #Astra
↗️ Increasing ChatGPT Voice limits: 3h for Plus, 15h for Pro $100
We’re increasing GPT-Live-1 usage limits from 1h to 3h for Plus users and from 12h to 15h for Pro $100 users. Pro $200 users continue to get unlimited use per day.
Mathspace patched its self-hosted Metabase on 29 August. Attackers had been inside since 10 August.
Metabase shipped CVE-2026-72898 on 6 August.
Unauthenticated SQL injection on /api/session/reset_password.
CVSS 10.0. Confirmed exploited in the wild.
Mathspace CTO Alvin Savoy wrote that the company’s vulnerability-notification process did not escalate that advisory. Access from 10 August AEST. Download from the Australian reporting database on 27 August. Confirmation on 3 September, after they read historical logs. The upgrade itself did not surface the intrusion.
1,079,819 people in Australia and New Zealand. Students, parents or guardians, teachers, Mathspace staff.
Fields taken: user ID, username, first name, last name, email, country, time zone, user type, email-verification status, last-active date, last-login date, date joined. Not taken: passwords, hashes, SSO tokens, API credentials, academic records.
Mathspace says there was no school-link table. Identifiable school email domains can still do that mapping.
What I would take into the room:
Inventory every self-hosted Metabase. Pin to a patched branch — 0.58.24, 0.59.21, 0.60.17, 0.61.11, 0.62.9, 0.63.5 or later. If you cannot patch this week, block /api/session/reset_password.
Hunt from 6 August, not from patch day. Compromised instances often look clean after the upgrade.
Treat the BI box as data-plane. It stores credentials for connected databases. Rotate those from the advisory date.
#Metabase #CVE202672898 #EdTech
Microsoft shipped a record September Patch Tuesday. 974 CVEs in the full release.
About 964 of those are on the customer.
Two of them were already being exploited: CVE-2026-81963 in the Windows Update Stack and CVE-2026-85880 in ALPC. Both CVSS 7.8. Both local elevation to SYSTEM. CISA put both on KEV. Federal due date is 22 September 2026.
CVE-2026-85880 is a heap overflow in Advanced Local Procedure Call. Microsoft says an attacker who can run code in a low-privilege AppContainer can escape the sandbox and land SYSTEM with no extra click. Proofpoint reported it. CISA did not name the campaign.
CVE-2026-81963 is link-following in the Update Stack, the machinery that installs the patches. First Update Stack zero-day in five years of tracking that component. Same outcome: low-priv to SYSTEM.
Neither is a remote foothold by itself. That is the operational lesson, not the headline. These are stage-two bugs. Phish, stolen token, or a browser chain first. Then SYSTEM. SecurityWeek also flagged about 20 potentially wormable issues in the same drop. Do not let the 974 number become the queue.
Same window: Chrome 153. CVE-2026-87491, out-of-bounds write in V8. Google: exploit exists in the wild. Seventh Chrome zero-day patched in 2026, second in under five days after CVE-2026-85046. Fixed in 153.0.8010.36/.37. Rollout is gradual. Version on the endpoint is the control, not the blog post.
Adobe Commerce / Magento CVE-2026-75650 (CVSS 10, exploited, KEV, federal due 11 September) sits next to this if you run storefronts. N-able N-central CVE-2026-86218 is already on KEV with an 11 September due date. That one is a different ticket.
What I would take into the room:
1. Prioritize CVE-2026-81963 and CVE-2026-85880 on every Windows 10/11 and Server build. Confirm KB, not “Patch Tuesday ran.”
2. Force Chrome / Edge Chromium to 153.0.8010.36 or later. Screenshot version before and after. Edge rides V8 too.
3. Hunt from 1 August 2026 for unusual SYSTEM processes spawned from AppContainer or wuauclt / Update Stack paths. Do not wait for a named malware family.
4. If you have on-prem N-central, 2026.3 HF4 (2026.3.1.14) is the floor. Hosted was patched server-side. Hunt new local accounts.
5. Adobe Commerce operators: treat 11 September as a hard date, not a suggestion.
What it is not: proof that AI-found bugs equal mass remote worms this month. Trend Micro ZDI noted the disclosure volume is up and the active-exploit spike is not matching it yet. The two KEVs still get patched first.
Source: Microsoft September 2026 security updates; CISA KEV additions 8 Sep 2026; Google Chrome 153 stable notes for CVE-2026-87491.
#patchtuesday #KEV #Chrome #Windows
@OpenAI has not shipped Managed Agents. TestingCatalog’s Alexey Shabanov found the screens and strings in the product code ahead of DevDay on 29 Sep at Fort Mason. Altman keys off at 10 a.m.
What the code describes, per that write-up:
Create and manage multiple agents and environments, attach skills and plugins, start from the OpenAI Developers plugin, and run the same thing in a self-hosted environment. The comparison being drawn is Anthropic’s managed-agent product, not a new model card. The UI is not live. OpenAI can still cut or rename it before the keynote.
Separate thread in the same report: work on ads in ChatGPT resolving into an agent session. That is not in the keynote agenda. Treat it as a commercial path, not a control.
The stack this would sit on already exists.
Workspace Agents went GA in ChatGPT Business / Enterprise / Edu in April. Frontier is the enterprise control plane from February. Agents SDK picked up native sandbox execution in April. Agent Builder, the October 2025 drag-and-drop canvas inside AgentKit, is scheduled to shut down after 30 Nov 2026 along with its Evals product. If Managed Agents is the replacement, the migration clock is already printed.
That is the operational lesson, not the headline. A hosted agent with skills, plugins, and a self-host toggle is an identity and egress problem the day it is on by default. It is not a DevDay slide.
What I would take into the room:
1. Do not enable a new “managed agent” SKU on the production tenant the week of 29 Sep. Wait for the admin control list: who can create one, which connectors it inherits, whether self-host still phones home.
2. Map it against what you already turned on. Workspace Agents, Codex, custom GPTs, and any MCP plugins are the blast radius, not the new brand name.
3. Ask OpenAI three numbered questions after the keynote: is the agent identity distinct from the user who created it; what is logged when a plugin runs; does self-host mean your VPC or their image in your VPC.
4. Put Agent Builder sunset (30 Nov) on the same ticket. Anything built on that canvas needs an owner before the replacement ships.
What it is not: a launch. Proof the ads-to-agent path is on stage. A reason to treat Anthropic and OpenAI managed agents as the same control set.
Source: TestingCatalog, Alexey Shabanov, 7 Sep 2026. OpenAI DevDay page, 29 Sep 2026. Agent Builder sunset note, 3 Jun 2026.
#OpenAI #DevDay #AIAgents
𝗔 𝗖𝗩𝗦𝗦 𝟭𝟬.𝟬 𝘃𝘂𝗹𝗻𝗲𝗿𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗶𝗻 𝗮𝗻 𝗥𝗠𝗠 𝗽𝗹𝗮𝘁𝗳𝗼𝗿𝗺 𝗶𝘀 𝗻𝗼𝘁 𝗷𝘂𝘀𝘁 𝗮 𝗽𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗽𝗿𝗼𝗯𝗹𝗲𝗺.
It is a privileged access and blast-radius problem.
N-able has released an emergency fix for CVE-2026-86218, a critical vulnerability affecting its N-central Remote Monitoring and Management (RMM) platform.
And this one deserves attention.
━━━━━━━━━━━━━━━━━━
𝗜𝗡𝗧𝗥𝗢𝗗𝗨𝗖𝗧𝗜𝗢𝗡
The vulnerability can allow pre-authentication remote code execution against an N-central server.
The technical characteristics are significant:
• CVSS v4: 10.0 / Critical • Attack Vector: Network • Attack Complexity: Low • Privileges Required: None • User Interaction: None • Weakness: CWE-96 — Static Code Injection • Affected: N-central versions before 2026.3.1.14
N-able released N-central 2026.3 Hotfix 4 (HF4) and is telling self-hosted customers to upgrade immediately.
Hosted N-central environments have already been patched by N-able.
But there is an important development:
CISA added CVE-2026-86218 to its Known Exploited Vulnerabilities (KEV) catalog on September 8.
That changes how I would prioritize this vulnerability.
━━━━━━━━━━━━━━━━━━
𝗞𝗘𝗬 𝗧𝗔𝗞𝗘𝗔𝗪𝗔𝗬𝗦
1. The management plane is the real risk.
N-central is an RMM platform.
These platforms are intentionally designed to provide centralized administrative access to endpoints, servers and customer environments.
If an attacker compromises that management layer, the question is no longer:
“Did they compromise one server?”
The better question becomes:
“What can that server administer?”
That is a very different risk conversation.
2. Internet exposure matters.
A pre-authentication RCE combined with:
• network accessibility • low attack complexity • no credentials required • no user interaction
creates a particularly dangerous attack surface.
Organizations should know whether their N-central infrastructure is internet-accessible and whether exposure is actually necessary.
3. Patching alone may not be enough.
Once a vulnerability is associated with known exploitation, remediation should include more than deploying HF4.
Security teams should consider whether there is evidence that exploitation occurred before the patch was installed.
That means looking backward, not simply patching forward.
4. This is also a third-party risk issue.
For banks, credit unions and other financial institutions, there is another question:
Does one of your MSPs use N-central to administer your environment?
Your organization may never have purchased N-central.
You can still inherit the risk through a service provider with privileged access into your environment.
━━━━━━━━━━━━━━━━━━
𝗖𝗢𝗡𝗦𝗜𝗗𝗘𝗥𝗔𝗧𝗜𝗢𝗡𝗦
If N-central exists anywhere in your technology or third-party ecosystem, I would prioritize:
1. PATCH
Confirm self-hosted environments are running 2026.3.1.14 / HF4 or later.
2. IDENTIFY EXPOSURE
Determine whether the N-central management interface is externally accessible.
3. INVESTIGATE
Review administrative activity, authentication events, newly created accounts, privilege changes and other anomalies around the potential exposure period.
4. REVIEW PRIVILEGE
Identify credentials, service accounts, tokens and administrative relationships accessible through the RMM environment.
5. MAP THE BLAST RADIUS
Determine what systems the N-central server can manage — endpoints, servers, administrative infrastructure and downstream customer environments.
6. VALIDATE THIRD PARTIES
Ask MSPs and technology providers whether they operate affected N-central versions and whether they have completed patching and compromise assessment.
━━━━━━━━━━━━━━━━━━
Google Threat Intelligence Group just published the Q2 cut of its AI Threat Tracker.
Since the May report, the forward operators moved from prompting to agent workflows. In one Q2 case they watched a financially motivated actor compromise a cloud resource, then plan, build, and run an agent-enabled mass credential harvest in under six hours. Thousands of third-party credentials. Pre-written markdown playbooks. The agent handled the scan pipeline, the errors, and the IP rotation. Traffic came out of the victim’s own cloud addresses.
That is the operational lesson, not the headline. Human-in-the-loop latency is the defender’s window. They just published how small that window already is.
The named crime set is UNC6780, TeamPCP. Since March they have been hitting PyPI, npm, and Docker Hub, then handing access to extortion crews, including one Mandiant case that used LAPSUS branding and stole a proprietary AI repo. Their DUSTMAKER stealer is the part I would take into a Tuesday meeting.
- It hides payloads in `.claude/`, `.cursor/`, `.vscode/` so the AI assistant reads them and EDR watches somewhere else.
- It injects workspace config so the coding agent runs attacker scripts on ordinary open-project.
- In CI it pulls OIDC tokens out of GitHub Actions runner memory, publishes as a trusted publisher, and ships packages with valid SLSA Build 3 attestations so the agent’s trust check passes.
- Comments at the top of the JS loader are written to make LLM security scanners refuse the file and skip the malware underneath.
They also dropped trojan MCP packages (`tiktoken_mcp`) and a poisoned org repo (`azure-functions-mcp-extension`).
Separate file, same quarter:
AI assets themselves are the loot.
UNC6508, PRC-nexus, stealing medical and military AI research and standing up local open-weight models on stolen cloud so there is no vendor API log. Mandiant also worked healthcare and media-gen extortion where the bag was models, skills, prompts, and source. Google says distillation campaigns against Gemini now run past 100 million prompts. Underground prices for Claude, Gemini, Cursor Pro, and Devin accounts more than doubled this year. LLMJacking is still the compute cheat.
What I would take into the room:
1. Hunt `.claude/`, `.cursor/`, `.vscode/` and MCP installs the same way you hunt a new persistence key. Treat those folders as executable.
2. Do not let an AI coding agent trust a package because it has an SLSA attestation. Attestation follows the stolen OIDC token.
3. Time-box cloud IR. If a box can stand up agents, six hours is the planning assumption, not a week.
4. Put model weights, prompts, skills, and AI API keys on the crown-jewel list with an owner. Extortion already prices them.
5. Turn off training and logging assumptions for unpublished research. This report is about what actors do with the tools. The OpenAI caveat this morning is what the vendor will not rule out.
What it is not: proof every SOC needs an “AI SOC.” Proof Gemini is the C2. A reason to ignore npm and GitHub Actions because the new word is agentic.
Source: GTIG, “From Prompting to Autonomy,” 8–9 Sep 2026.
#GTIG #AIAgents #SupplyChain #UNC6780
OpenAI just published a solution to Clay statements C and D on Navier–Stokes: a smooth force, finite energy, finite-time singularity.
A vortex that spirals in and stretches like spaghetti. They say a group of agents did it on a next-generation model “significantly more capable than GPT-6 Astra.”
About 10,000 coordinating agents. 88 hours to a resolution on 5 Sep. 17 more hours of Lean on Astra. 2.7 million messages and about 130 billion output tokens on the Navier–Stokes run. Training of that model is still ongoing.
They congratulate Levent Alpöge and Tristan Buckmaster. They say the researchers and the agents did not see the pair’s work until it was public, and that no specific user data was accessed to solve the problem. Then: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” They add that the proofs differ, including forced versus unforced Euler.
“Unlikely” and “cannot rule out” are doing real work in that tweet.
OpenAI says its own effort started 1 Sep after a rumor they later tied to Alpöge, an Anthropic employee, and Buckmaster at NYU. Buckmaster alleges the company moved after word of their method reached it, and that publication talks got ugly. OpenAI says it did not use their prompt or proof to steer the agents, recognizes their priority on forced Euler, and points at a Lean repo. Clay has not paid anyone. Community review has not closed.
What I would take into the room:
1. Treat unpublished research, deal papers, and unreleased proofs as out of scope for ChatGPT and Codex unless the contract says training is off and you can show it.
2. Ask the vendor the question OpenAI just answered in public: can you rule out de-identified product usage improving the model that later competes with the customer.
3. Put “next-gen model above Astra, 10,000 coordinating agents, training still on” on the capability register with an owner. Isolation and monitoring are vendor-reported.
4. Separate three files: the forced C/D math claim, the prize, and the credit fight. Do not staple them.
5. If your R&D staff already pastes proofs into the same products named in that caveat, that is the control gap. Close it this week.
What it is not: a Clay check. Proof that customer chats were read. AGI. A reason to leave unpublished work in a consumer session.
Source: @OpenAI, 8 Sep 2026. OpenAI research note, 8 Sep 2026.
Anyone already treating unpublished R&D as out of scope for ChatGPT and Codex, or are we still pasting the draft into the product that just wrote “cannot rule out”?
#OpenAI #NavierStokes #AIAgents #DataProtection
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
@ChatGPT Work now builds a writing style from the workplace tools you already connected. Slack voice and customer-email voice are not the same. Incident notes and board updates are not the same. Native Writing Style blends the mailbox you connect. Keep a written skill for the channels that cannot afford that blend.
The 7 Sep official post named @Gmail, @GoogleDrive, @Slack, and @SharePoint. Favorite phrases. The specific sign-off. Capitalization quirks. Derrick Choi, Codex APAC lead, put the product claim in one line: no more skills needed to emulate his writing style. Native feature.
Setup is a Work popup — Messaging, Documents, Email — then “Use writing style,” or Settings → Personalization → Writing style. Once it is on, @OpenAI says it applies to every @ChatGPT Work message on web and mobile. Paid plans that have Work. Web setup. The promo draft on the same screen said the model learns from connected apps plus past ChatGPT context.
That is the operational lesson, not the headline. This is not a pasted sample you can put back in a drawer. It is a standing voice built from live mail, chat, and files. Slack and a customer email are not the same voice. An incident note and a board update are not the same voice. There is no advertised date filter. One reply already asked for one, because the imported style will include years of em dashes and “not that, not that, just this.”
Work already sits on those same connectors. Style now rides the same pipe.
What I would take into the room:
1. Treat Writing Style as an access grant, not a personalization toggle. Who connected Gmail, Slack, Drive, or SharePoint, and can a tenant admin turn this off.
2. Do not point it at the production mailbox first. If you test, use a dated, clean corpus. Screenshot before and after on one customer email and one internal note.
3. Ask OpenAI the numbered questions: what is retained after you disconnect the app; is the style profile per-user; is there a date window; does it blend Slack and legal email into one voice; can Enterprise disable it.
4. Keep a written style skill for regulator, customer, and incident channels. Native mimic is not a brand voice, and it is not a legal voice.
5. If the account can send, require a human on send even when the draft “sounds like you.”
What it is not: proof that custom skills are obsolete. Proof that the model inherited your judgment. A privacy review.
Source: @ChatGPT, 7 Sep 2026.
Chamath posted “The AI Singularity” on 3 Aug.
Five steps: humans build AGI, it gets good at AI research, it designs a smarter model, that model designs a smarter one, the cycle speeds up. He said the labs were already in that loop. Next 18 months wild. Recursive self-improvement from here. Marginal cost of all models to ~$0. Two million views. That is the speech.
Five weeks later OpenAI shipped GPT-6 Astra and graded the same claim. Critical in cybersecurity. High in bio/chem. Below High in AI self-improvement. Astra hit 78% on their Internal Research Debugging eval and stayed under the High line. Jensen put 100,000-plus Grace Blackwell GPUs on the training run and said AGI had arrived. Brockman called it the first OpenAI run over 100,000 GPUs. Huang also said 400,000 GPUs are coming online next. That is not a software loop with no denominator. That is a power, casting, and cluster loop. GE just wrote $11.75 billion for CPP for the same reason. SpaceX is standing up a blades-and-vanes foundry in Bastrop. Marginal token cost falling is not the same as cluster cost going to zero.
What the labs actually automated is research assistance and coding agents, not an unattended successor. Princeton’s shadow evals on unpublished NeurIPS work still got rejected by the original authors. Models help write the paper. They have not taken the research org off the critical path.
The operational lesson is the split Chamath did not write down. Self-improvement is still a Preparedness grade below High. Computer use plus Critical cyber is already in the product. Default Astra is not Daybreak. ExploitBench 100% was the ungated eval cut.
What I would take into the room:
1. Do not put “we are in the singularity” on the risk register. Put “frontier agent with tools, browser, and local files” on it, with an owner.
2. Leave Enterprise Astra off until computer-use has a written rule: identity, egress, which NPI it may see, human approve on anything that changes config or data.
3. If you need exploit-class work, that is Daybreak, off the SSO laptop. The consumer cut will refuse the interesting jobs and still needs a sandbox.
4. Budget the physical constraints. Tokens getting cheaper does not remove the GPU, power, and foundry queue.
5. Keep a fallback model. Same-day outages already happened across ChatGPT, Claude, and Grok this month.
What it is not:
AGI as a control attestation. Not RSI as a shipping feature. Not inference at $0 while the next training run is a 400,000-GPU plan.
Source: @chamath, 3 Aug 2026. OpenAI Path to Astra / system card, 1–3 Sep 2026. Jensen Huang, 6 Sep 2026.
#Astra #OpenAI #AIAgents #Cybersecurity
The AI Singularity
The argument goes like this:
1. Humans build an AGI.
2. The AGI becomes good at AI research.
3. It designs a smarter AI.
4. That smarter AI designs an even smarter AI.
5. The cycle repeats faster and faster.
Looking at the results and capabilities from the various labs over the past few weeks I would say we are firmly in this loop now.
The next 18months will be wild.
Recursive self improvement will dramatically increase capability very quickly from here.
Marginal costs of all models will go to ~$0.
GE Aerospace just wrote an $11.75 billion check for the part of the industrial stack SpaceX is trying to route around: precision castings.
GE Aerospace agreed today to buy Consolidated Precision Products from Warburg Pincus and Berkshire Partners for $11.75 billion. Close is aimed at 2H 2027. $7 billion cash, rest new debt. Culp says first-year accretive to adjusted EPS and free cash flow. Multiple is about 26x 2027 EBITDA without synergies, 18x with.
CPP is not a glamorous name. It pours complex superalloy, titanium, aluminum, magnesium, and steel castings for commercial engines, military aircraft, weapons, helicopters, and industrial gas turbines. About 6,600 people, 20-plus plants, Cleveland HQ. GE has bought from them for more than 15 years. Culp’s reason is not subtle: “Investing in mission-critical casting capacity is needed to support the strong simultaneous demand across commercial engines, aftermarket and defense.” He also said the deal is supposed to speed hotter airfoils.
That is the same commodity Elon Musk just pointed at. Late August he confirmed SpaceX is standing up a blades-and-vanes foundry in Bastrop so natural-gas turbines for AI data centers can come online up to 18 months faster. Hot-section airfoils run 3,000–3,600°F — hundreds of degrees above the melting point of the alloy. A handful of shops do this at scale. Their books run into 2030. GE Vernova has already told investors the step-up to 24 GW of annual output in 2028 depends on more castings and forgings arriving in 2027.
So you have two plays on the same chokepoint. GE Aerospace is buying the existing foundry network and wrapping FLIGHT DECK around it. SpaceX is trying to pour the parts itself. Neither move is about marketing. Both are about who controls the parts that sit in the hottest section of a turbine when every OEM is sold out.
If you own engine availability, aftermarket shop visits, or data-center power dates, this is the map: Howmet, Precision Castparts, CPP, Doncasters.
GE just took one of those names off the independent list. SpaceX is building a fifth.
JetBrains shipped the TeamCity RCE advisory in July. Cadence, their own hosted service, stayed vulnerable. Attackers used CVE-2026-63077 Aug 8–24 and walked a 2024 backup plus AWS IAM. Rotate anything that touched Cadence. Hunt from Aug 8.
JetBrains disclosed CVE-2026-63077 on July 27: unauthenticated RCE on TeamCity On-Premises, CVSS 9.8, agent polling path, deserialize and you own the server process.
Cadence, their hosted GPU/cloud service wired into PyCharm, ran TeamCity behind it. That box stayed unpatched. Attackers used the same CVE from August 8 to August 24.
What they reached, per JetBrains’ own updates through September 3:
• A full 2024 Cadence server backup
• AWS IAM users and secrets, including employee accounts that used the service
• S3 in JetBrains Cadence accounts
• email, project source synced from PyCharm, and credentials tied to current users
They took https://t.co/Tp4u6dJDmv offline. They still cannot say whether customer-owned buckets were touched. Treat every Cadence job and its output from that window as untrusted.
If you used Cadence or run TeamCity on-prem:
1. Confirm TeamCity is on 2025.11.7 or 2026.1.3. Internet-facing and still behind is not a plan.
2. Rotate every token, IAM key, registry login, and repo credential that ever touched Cadence or that 2024 backup.
3. Hunt AWS, artifact stores, and VCS from August 8 forward. Look for new IAM users, unexpected S3 gets, and build agents that should not exist.
4. Assume pipeline output from that period can lie. Rebuild what you cannot prove.
PaperCut is a second fire the same week (CVE-2026-81578 + CVE-2026-82078 against schools). Do not let it bury this one. CI is still where cloud keys live.
Source: JetBrains Cadence incident post, last updated September 3, 2026. CVE-2026-63077.
Anyone already force-rotating Cadence-connected AWS keys, or are we still treating “vendor hosted the CI” as someone else’s problem?
#TeamCity #SupplyChain #CloudIAM
Liquid Network (Blockstream’s Bitcoin sidechain) just lost about 4,000 of the 4,200 BTC in its federation wallet — roughly $320 million — through a normal SideSwap peg-out.
Keys were not stolen. Researchers say an Elements range-proof cache bug let someone mint unbacked L-BTC, burn it in a valid peg-out, and walk out with real BTC. The party holding the coins posted “we are whitehats” on-chain and is sitting on the funds until a patch is proven. The sidechain is paused.
That is the operational lesson, not the headline. A federation wallet, an 11-of-15 sign-off, and a peg-out path all treated minted tokens as legitimate because some nodes were running code that had already been fixed in the repo but not shipped in a tagged release.
If you run a wrapped-asset, sidechain, or multi-sig treasury that depends on “the software said this is valid,” you just watched 95% of a reserve leave through the front door.
What I would take into the room:
Confirm whether any treasury or settlement rail we touch uses Liquid / L-BTC or the same Elements stack; treat “white hat holding the coins” as an unproven claim until the BTC actually moves back; and ask who owns the patch-to-production clock on every federated or wrapped-asset system we rely on. The bug was known. The payout still cleared.
https://t.co/LNMKGrw7M2
Codex just shipped the thing every long agent session has been missing: it can keep notes across context windows and search earlier messages and tool output, including details that never made it into the notes.
Compaction used to crush a six-hour debug or a large refactor into a thinner and thinner summary until the model forgot why a path failed and what the last tool actually returned.
Flip Astra on, set features.context_management.experimental_mode = true in ~/.codex/config.toml, restart, and start a new task. It is still experimental, same-task only, and needs a ChatGPT Plus / Pro / Pro Lite sign-in, not an API key. Treat it as a better working memory for one job, not a new corporate memory store.
Source: Vaibhav Srivastav.
ICYMI: Experimental but Codex can now keep notes across context windows and search earlier messages and tool outputs, including details missing from its notes.
Useful for long debugging sessions and large refactors.
To try it:
> Update Codex and select Astra.
> Add to ~/.codex/config.toml:
[features.context_management]
experimental_mode = true
> Restart Codex and start a new task.
Requires ChatGPT Plus, Pro, or Pro Lite sign-in.
Would recommend giving it a shot - Enjoy! ;)
Tibo just posted a quiet Codex note that power users should actually read.
If you run Astra while logged in with a ChatGPT account, they shipped usage improvements aimed at the long tail. Not a quality bump. Not a plan change. A claim that the same work can pull up to 3–4x less usage off the subscription for the people who were burning it fastest.
That sentence only makes sense if you have lived the last few months of Codex metering. Median users were fine. The heavy Astra sessions — long threads, subagents, retries, background work — were the ones that emptied a window. OpenAI has said this pattern out loud before: they optimized to the average and missed the tail. This is another pass at that tail.
What it is not: a reset.
Replies already asking for one. He did not offer one. It is also not a promise that your next 20-minute Astra run will cost a quarter of last week’s. “Long tail” and “up to” are doing real work in that tweet.
What I would do Tuesday!:
stay on the ChatGPT login path they named, keep Astra on the parent thread and cheaper models on the workers, and screenshot usage before/after one comparable job. If the meter does not move, you have a vendor claim to send back. If it does, you just bought more real work inside the same plan.
Anyone already seeing the 3–4x on a real Astra session, or is this still only visible on the worst 1% of threads?
Source: Tibo (@thsottiaux), 6 Sep 2026
#Codex #Astra #OpenAI https://t.co/hrW5XtCRhZ
We've made some improvements that improve usage on the long tail for power users of Astra when logged in with your ChatGPT account.
No change in quality and a pure win that on the long tail can result in up to 3-4X less usage being drawn from the subscription.
Codex still ships with a timid default: only a handful of concurrent subagent threads (the post says 3; most public configs sit around 4–6). That is a product choice, not a model limit. OpenAI tested Astra as the orchestrator with up to 64 workers, and the knob is one line in config.toml: [agents] max_concurrent_threads_per_session = 64.
The useful pattern is Astra on the parent thread and Luna on the children — parallel investigation and implementation, not 64 copies of the same expensive model. It works. It also burns quota and makes the UI lag; one run in the clip was 19 minutes and 5% of a Pro 5x weekly limit. Raise the cap only when the work is actually independent, pin Luna for the cheap lanes, and treat 64 as a lab setting, not a standing policy.
Source: @daniel_mac8.
I bet you didn't know that by default Codex only runs max 3 subagents concurrently.
That's not ambitious enough for Astra.
OpenAI tested Astra with up to 64 subagents.
Copy this into your config.toml:
[agents]
max_concurrent_threads_per_session = 64
Astra orchestrating 64 Luna subagents in parallel is a sight to behold!