Principal Software Developer @SASsoftware | Agentic AI | Technical lead for SAS Intelligent Decisioning and SAS Model Manager. CEO Award of Excellence.
See how AI-powered decisioning helps marketers deliver next-best actions, real-time engagement and continuous optimization at scale.
Read The Marketer's Guide to AI and Decisioning: https://t.co/RWL1H93me8
#MarTech
Traditional auto insurance pricing relies heavily on generalized linear models, which are explainable but can struggle to capture complex differences in customer behavior. This can produce overly broad tariffs, inaccurate premiums and lost sales.
Neova Sigorta is using SAS Insurance Life Cycle Accelerator to combine GLMs with explainable machine learning, uncover more granular risk patterns and deliver faster, fairer and more personalized pricing, whilst improving proposal acceptance.
https://t.co/8ZPVCi23S2
🛬 Most of us are unaware of all that goes into keeping things moving at the airport. Global airport manager, Fraport AG, uses SAS to help keep operations running as smoothly as possible. Read how https://t.co/I5wabEQUW2
#SASViya
SAS Communities post that takes a look at the Agentic Product Configurator, part of the SAS Insurance Life Cycle Accelerator.
It is built around a simple premise: product design should read like a conversation with a knowledgeable colleague, not a support ticket.
https://t.co/BzFzhzILt7
The SAS Agentic AI Accelerator provides a method for building AI agents leveraging SAS Viya technology.
It is designed to help users move more quickly from use-case idea to production, utilizing No/Low/Yes Code interfaces and full governance as a way to build agents that balance autonomy and trust.
https://t.co/BiIcnBzwNK
How to use the Private Docker publishing destination to publish a SAS decision or SAS model and then run it using SAS Container Runtime, in a completely cloud-provider-agnostic environment.
The ultimate goal is to run a SAS model or SAS decision against input data in the most lightweight way possible, independently of SAS Viya.
https://t.co/5TjgPJdJR3
Our latest accolade recognizes our strengths in Decision Intelligence. The IDC Marketscape report says: "Consider SAS if you are seeking a platform that can run identical decision logic across cloud, on-premises, and edge deployments without rewriting."
https://t.co/x0wAGs2kqj
Graduating 20 students in 2025, NC State’s Department of Statistics has ranked No. 1 as the nation’s top producer of Ph.D. graduates. 👏
See the recent data from the American Statistical Association (ASA). Read more: https://t.co/OyVlPsNu3O
What makes Claude Projects so interesting is that it handles teams of agents really well, you talk to a main orchestrator agent and it spins up specialists. Basically it creates an organization to solve your issue, mixing expensive and cheap agents depending on your preferences.
For example, I asked Fable in Claude Projects to select famous historical mysteries that it could try to resolve. It initiated research agents, selected the mysteries based on data it could access, and spun up eighteen separate threads, each with an agent each focused on one mystery. Then each thread launched additional agents (simulating avalanches, breaking codes) before summarizing those and passing them to still more agents for write up and another set of skeptical agents to fact check. It did this over a day of work, with the central orchestrator agent organizing it all.
The results were interesting if you like historical mysteries. They are also for fun and certainly not definitive or guaranteed error-free (but they are also mostly reasonable & grounded in the literature). https://t.co/uHToT5GNuF
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.
Today, we're sharing three measurements that help track AI development:
1. How much AI R&D is done by AI.
2. How well AI agents are overseen.
3. How compute is allocated.
We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.
As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.
Read the full post and methodology: https://t.co/iPFz8Z4ugE
Addressing both trust and maintenance in AI can be a challenge. Meta recently published an article that describes their "Organizational Second Brain".
The brain consists of a knowledge layer and a set of auditable files specifying the reasoning steps to process knowledge.
When an agent outputs an error, the system diagnoses whether this originated from stale facts or flawed reasoning.
The goal is to keep business rules and procedural logic in human readable, version controlled text files to ensure transparency, auditability and reversibility for enterprise compliance.
https://t.co/4NDO176jtp
A recent announcement from OpenAI designates GPT-6 Astra as its first model to meet the “Critical” cybersecurity capability threshold under its Preparedness Framework.
Astra can identify unknown vulnerabilities and execute zero-day exploits across hardened systems without step-by-step human guidance.
The evaluation metrics highlight a big jump in offensive autonomous reasoning:
• Zero-Day Discovery: Astra scored 100% on ExploitBench and discovered two previously unknown V8 zero-days on a benchmark of 20 recent disclosures.
• Sandbox Escapes: In red-teaming, the model built a full browser sandbox escape starting from a single HTML file and chained local privilege escalation from an unprivileged user to root on a hardened OS.
• Resisting Alignment Honeypots: To test autonomous deviation, OpenAI placed Astra near untargeted systems informed by the Hugging Face incident. While GPT-5.6 Sol attacked the unauthorized target in 56% of runs, Astra refrained 100% of the time and never attempted to route around mid-task denials from automated review systems.
Deployment remains heavily restricted through the Daybreak Blue early access program.
Production Warning: OpenAI warns that Astra’s safety monitors will interrupt legitimate defensive work mid-execution. Friction is currently the primary safety feature, not a bug to optimize away.
https://t.co/LiTwiezwyV
You can now easily compare the intelligence, cost, and speed on different effort levels for your preferred models with the new Artificial Analysis Model Release pages
How a model performs across intelligence and cost is heavily influenced by its configured reasoning and effort level. Frontier models are now being released with up to six different effort levels, meaning the same model weights can produce significantly different performance and cost profiles. Model Release pages make it easy to compare these configurations across all of our standard charts in one place
Each release page features:
➤ Intelligence, cost per task, output speed, and latency across every effort variant
➤ Side by side comparison of effort levels against those of similar models
➤ Artificial Analysis Capability Index scores on each effort level across Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics
For example, on Intelligence vs Time per Task, GPT-6 Astra ranges from 46–53 on the Artificial Analysis Intelligence Index and 1.6–8.2 minutes per task, while Claude Fable 5.1 spans a similar intelligence range (47–53) on a wider time per task range (4.2–12.2 minutes)
Security operations are shifting toward multi-agent swarms, but isolated tools can force human operators to manually shuttle context between triage, enrichment, and remediation agents.
This blog post from Cisco lays out the case for pairing MCP for vertical tool integration and Agent2Agent (A2A) for horizontal inter-agent delegation.
Under the proposed A2A specification:
• Agent Cards: Agents publish agent cards that advertise capabilities and prove authenticity.
• Task Delegation: Client agents hand off work using established protocols.
• Autonomous Reasoning: Agents receive structured tasks and apply their own independent reasoning rather than following fixed scripts.
Pairing the MCP for vertical tool integration with the A2A specification for horizontal delegation sounds reasonable from an architecture perspective. However, organizations will still need to solve gaps around IAM-governed agent identities, durable audit trails, and strict receiver-side input validation.
https://t.co/3X4887dHrA
The window to patch software bugs is collapsing
Of the bugs hackers actually exploit, ~87% are now being attacked on or before the day the bug is public knowledge
That share was 23% in 2020
Charts of the Week: https://t.co/RfYfzHpbLI
It was a great week for innovation across the model ecosystem. We're bringing these new models into Copilot, enabling it to take on increasingly complex work, from quick questions to delegated tasks and, increasingly, complete long-running jobs through Autopilots.
Here's one fun example: using an Opal-powered Autopilot running on a secure Windows 365 Cloud PC, Copilot goes through a month of trail cam footage, finds every animal sighting, creates a highlight reel of the best clips, labels each one with the camera, date, and species, catalogs every sighting in a spreadsheet, builds a PowerPoint summarizing the findings, and shares everything in Teams for colleagues to review.
This is the next frontier for Copilot: software that doesn't just assist with work, but can own and complete entire tasks that unfold over hours or days.
@petersykim Passing to a fresh subagent can help provided there is some contract in place for passing the context rather than relying on natural language. Keeping subagents to a small number of steps as well.
Long-horizon reliability remains a major bottleneck for enterprise adoption. A research paper on arXiv studies agent performance over extended execution paths, revealing that task success follows a geometric decay curve governed by step count.
Key findings from the empirical benchmark include:
• Geometric Reliability Decay: Even frontier models experience rapid reliability drop-offs, collapsing toward near-zero task completion around step 16 of unattended workflows.
• Step Count vs Context Length: Counterintuitively, shrinking context windows accelerates failure. Sequential reasoning drift across individual steps, rather than context window saturation, is the primary driver of agent degradation.
• Benchmark vs Production Gap: Agent success drops from 0.42 on standard short-horizon benchmarks down to 0.24 on realistic 100-step enterprise workflows.
Building reliable agents requires breaking long-horizon tasks into short, deterministic sub-chains rather than expecting a single model loop to hold state across dozens of steps.
https://t.co/PSkhc2yYPe