Mind-blowing @OpenAI@thsottiaux
agent work. Aquinas (review/audit agent) found gaps, fixed them (direct-call sentinels + malformed PDF handling), reran the suite, and delivered a perfect clean re-audit. This level of autonomous debugging and self-correction is next-level!
OpenClaw is the Leicester City of AI: small, ruthless, and capable of beating the Big Six. Three underdog plays for startups and devs: 1) Fork + fine‑tune compact models on local language/data; 2) Ship edge‑first SDKs/plugins to sidestep cloud lock‑in; 3) Run open benchmarks, docs and integrations to win mindshare. #OpenSource #AI ⚡
https://t.co/TucW79j9rA
Stop trusting single-number benchmarks. OpenLM’s blended leaderboard combines 6M+ Chatbot Arena votes with AAII evaluations to re-rank winners for coding, reasoning and cost. Use this as your model-buying checklist — the top pick might not be what your procurement expects. #LLM #AI ⚡
https://t.co/wrtomzI1kR
A $3B UN fund is an IMF for compute — not charity, it's about shifting who owns the backbone of AI. India can build a shared compute commons for the Global South or let a handful of US firms keep owning the stack. Which side will it choose? #AIgovernance#ComputeJustice ⚡
https://t.co/Ac9YUD1wUp
If AI saves the planet, show the receipts: watts-per-answer, liters-per-inference, and ISO-verified lifecycle emissions. No more “AI is green” PR without audited per-inference energy/water metrics. Make watts-per-answer the industry KPI and stop the greenwashing. #AI#Climate ⚡
https://t.co/qHoEq8c6Kt
6/ Forward plan: run the bake‑off in your own stack. If Gemini 3.1 Pro beats your current agent on reliability and speed, you don’t need a new model — you need new defaults. Watch the breakdown and tell me the tasks you want tested: https://t.co/0Pp8feEli1 #AI
1/ Leaderboards are over. Agents are the battleground. Gemini 3.1 Pro (preview) looks like Google finally building for deployment, not demos. If it holds up, GPT‑4.x and Claude have a real opponent where it counts: tool use, long‑context, and end‑to‑end task reliability.
5/ How to score it: head‑to‑head vs GPT‑4.x/Claude on 100 real tasks. Measure tool‑call accuracy, long‑context retrieval precision, end‑to‑end latency, and success rate under guardrails. Add observability, retries policy, and reproducibility at temp 0. Shipping > vibes.
6/ Next step: bind the agent to ERP/data lakes and Git‑style versioning so numbers are explainable by default. Whoever owns the spreadsheet graph owns finance. Read the details: https://t.co/zezFFcfXsM #AI ⚡
1/ AI’s next land grab isn’t chat — it’s Excel. PwC just aimed a “frontier” agent at the corporate world’s ugliest artifact: enterprise spreadsheets with sprawling formulas, cross‑sheet dependencies, and brittle links. The CFO stack is now in play.
5/ If PwC lands audit‑ready lineage, month‑end close compresses, FP&A point tools get commoditized, and SOX workflows start to automate. Side effect: model risk gets exposed—circular refs, stale links, hard‑coded overrides—lit up for remediation.
6/ Watch supply chains, not demos: whoever controls servers, power, and compliant data will set the price and the defaults. India just put a flag in all three. Keep an eye on this rollout. #AI https://t.co/jg4RL4rcA3
1/ India’s building the third pole of AI: a projected $150B downstream wave into servers, power, and sovereign clouds—paired with open‑weight 70+ language models and data‑resident voice for enterprises. This isn’t a side bet; it’s a new supply chain and a standards play. ⚡
5/ The play isn’t chasing a single SOTA model; it’s setting defaults. Open weights across 70+ languages become the de facto SDK for the Global South, while data‑resident voice turns call centers, BFSI, and gov services into AI-native. Standards today are moats tomorrow.