I am evaluating the recently released grok 4.5 LLM via its CLI (the Agent).
First task:
Test if it follows guidelines in my RDF-based harness at session initialization time.
Result:
Passed!
See comments for link to HTML doc it generated demonstrating this test.
/cc @elonmusk@milichab
Apparently the answer is "no", judging by the attack announced after I posted this.
Not a great look for npm after touting its ingestion-time scanning.
Santa definitely has ADHD.
He has all year to do one thing and doesn’t start until the night it’s due but gets it done anyway in an 8-hour hyperfocus frenzy powered by god like amounts of sugar, carbs and adrenaline
The new @code release supports BYOK without a GitHub sign-in, making it easier to use your own language models or run workflows using local models like Ollama.
Check out what’s new:
🌐 Integrated Browser now includes device emulation for faster responsive testing, plus one-click screenshots to attach browser context directly to chat
🛠️ Custom Endpoint is now in Stable, so you can connect enterprise or self-hosted AI endpoints directly in VS Code
I admit though, I've already moved past the itsec, embracing a new field I could perhaps describe as "Knowledge Security"...
I'm dying to write more about it - an endeavour I've been working on for the past 3+ years - but it's still not ready... Hope to write more soon! 😃
@PatrickMoorhead We could start by rebooting indestructible Blackberry & Nokia designs that were once manufactured in North America and Europe. There is even demand for simulacra, e.g. monthly batches from this German shop sell out within minutes, https://t.co/paFlCVZ8Iz
I think we've been chasing the wrong definition of AGI.
It's not about giving better answers.
It's about whether the AI prepares before it answers.
That's why OpenHuman's Super Context caught my attention.
Most AI tools start every new chat from zero. You paste docs, explain the project, add context again.
With Super Context, OpenHuman gathers the relevant context from your machine locally before replying, so your very first message feels like you've already been talking for 10 turns.
That's a much more interesting direction for AI than simply making bigger models.
Most non-invasive brain-to-text decoding has sat at single-digit accuracy — basically unusable outside a lab. Meta's FAIR lab just moved that number, with no implant and no surgery.
They released Brain2Qwerty v2 — an end-to-end pipeline that decodes typed sentences from real-time magnetoencephalography (MEG) recordings. It pairs a convolutional encoder, a transformer, and a character-level language model, trained on ~22,000 sentences from 9 participants recorded for 10 hours each.
Here's what's actually interesting:
→ 61% average word accuracy (39% WER) — up from 8% for prior non-invasive methods
→ Best participant: 78% word accuracy, with over half of sentences decoded at one word error or less
→ End-to-end deep learning replaces hand-crafted event-detection pipelines; fine-tuned LLMs add the semantic context that cleans up noisy character predictions
→ AI agents iteratively refined the decoding pipeline — but engineers still selected the final training configs manually
→ Accuracy improves log-linearly with data, suggesting the gap with surgical implants could narrow through scaling alone
Full analysis: https://t.co/6CH8A2pNjw
Paper: https://t.co/9wvv2zWKGM
Repo: https://t.co/XmSYGyPLZF
Technical Details: https://t.co/uhJt3I0LKg
@AIatMeta@Meta_Engineers
Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
There's something missing from human and AI collaboration: a way to exchange reviews and history. Changedown puts that context in the Markdown character stream, giving humans clearer trust and agency, and ensuring AI agents encounter state in the world. https://t.co/ww0SXuWHNT
Almost forget, thank you @Google for your "active" and "tireless" responses. Your informative template emails have successfully caused 160,000 users and over 800 five-star reviews to vanish completely.
We know you have a LOT of organizational requests 🙏 and we ARE working 👷♀️ on them (🗂️ folders anyone?) But, in the meantime, here's a little something 🎁 to whet your appetite 🍽️... you 🫵can now choose the emoji for your notebook 📒! 😍🥳😎
Go on, express yourself! 🎨🎭🌻
Tired of Wikipedia’s "notability" police? 🛑 Check out https://t.co/cn15bPHl0t by Ciro Santilli. It’s Wikipedia + GitHub for all.
✅ Own your data (Git/Plaintext) ✅ Native LaTeX for deep science ✅ No censorship or forced consensus
Freedom for your knowledge! 🚀 #Tech