@RidgelineCyber@Bluewall Are you sure that non-managed/compliant devices would be blocked? AFAIK, token protection requires just the Entra ID registered status
I've been researching the Microsoft cloud for almost 7 years now. A few months ago that research resulted in the most impactful vulnerability I will probably ever find: a token validation flaw allowing me to get Global Admin in any Entra ID tenant. Blog: https://t.co/jD6EaGtsn3
2. Monochromatic icons for the categories in the stats sections are harder to distinguish (same as all Google apps look the same). I’d like to see them colored again, as in other screens of the app
1. The iOS widget is now broken. It shows an empty black rectangle. The app opens when tapped, but I miss the possibility to check my balance at a glance 😓
Apple gets it. Robots are going to be everywhere, but they won’t look like robots. Check out their new paper ELEGNT.
I believe this is the future of everyday objects: helpful and human.
why did R1's RL suddenly start working, when previous attempts to do similar things failed?
theory: we've basically spent the last few years running a massive acausally distributed chain of thought data annotation program on the pretraining dataset.
deepseek's approach with R1 is a pretty obvious method. they are far from the first lab to try "slap a verifier on it and roll out CoTs."
but it didn't used to work that well. all of a sudden, though, it did start working. and reproductions of R1, even using slightly different methods, are just working too--it's not some super-finicky method that deepseek lucked out finding. all of a sudden, the basic, obvious techniques are... just working, much better than they used to.
in the last couple of years, chains of thought have been posted all over the internet (LLM outputs leaking into pretraining like this is usually called "pretraining contamination"). and not just CoTs--outputs posted on the internet are usually accompanied by linguistic markers of whether they're correct or not ("holy shit it's right", "LOL wrong"). this isn't just true for easily verifiable problems like math, but also fuzzy ones like writing.
those CoTs in the V3 training set gave GRPO enough of a starting point to start converging, and furthermore, to generalize from verifiable domains to the non-verifiable ones using the bridge established by the pretraining data contamination.
and now, R1's visible chains of thought are going to lead to *another* massive enrichment of human-labeled reasoning on the internet, but on a far larger scale... the next round of base models post-R1 will be *even better* bases for reasoning models.
Worth repeating:
Do not confuse retrieval with reasoning.
Do not confuse rote learning with understanding.
Do not confuse accumulated knowledge with intelligence.
Do you know what color chucknorris is?
Try it right now as a background color in HTML. On most browsers, you should see a crimson shade of red.
What is going on?
It's actually a fascinating artifact from the Netscape 4 era👇