🚀 We just released our dataset on @huggingface!
QueST-PartNetMobility-SAPIEN built for the tracking problem everyone avoids: what happens after frame 60?
That's the difference between "tracking" and "long-horizon tracking."
📈 2,500+ downloads in 48hr
❤️https://t.co/G2PsBUX6I2
Seems like the matching algorithm never allows me to get in 😭 ......
GSoc, MATS and now SPAR as well
Fortunately this time I was in touch with the mentors so would be working on a different project
Agents can reason. But can they prove they have enough evidence to act?
We built EviPlan as an evidence firewall for autonomous agents.
Evidence → Freshness + Provenance + Consistency + Policy → ALLOW / RECHECK / BLOCK
Now live as a Telegraph Miner 🚀
@Telegraphprotoc
Our hackathon is live, and over 1,300 builders have already registered. It's only day one, and use cases are already piling in from the community, built by different teams, with many more still to come. Here are a select few:
Degenlens by @drained99 : built as a Telegraph Miner across three intents (ONCHAIN_TX_LOOKUP, WALLET_BALANCE_CHECK, FRAUD_DETECTION), turning raw on-chain activity into verifiable gambling intelligence: deposits, withdrawals, wallet exposure, and fraud patterns like wash trading and sybil activity, each answer scored, and ranked by WASM evaluations from Track 2 and finalized with a verifiable signal hash, so casinos, analysts, and agents can all trust the same number.
EviPlan by @MayankAnan81495 : registering as a Telegraph Miner under the network's risk-and-trust intents, so an agent can check temporal freshness, source provenance, and state consistency against a ranked, verified signal before it's allowed to execute an action like an on-chain trade, a DevOps command, or an IoT trigger.
Anchor by @I_am_SamY01 : registering as a Telegraph Miner to serve Aave v3 wallet-risk signals, health factor, liquidation distance, risk classification, and data freshness, so any agent on the network can query a verified, ranked version of that signal instead of trusting a single source.
Three separate builders, three very different problems, and this is just the start.
Season I is open across three tracks: Miner, Script Author (Evaluations), and Application, with $15,000 in prizes across the series.
Join our community now: https://t.co/l8PNzr11Jf
Register for track 1, 2 & 3: https://t.co/2LVFcfwtdn
Applied for the autumn cohort by chance ....honestly the application process takes a lot of time though it's was super helpful for gaining experience in ai safety research ...went till stage 3.5 but ultimately rejected
Moved to stage 2 for this cohort ....
🚨 MATS Winter 2027 applications are now open.
Fully-funded, 12-week fellowship for aspiring & established AI alignment, interpretability, security, governance researchers & field-builders
📍 Berkeley/London
📅 Jan 19–Apr 10
💰 $6.4k/mo + $8k - 16k/mo compute
Apply by Sep 6 ↓
Our hackathon is live, and over 1,300 builders have already registered. It's only day one, and use cases are already piling in from the community, built by different teams, with many more still to come. Here are a select few:
Degenlens by @drained99 : built as a Telegraph Miner across three intents (ONCHAIN_TX_LOOKUP, WALLET_BALANCE_CHECK, FRAUD_DETECTION), turning raw on-chain activity into verifiable gambling intelligence: deposits, withdrawals, wallet exposure, and fraud patterns like wash trading and sybil activity, each answer scored, and ranked by WASM evaluations from Track 2 and finalized with a verifiable signal hash, so casinos, analysts, and agents can all trust the same number.
EviPlan by @MayankAnan81495 : registering as a Telegraph Miner under the network's risk-and-trust intents, so an agent can check temporal freshness, source provenance, and state consistency against a ranked, verified signal before it's allowed to execute an action like an on-chain trade, a DevOps command, or an IoT trigger.
Anchor by @I_am_SamY01 : registering as a Telegraph Miner to serve Aave v3 wallet-risk signals, health factor, liquidation distance, risk classification, and data freshness, so any agent on the network can query a verified, ranked version of that signal instead of trusting a single source.
Three separate builders, three very different problems, and this is just the start.
Season I is open across three tracks: Miner, Script Author (Evaluations), and Application, with $15,000 in prizes across the series.
Join our community now: https://t.co/l8PNzr11Jf
Register for track 1, 2 & 3: https://t.co/2LVFcfwtdn
crazy to see the adoption of frontier models in robotics ..
1) when model planning leads to action decisions are the evidences grounded ... this is where i think physical grounded models will come in
2) What's going to be the memory layer for it as LLMs are bad at forecasting WM
Most unsettling line in Anthropic’s Risk Report isn’t about killer AI.
It’s "Their own evaluations for measuring AI-driven R&D acceleration are starting to saturate" – Meaning the tests can no longer reliably capture capability increases.
At the same time, Anthropic says Claude now writes a large majority of the code merged into its production codebases & its internal AI-assisted R&D is already significantly faster.
They’ve also raised high-stakes misalignment risk from “very low” → “low.”
The combination is fascinating.
We’re getting better at building models faster than we’re getting better at measuring what those models can actually do.
Got it now where I was going wrong.... it's the foundation engineering of the SOTA you are building upon is what matters the most for getting the results
Showing that something fails is not enough ..
Our Paper is now available on Arxiv for reading :
QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon Tracking
Mayank Anand, Mohammad Saqlain, Kyan Mahajan, Priya Shukla, Gora Chand Nandi, Andrew Melnik
https://t.co/Gk2gstQ6a9