โ๏ธ Data Scientist & ML Engineer | Comput. Physics, Remote Sensing.
๐ Cofounder @ParseAI | AI title research for real estate & energy transactions.
Software horror: litellm PyPI supply chain attack.
Simple `pip install litellm` was enough to exfiltrate SSH keys, AWS/GCP/Azure creds, Kubernetes configs, git credentials, env vars (all your API keys), shell history, crypto wallets, SSL private keys, CI/CD secrets, database passwords.
LiteLLM itself has 97 million downloads per month which is already terrible, but much worse, the contagion spreads to any project that depends on litellm. For example, if you did `pip install dspy` (which depended on litellm>=1.64.0), you'd also be pwnd. Same for any other large project that depended on litellm.
Afaict the poisoned version was up for only less than ~1 hour. The attack had a bug which led to its discovery - Callum McMahon was using an MCP plugin inside Cursor that pulled in litellm as a transitive dependency. When litellm 1.82.8 installed, their machine ran out of RAM and crashed. So if the attacker didn't vibe code this attack it could have been undetected for many days or weeks.
Supply chain attacks like this are basically the scariest thing imaginable in modern software. Every time you install any depedency you could be pulling in a poisoned package anywhere deep inside its entire depedency tree. This is especially risky with large projects that might have lots and lots of dependencies. The credentials that do get stolen in each attack can then be used to take over more accounts and compromise more packages.
Classical software engineering would have you believe that dependencies are good (we're building pyramids from bricks), but imo this has to be re-evaluated, and it's why I've been so growingly averse to them, preferring to use LLMs to "yoink" functionality when it's simple enough and possible.
๐ข๐ป๐ฒ ๐บ๐ฒ๐บ๐ผ๐ฟ๐ ๐ฐ๐ฎ๐ปโ๐ ๐ฟ๐๐น๐ฒ ๐๐ต๐ฒ๐บ ๐ฎ๐น๐น.
We present ๐๐ผ๐๐ฒ๐ฅ, a new ๐ต๐๐ฏ๐ฟ๐ถ๐ฑ ๐บ๐ฒ๐บ๐ผ๐ฟ๐ architecture for long-context geometric reconstruction.
LoGeR enables stable reconstruction over up to ๐ญ๐ฌ๐ธ ๐ณ๐ฟ๐ฎ๐บ๐ฒ๐ / ๐ธ๐ถ๐น๐ผ๐บ๐ฒ๐๐ฒ๐ฟ ๐๐ฐ๐ฎ๐น๐ฒ, with ๐น๐ถ๐ป๐ฒ๐ฎ๐ฟ-๐๐ถ๐บ๐ฒ ๐๐ฐ๐ฎ๐น๐ถ๐ป๐ด in sequence length, ๐ณ๐๐น๐น๐ ๐ณ๐ฒ๐ฒ๐ฑ๐ณ๐ผ๐ฟ๐๐ฎ๐ฟ๐ฑ inference, and ๐ป๐ผ ๐ฝ๐ผ๐๐-๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป.
Yet it matches or surpasses strong optimization-based pipelines. (1/5)
@GoogleDeepMind@Berkeley_AI
Qwen3.5 dense (smol ๐ค) models just dropped
- natively multimodal
- 0.8B ยท 2B ยท 4B ยท 9B (+ base variants)
- 262K context extensible to 1M
- built-in thinking
fine-tune them with TRL out of the box โ SFT, GRPO, DPO and more!
Weโre excited to release TorchLean which is the first fully verified neural network framework in Lean. The Lean community has largely focused on pure mathematics. TorchLean expands this frontier toward verified neural network software and scientific computing. With the recent release of CSlib, we see this as another step toward a fully verified ML stack.
We support features:
1. Executable IEEE-754 floating-point semantics (and extensible alternative FP models) verified tensor abstractions with precise shape/indexing semantics
2. Formally verified autograd system for differentiation of NN programs Proof-checked certification / verification algorithms like CROWN (robustness, bounds, etc.)
3. PyTorch-inspired modeling API with eager-style development + export/lowering to a shared IR for execution and verification
Project page: https://t.co/YHpqhRbMQe
Paper: [2602.22631] TorchLean: Formalizing Neural Networks in Lean
Work done @Robertljg, Jennifer Cruden, Xiangru Zhong, @huan_zhang12 and @AnimaAnandkumar.
#MachineLearning #ScientificComputing #Lean
๐จ BREAKING: Alibaba just handed the AI agent community a production-grade sandbox for free.
OpenSandbox is a full-stack platform for running untrusted agent code safely:
โ Unified APIs across multi-language SDKs
โ Docker and Kubernetes runtimes purpose-built for agents
โ Browser automation, VS Code desktop, and network isolation included
โ Designed for coding agents, GUI agents, evaluation, and beyond
Not a side project. Built by Alibaba. Open source.
1.5k stars (+1,100 this week). The secure agent infra you didn't have to build yourself.
Use local models on remote devices you controlโas if they were local.
- Introducing LM Link from @LMStudio.
- Encrypted, identity-based access to your own LLM hardware.
- No public endpoints. No API key sprawl.
https://t.co/SGJC0JhlbU
New Google paper challenges how we measure LLM reasoning.
Token count is a poor proxy for actual reasoning quality.
There might be a better way to measure this.
This work introduces "deep-thinking tokens," a metric that identifies tokens where internal model predictions shift significantly across deeper layers before stabilizing.
These tokens capture "genuine reasoning" effort rather than verbose output.
Instead of measuring how much a model writes, measure how hard it's actually thinking at each step. Deep-thinking tokens are identified by tracking prediction instability across transformer layers during inference.
The ratio of deep-thinking tokens correlates more reliably with accuracy than token count or confidence metrics across mathematical and scientific benchmarks (AIME 24/25, HMMT 25, GPQA-diamond), tested on DeepSeek-R1, Qwen3, and GPT-OSS.
They also introduce Think@n, a test-time compute strategy that prioritizes samples with high deep-thinking ratios while early-rejecting low-quality partial outputs, reducing cost without sacrificing performance.
Why does it matter?
As inference-time scaling becomes a primary lever for improving model performance, we need better signals than token length to understand when a model is actually reasoning versus just rambling.
Paper: https://t.co/Yj0bPdiLni
Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
Diagonal panning video from a very detailed shaded map of Genova, Italy. The full map is 36000 x 16700 pixels with 0.5 m resolution and is available at https://t.co/hs0XhI8Ann
Data source: https://t.co/ewV3IwUWni [Comune di Genova].
#maps#shadedmaps#cartography#shadedrelief
Do we really need billion-parameter transformers for time series forecasting?
What if we did it in under 3M parameters?
Introducing Reverso: time series foundation models as small as 200K paramters significantly pushing the Pareto frontier in zero-shot time series forecasting!
100,000+ models trained with Unsloth have now been open-sourced on Hugging Face!
Popular fine-tuned LLMs you can run locally:
1. TeichAI - GLM-4.7-Flash distilled from Claude 4.5 Opus (high)
2. Zed - Qwen Coder 7B fine-tuned for stronger coding
3. DavidAU - Llama-3.3-8B distilled from Claude 4.5 Opus (high)
4. huihui - gpt-oss made โabliberatedโ
Links to models:
1. TeichAI: https://t.co/OwJ3xfZ9BC
2. Zed: https://t.co/ssfLx1F67F
3. DavidAU: https://t.co/6gN8RmnDio
4. huihui: https://t.co/vg4AcsvaPC
See all the 100K latest models fine-tuned with Unsloth here: https://t.co/3abm6sql2h
See How the World Has Changed Over Time ๐๐ฐ๏ธ
Ever wondered how cities, infrastructure, or landscapes evolved over the years?
World Imagery Wayback makes it easy.
This free service archives satellite images since 2014, allowing you to track territorial changes over time by simply selecting a date on the timeline.
Itโs especially useful for GEOINT specialists and anyone doing retrospective analysis:
โ Track urban growth and industrial expansion ๐๏ธ
โ See when new roads, bridges, or infrastructure appeared ๐ฃ๏ธ
โ Analyze changes in farmland, ports, or airfields ๐พ
โ Document long-term impacts of natural disasters or major construction ๐๏ธ
A simple, visual way to understand how our world transforms, year by year.
๐ Website link: https://t.co/FDPvHcoCzm
__________
P.S. โป๏ธ Repost if you found this helpful.
If you liked this post and would like to learn more methods and techniques to discover information about people, check out my OSINT Mastery course https://t.co/o7UQQR6SNU