Google trained an AI to predict your neighbourhood's income by counting the coffee shops, bus stops, and high-rises on a map. Nobody told it what income was.
The model is called S2Vec, and it was published by Google Research as part of their Earth AI initiative. It takes the built environment (every building, road, park, and business in an area) and converts it into a layered image. Three coffee shops and one park in a grid cell become pixel values. The AI then reads that image the same way a computer vision model reads a photograph.
The training method is the part that matters. S2Vec uses masked autoencoding: you show the model a patch of a city with chunks missing, and it learns to fill in the gaps. Show it a cluster of high-rise apartments next to a subway station, mask out a section, and it predicts a grocery store belongs there.
Do that millions of times across the globe and the model learns the deep spatial grammar of how cities organise themselves. No human ever labels a region as "financial district" or "suburban residential." The model figures out those groupings on its own from the geometry of what's built where.
The output is an embedding, a string of numbers that acts as a mathematical fingerprint for any location on Earth. Feed those embeddings into a prediction task and S2Vec can estimate population density, median income, and carbon emissions for regions it has never seen before.
On zero-shot geographic extrapolation (predicting for regions entirely absent from training data) S2Vec was typically the best-performing individual model.
It matched or beat satellite imagery baselines like RS-MaMMUT and outperformed GEOCLIP on socioeconomic prediction. The best results came from combining S2Vec with satellite image embeddings. Built environment data alone couldn't capture vegetation, terrain, or transportation patterns well enough for environmental tasks like tree cover and elevation. But fused together, the two modalities outperformed everything else.
The standard approach to geospatial ML has been hand-crafting indicators for every new problem. Predicting air quality meant building a bespoke feature set. Estimating housing prices meant building another one. S2Vec replaces that with a single general-purpose representation that transfers across tasks.
The training data is map features, not satellite pixels.
That distinction is pretty important to understand. It means: map data updates faster, costs less to process, and covers built infrastructure at a resolution satellite imagery can't always match.
A satellite sees rooftops. S2Vec knows there are three cafes, a pharmacy, and a bus stop underneath them.
Google's broader Earth AI pipeline now has three foundation models working in parallel.
1. PDFM for population dynamics.
2. RS-MaMMUT for satellite imagery.
3. S2Vec for the built environment.
Stack them and you get a system that can read a neighbourhood the way a local understands it.
More info on it here: https://t.co/vVJlLlfhc7
Project SmallField & Next 10B Molecule Geometries
A few weeks ago I put USearchMolecules into 3D, and said the toolkits behind it would follow in a separate wave. This is the first of them - project SmallField, a GPU-accelerated 3D conformer generation engine for small molecules in Chemistry and Biology - and I'm particularly conflicted about it. The reason is, of course - AI, LLMs, and agentic coding.
I started this project last year. First of all, I wrote almost none of the code myself, working more like a Technical Manager - mostly providing guidance. Second, making things worse, it reimplements two decades of work from RDKit - a BSD-3 project I deeply respect - and uses RDKit itself as the oracle for every correctness claim I make. A reimplementation is only ever as good as what it reimplements. Third, it doesn't invent anything - focusing on the academia/industry-standard ETKDGv3 and MMFF94 algorithms, just porting them to GPUs carefully: mixed precision where the geometry survives it, layouts that fit in SRAM, lower register pressure, and enough pipelining to keep the CPU and the GPU both busy. And if it wasn't enough, there is of course a team at NVIDIA working on something similar - nvMolKit. I tried to reformulate those ideas into something more like a Pull Request, but it didn't feel natural - the result was less a patch than a parallel implementation. What it's like to ship a codebase you didn't type is its own post, and I'll write it once I've stopped arguing with myself about it.
Why the hell did I release it then?
I wanted to scale-up my Unum infrastructure work in the CompBio & CompChem domains, quite literally adding extra dimensions to USearchMolecules - expanding from 1D fingerprints to 3D molecule shapes. Sadly, I raise money about as well as a brick swims, so I can't yet hire a team of people to help me build. All I have is reliable partners like Claude and @nebiusai 🫶. So over the course of the last 9 months, through multiple failed attempts, I got to a pipeline whose MMFF94 energies match RDKit's on identical coordinates, at roughly 10x the throughput.
On an NVIDIA DGX-H100, when configured properly, USearch v2 sustains ~100K queries per second across 2 CPU sockets. I needed the 8x H100 GPUs to drive relaxations at the same or higher throughput to fully saturate hardware. At this point for Enamine REAL-class molecules and 3 conformers per molecule - generation runs at ~250K conformers/s on that 8x H100 node. 3 conformers per molecule was enough diversity for GDB13, but for REAL, the next run will widen the scope, resulting in 100B shapes.
The first 10B shapes are already available on Hugging Face, AWS OpenData, and Nebius. That's ~460 billion atomic coordinates, within a factor of three of the entire AlphaFold database - a yardstick, not a claim, since a folded protein is by far the harder object. Not to say that I have something cool coming up for proteins 😉
The next waves will (1.) anchor reference-quality properties on a small set of well-characterized molecules and interpolate in the chemical space, (2.) probably replace ETKDGv3 and MMFF94 with something better, and (3.) probably still avoid neural nets for the geometries themselves. Always happy to compare notes with anyone pointed at discovery on this scale - CZI, Arc Institute, or elsewhere 🤗
SmallField toolkit: https://t.co/JgLVyqRmkR
USearchMolecules dataset: https://t.co/eAlVuO4nVv
pleased to announce an update of https://t.co/IN9SXP0bXn, which has been mothballed for a couple of years when the previous grant ran out. thanks to support from the Alfred P. @SloanFoundation we have updated the patent-to-paper citations through 2025.
one change in this release is that the citations are limited to patents granted by the USPTO. this will be our scope going forward.
this release also fixes a bug whereby the # of in-text citations for 2021 and 2022 were artificially low.
lastly, the system has been ported from the old pile of Perl scripts to Python, so you can more easily download the code and replicate what we did. you will need a server with about 400G of memory to run it.
happy researching!
New research from the Growth Lab and @CSHVienna, published in @NatureComms, examines the two branches of economic complexity (the Principle of Relatedness and complexity indices).
The research offers clarity about the formal relationship between them, and suggests that prominent complexity metrics describe overt changes in activities that economies undergo, rather than infer information about capabilities.
These changes in how we view complexity metrics may affect their use in policy settings. 🧵 1/15
🔗 https://t.co/TbTGBdvCM5
Introducing a new way to see our planet! 🌍 🛰️ Using Google DeepMind’s new geospatial AI model, AlphaEarth Foundations, we’re making satellite data more analysis-ready, packing a year of info into each pixel. https://t.co/bxqV7bnXkG #EarthEngine#AIforGood
Played with AI tool Illuninate from Google, which generates a 5min podcast automatically for the joint work "Evaluating the principle of relatedness" with @FrankNeffke on Research policy.
Podcast at: https://t.co/HNttme6daS
Article at: https://t.co/E2DcCMf0fB
We are happy to release DuckDB v1.1.0 “Eatoni”.
The new release packs a ton of new features: friendly SQL extensions, performance improvements and spatial features. It also includes improvements towards supporting Community Extensions.
See our blog post: https://t.co/nl4AQup2Va
Understanding and predicting human mobility in urban areas is very important. Unfortunately, there are no large, standardized, and open datasets to compare different models. Here, we introduce a new open, anonymized dataset of 100k individual mobility trajectories. Check it out!
📢 We've added 2024 data to "Metroverse" - our urban economy navigator tool.
Uncover the economic composition of your city. This snapshot of #Johannesburg 🇿🇦reveals that 32% of city employees are in the professional & business services industries.
https://t.co/tft5SXvEF9
📢 We're hiring!
Seeking postdoctoral researcher to spearhead research on #GreenGrowth, analyzing supply chains for #CleanEnergy technologies and leveraging this data to influence strategic industrial policy.
#EconJobs#EconTwitter@ricardo_hausman
https://t.co/8ds1915meh