Hello world!
I'm part of an innovation first team at @TomTom where we build AI-native location tools. MCP servers, maps SDKs, agents that understand geography.
Starting to share what we ship, what I build on the side, and what I learn along the way.
Let's go 🚀
La crisis de vivienda en España: lo que creo que puede funcionar para solucionarla... y lo que probablemente no sirva o sea directamente contraproducente. En siete minutos y medio.
Dentro video 👇
PRESENTAMOS https://t.co/Mw8ZZh8rV1
Estos días ha habido mucho revuelo con una iniciativa comunitaria para mejorar los servicios digitales en España.
@disamdev y yo llevamos meses con una idea parecida, y semanas trabajando en ella. Ahora hay 2 proyectos. La idea es llegar a todo. 🧵
Embeddings do not solve everything. Focusing on specific outcomes is what improves them. Excited to keep pushing what we learn out to the geospatial community.
Almost two years ago I read my first text on geospatial embeddings: representing places as vectors. Since then I have been exploring how to scale and generalize geospatial data access. Some things I learned along the way 🧵
So what are they good for? Clustering, semantic similarity, change detection. And as a data product: they are amazing pre engineered training features for ML models. Better shape and faster, cheaper convergence than long feature engineering pipelines.
Every major vision AI model looks at satellite imagery the same way you look at a photograph.
It's flat, without a concept of height, slope, or elevation.
GPT-4o scores 21.61% on tasks that require understanding how tall a building is or which direction water would flow downhill. Gemini-2.5-Flash gets 28.45%. The best open-source models hover around 22%.
A team built a benchmark called GeoHeight-Bench to test whether large multimodal models can reason about vertical geometry from satellite images. Seven task types across pixel, object, and scene levels: elevation retrieval, height ranking, terrain slope, flood inference, landslide identification. The kind of questions that matter when you're mapping disaster risk or planning urban infrastructure.
The results are bleak. On the harder benchmark (GeoHeight-Bench+), GPT-4o drops to 15.12%. LLaVA-NeXT-34B, a 34-billion parameter model, manages 15.42%. On landslide and flood inference tasks, pure optical baselines score near zero. They can count buildings and classify land cover. They can't tell you which buildings sit on a ridge and which ones sit in a flood plain.
The core problem is a modality gap. RGB satellite images encode colour and texture. They don't encode elevation. Digital elevation models (DEMs) do, but they're not available at inference time for most real-world applications. So the question becomes: can you teach a model to hallucinate height from flat images alone?
GeoHeightChat does this in two stages. First, a lightweight adapter (GeoAdapter) learns to predict dense geometric features from RGB images by distilling knowledge from a teacher network that has access to actual elevation and land-cover data. The adapter uses a zero-initialisation strategy so it starts as an identity function and gradually learns geometric shifts without destroying the pretrained CLIP features. Second, those hallucinated geometric features get fused into a LLaMA-2-7B backbone through an adaptive residual connection for instruction tuning.
At inference, the elevation data disappears entirely. The model takes a standard satellite photo and reasons about height, slope, and terrain as if it had a DEM in front of it.
GeoHeightChat scores 44.14% overall on GeoHeight-Bench. On GeoHeight-Bench+, it hits 65.59%. On height ranking, it reaches 94.87% where GPT-4o scores 5.74%. On elevation retrieval at pixel level, 76.35% against GPT-4o's 0.00%. The ablation confirms that height supervision alone delivers most of the gain (35.11% overall vs 12.58% baseline), while adding semantic class boundaries pushes it to 42.94%.
A 7-billion parameter model with a bolt-on geometric adapter is outperforming models 5x its size on spatial reasoning tasks that actually matter for Earth observation. The vertical dimension was always there in the data. No one had forced a vision-language model to learn it from RGB alone until now.
Link to full paper: https://t.co/z4Ssm8PJFt
Then the most common one is to give LLMs access to geospatial tools via MCP: maps, routing, POIs, spatial analysis. But is tool access really geospatial awareness? Or is it even the best way to do so?
I feel like a lot of people were already doing something very similar to this with quick html mocks and multiple options. The same way the Claude Design per say did not catch my attention I am curious to try this
Claude Code can design now. The new /design skill (research preview) brings Claude Design's artboard workflow into the CLI and Desktop, built on artifacts.
Run /design to get editable artboards for your UI — pick one, tweak it, then have Claude implement it.
This is my "feel the AGI" moment: I used GPT-5.6 Sol to train my own autocorrect model that outperforms GPT-5.6 Sol (wtf??)
I have no ML background. I have no idea what I'm doing. I just kept pushing Sol until it spat out a SOTA model. And I spent $0.
The motivation: Years of talking to AI have made me terrible at typing. Rather than fix my skill issue, I decided to throw more AI at it. My idea was: instead of autocorrect that interrupts my flow, I want to type fast with mistakes and have AI clean it up after.
I wanted the smallest local model possible, for speed, for battery life, for science! So I decided to train my own.
Inspired by @karpathy’s autoresearch, I ran Codex /goal with this setup: pick an experiment, try it, record the results to a doc, throw it out if it fails, and plan the next experiment without repeating failures. I gave a few examples that had to pass, tight latency targets, and let it run.
Sol did some amazing things.
First, it scanned benchmarks and shortlisted base models: Qwen 3.5, Gemma 4, Liquid LFM 2.5. It found a dataset on HuggingFace for typed text.
Then it built a simulator for fingers striking a Mac keyboard, modeling the physical layout with a Gaussian distribution around each key. It simulated striking the wrong key, wrong order, fat-fingering, etc.
With the models + data + simulator, it fine-tuned using MLX right on my MacBook. It had a working prototype within an hour! But accuracy was pretty poor.
—
Problem 1: Tokenization
Sol read papers, ran tests, and identified that the tokenizer was the bottleneck. Tokenization makes typos hard for the model to see, so it memorizes mappings instead of using its language priors.
Sol tried ByT5, Google’s tokenizer-free byte-level LLM. This made a big improvement, but the model is old and lacked the knowledge needed to reach Sol performance.
Sol dug deeper and realized a tokenizer-free model isn’t needed; instead, it used T5Gemma, an encoder-decoder model. This can understand the input deeply before producing output, and furthermore, Sol could post-train the encoder to improve performance. This gave a much higher ceiling.
—
Problem 2: Loss function
Now the model was correcting some typos perfectly, but ignoring most. Sol realized that standard cross-entropy loss was teaching the model to avoid edits, because the vast majority of characters in the training data were left unmodified.
The fix was wild: Sol wrote a custom loss function that byte-aligns the source and target strings, uses a dynamic programming algorithm to compute the minimum edits between the two, then weights correct edits much higher than copies. After a lot of tuning, this dramatically improved accuracy.
—
Problem 3: Autoregression
One failure mode remained: if the model made a mistake, it couldn’t backtrack. It could only predict the next token. Teaching it to “think” like a reasoning model would solve this, but would be far too slow.
Sol found a beautiful solution: instead of greedily predicting the next token, beam search over all possibilities. This parallelizes the exploration instead of one linear chain-of-thought. At the end, choose the path with highest cumulative log probability.
This worked great, but made the experience worse, since the user wouldn’t see progress until the whole search was done. To fix this, Sol made a clever observation: after each search step, the longest common prefix among surviving branches is guaranteed to appear in the final result, so it can be displayed immediately. As the search progresses, weaker paths are dropped and the prefix grows, so the user sees continuous progress.
Sol built all this as a custom MLX pipeline that does the parallel decoding on the MacBook GPU, with just ~40ms TTFT. It’s crazy fast and entirely local.
—
Final eval (error reduction rate, higher is better):
- Apple autocorrect: 49.66%
- GPT-5.6 Luna: 82.47%
- GPT-5.6 Terra: 87.64%
- GPT-5.6 Sol: 90.56%
- Our model (1.7B): 91.02%
Final cost:
- 1 quota reset (thanks @thsottiaux)
- $0
(And yes, I verified there's no cheating. In fact, we test words scrubbed from the training data to prove the model isn’t memorizing)
There were a ton more details and tangents I could write about: contrastive learning, GRPO, DPO, dynamic masking, and more. Sol is a fascinating and creative model. It blew my mind so many times.
Don’t let a lack of experience stop you: Sol makes AI experiments accessible to anyone!
📷 Zonas buenas y menos buenas para observar el eclipse en territorio de España. Líneas azules: altura del Sol sobre el horizonte. Flechas rojas: dirección del Sol en el momento del máximo.
📕 Revista N.º 325-326 (Julio-Agosto 2026)
Agenda Cenit
Antonio Bernal