Ten Grok Bot templates for the actual founder loop: survive, decide, ship, review, sell, post. Plus a shared memory repo so the whole org gets smarter at once, not ten private diaries.
A pile of bots is not a startup. A pile has no graph, no memory. https://t.co/igEVrUlOKv
Paper: SemNav: Semantic Navigation for Repository-Level Issue Localization
Your coding agent is hunting the files that matter for a bug report. The repo is huge. It keeps opening whole source files and guessing.
SemNav is a harness for that hunt (the code around the model that decides what it can open, how it jumps between symbols, and which suspects stay on the list).
Ordinary search seeds a broad candidate list. The agent revises that list as it learns more.
Go-to-definition jumps to the real callee across files instead of searching among name collisions.
Short issue-conditioned cards summarize what a function does for this bug, with a few supporting snippets, so the agent does not have to keep the whole file in working memory.
A candidate workspace keeps each suspect with the evidence that still justifies it.
Same Gemma 4B model on Software Engineering Benchmark Lite (SWE-bench Lite).
Share of issues where the correct file appears in the top 10 guesses:
Strongest prior: 68.33%.
SemNav: 82.67%.
Those better file picks feed the repair step. With the same localization model feeding a fixed repair backend, resolved issues move from 44.00% to 52.33%. Most of the lift is the localization harness, not a bigger writer.
@TheAhmadOsman@victormustar Seems like all chinese labs are just postponing releases since Astra and Opus 5.5.
Sol 6.1 and Gemini 4 Argon are the only model relevant on pareto frontier with opensource
This is significant breakthrough in context engineering
Basically turns context into active learnings stored and fetched from memory on priority and relevance.
Need a detailed walkthrough on this
‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA!
Introducing 🩵Context Language Models (CLMs)🩵
- Natively manage their own context
- Treat context as a file
- Learn policies in CLM weights, no harness