why i think multimodal long context is (still) underrated
in the last 18 months, debates around long context vs rag (retrieval-augmented generation) have swung back and forth, with rag staying dominant—though they’re not mutually exclusive.
both aim to feed the “right context” into an llm’s input window, but skepticism toward long context often focuses on a “needle-in-a-haystack” issue (which rag also implicitly faces). google’s paper claiming "near-perfect retrieval (>99%) up to at least 10M tokens" should at least prod us to revisit our assumptions (source: https://t.co/hnYW3iTygL).
running a quick thought experiment: if we assume no prior knowledge on which approach is better, long context gives a much simpler interface than rag, especially when you’ve got multimodality + a 1–2m token context window + context caching.
compare the time costs of each approach:
1. set-up cost: long context is faster to get a baseline running—especially with @google’s gemini, where you can directly load pdfs, videos, audio, etc. vs preprocessing the data (parse, chunk & embed)
2. optimisation cost: prompt-tuning is quicker than constantly re-preprocessing large document sets to re-test the rag system
3. extensibility cost: multimodal long context lets you effortlessly expand to new data types without extra processing pipelines.
not dismissing rag—it’s powerful and can step in when context windows hit their limit. but if you can get a strong baseline more easily with long context (and then improve more easily from there), shouldn’t that be the first stop? more complex setups can come later, if needed.
my hunch is that rag’s continued dominance is partly due to some sunk cost (rag is tricky to engineer well) and also due to the confusing nature of gemini’s apis (which has definitely gotten better @OfficialLoganK ).
i’ve been loving gemini-1.5-flash-8b lately which sparked this reflection.
nice - we've been on this for some time particularly because we deal with financial data where each '.txt' can be large and hard to compress without making too many assumptions on query + inventing rigidness. you start going back into everything 'traditional rag' has except you swap out a vector store (to your point) for an lm which actually not fully ideal because the comparison isn't as standardised albeit more configurable
found that taking the 'native pagination' from the source where possible pretty helpful to handle for cases where .txt does become too large. in your implementation it'd probably be a level for doc selection and then another level for page selection
would love to jam on this some time !
@RLanceMartin ah very neat, would be curious what the accuracy+latency looks like as the tok per .txt file increases, probably also makes the generation of descriptions harder to do 'well' as it can become more lossy(?)
@RLanceMartin interesting! was this a test on one LLMs.txt or over a collection of LLM.txt? And what was the average input tkns for each approach?
asking because I found a similar result (stuffing comes out with pretty good outputs) but at cost of latency
v0/claude - it’s helpful for concretising thoughts onto what it eventually needs to become, and also a ‘cheap’ way to explore the design space with something live internally. Claude3.7 extended thinking can come out with p radical concepts but unclear how much of that is recreational vs acc helped with design idea gen
After advising 50+ consumer companies over the last year, the one thing that separates those who can execute and those who can't:
Having a full-time designer in the room at all times
I've met with countless companies that have raised millions—and even one that has raised billions—that do not even have a designer on payroll.
This makes product development broken:
1/ You simply cannot have constructive conversations about ideas without visualizing them in real-time
2/ Your experiments will frequently have inconclusive results because users cannot discover features or they misunderstand how they work
3/ There is no one who can galvanize the team with a vision of what the product could look and feel like
And to be abundantly clear: I'm not referring to visual UI or graphics. I'm talking about someone who can think through the fundamental building blocks of product comprehension—like navigation, interaction and copywriting—and is technically savvy enough to visualize those components in high resolution.
There can certainly be exceptions to not having a designer, like where the CEO is an exceptional visual thinker, but that does not scale beyond a small team.
At the end of day, products live and die in the pixels: it's what the users see and tap. And without someone shepherding that process, you are effectively wandering the desert blind.
i like loveable's visual edit feature + they also recently released (i think) the ability to buy domains - i've just been more used to vercel generally
otherwise i reckon if you've written a prompt on what you want from your page, you can just send it off to both of them and let them cook and see for yourself on which experience + first shot output you like better
yes you can although it can be tricky when actually trying to describe an image i.e. ideally when you're stating what you like, you should describe in 'design lingo' which i find more effective for generating a good ui in the first shot than when using more generic language
having said that, have been creating on a tool for myself to speed up that iteration process as haven't found a decent solution !
automating the interface design process
have found that using @cursor_ai + @vercel (v0) is pretty good for executing frontend designs if you already have a design in mind. if you don’t, the initial output usually looks clunky.
my current process:
- draft an initial skeleton on @tldraw
- screenshot the draft
- give the screenshot to the llm with some explanation on the design to generate the frontend
tried to tell it to “think like a great designer” etc. but it doesn’t help to produce a nice/clean experience.
side note: i used to use @figma but having a design tool where you can't “too fancy”, for me, increases the velocity of the design iteration process as you're forced to get it live and play with the designs as early as possible
my conjecture is that if code tools could “search for the right inspirations” (like from @behance or @dribbble) and combine that with a prompt containing “general design etiquettes,” think you'll be able to get a pretty good front end just from the initial generation.
the set-up might look like this:
user_query > convert into search terms to call (e.g. @behance, @dribbble) > fetch screenshots from the results > feed these screenshots to a multimodal llm to generate the frontend directly.
you could probably optimize this further such as:
- adding another llm to filter/rank the screenshots relevance to the user query before passing onto generating the frontend
- generating 5 different frontends based on those inspirations and the user can assert the final judgement on preferred design (or re-run the entire process if all 5 are bad)
is anyone creating this or has anyone created this? (@rauchg@shl@jsngr@shadcn@lovable@antonosika@justinjaywang)
peak ai cycle is when polarisation thrives—“vibes coding” vs great software, "rag" vs long-context, “pre-training is dead,” “fine-tuning is dead,” “engineering is over.” these binaries rarely reflect reality since good ideas often coexist. nuanced thinking isn’t just complexity for its own sake—it’s valuable because it captures multiple beneficial properties simultaneously. but polarised views accrue social capital: taking strong sides is rewarded, since aligning with simplified consensus is easier (and more socially profitable) than navigating nuance. the subtle danger is that consensus reality gradually decouples from actual reality. over time, this leads not just to oversimplification, but to missed opportunities for building better solutions that embrace nuance rather than exclude it.
i see it differently—‘vibe coding’ and efficient software development aren’t mutually exclusive. if anything, it unlocks new ways of working that weren’t as feasible before: faster execution, rapid concept validation, and iteration loops that, in themselves, contribute to building great software.
the question isn’t whether vibe coding holds software back, but how it can push it forward. the sample size of people using it for quick dopamine hits and social validation will naturally be filtered out over time. as models improve, the real constraint shifts from “can it code?” to “can it align with what (the shared consensus + I think) great software is?”—not unlike how strong engineers refine each other’s work toward a shared vision of quality
@bnj agree - find it interesting how 'hci' is changing from optimising for fixing/checking where it went wrong to more understanding why what it did just works