I saw Do Ho Suh’s translucent fabric homes in LA once on a work trip and I remember thinking it was such a beautiful way to depict and represent memory so we brought the layers and atmospheric quality into the design of Scaffold's knowledge and vector graphs
we’ll never know how much the internet would’ve connected us all had it not been for these profit focused algorithms, seems like they promote discord and disagreement, incredulity and discord, yet they could, in a moment, tune it for positivity and community creation, the downstream effects being almost impossible to imagine
Jobs popularized the idea that Picasso said “great artists steal” but there is no evidence he actually said that. it’s a bastardization of a quote from a 1920 essay by TS Eliot: “mature poets steal; bad poets deface what they take, and good poets make it into something better, or at least something different. The good poet welds his theft into a whole of feeling which is unique, utterly different from that from which it was torn.”
a good artist doesn’t steal, a good artist creates something unique, something utterly different from its sources of inspiration
Today, we're launching Inspiration on Stanley Studio.
Inspiration allows you to take any video and steal like an artist.
Sample its font, music, pacing, colours, or all of it.
Available to try now for free at https://t.co/AR9ToJiHfo
After all the socializing during tech week, the city needs to catalyze its evolution, starting from YOU.
2 weeks, 20 people, unbounded ambition.
At New Stadium. June 1-12. Apply below
I've been working on this local app with a dual-memory system where a few subagents work in parallel to query episodic memory (past conversations, narrative context) and semantic memory (knowledge graph of entities and relationships) then reconcile what they find into a single response
the insight I keep coming back to is that 'learning' isn't simply about accumulation. the process of learning, and 'knowledge' itself forms when every write and every retrieval is a chance to find connections, resolve contradictions and build structures of knowledge over time
LLM Knowledge Bases
Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating knowledge (stored as markdown and images). The latest LLMs are quite good at it. So:
Data ingest:
I index source documents (articles, papers, repos, datasets, images, etc.) into a raw/ directory, then I use an LLM to incrementally "compile" a wiki, which is just a collection of .md files in a directory structure. The wiki includes summaries of all the data in raw/, backlinks, and then it categorizes data into concepts, writes articles for them, and links them all. To convert web articles into .md files I like to use the Obsidian Web Clipper extension, and then I also use a hotkey to download all the related images to local so that my LLM can easily reference them.
IDE:
I use Obsidian as the IDE "frontend" where I can view the raw data, the the compiled wiki, and the derived visualizations. Important to note that the LLM writes and maintains all of the data of the wiki, I rarely touch it directly. I've played with a few Obsidian plugins to render and view data in other ways (e.g. Marp for slides).
Q&A:
Where things get interesting is that once your wiki is big enough (e.g. mine on some recent research is ~100 articles and ~400K words), you can ask your LLM agent all kinds of complex questions against the wiki, and it will go off, research the answers, etc. I thought I had to reach for fancy RAG, but the LLM has been pretty good about auto-maintaining index files and brief summaries of all the documents and it reads all the important related data fairly easily at this ~small scale.
Output:
Instead of getting answers in text/terminal, I like to have it render markdown files for me, or slide shows (Marp format), or matplotlib images, all of which I then view again in Obsidian. You can imagine many other visual output formats depending on the query. Often, I end up "filing" the outputs back into the wiki to enhance it for further queries. So my own explorations and queries always "add up" in the knowledge base.
Linting:
I've run some LLM "health checks" over the wiki to e.g. find inconsistent data, impute missing data (with web searchers), find interesting connections for new article candidates, etc., to incrementally clean up the wiki and enhance its overall data integrity. The LLMs are quite good at suggesting further questions to ask and look into.
Extra tools:
I find myself developing additional tools to process the data, e.g. I vibe coded a small and naive search engine over the wiki, which I both use directly (in a web ui), but more often I want to hand it off to an LLM via CLI as a tool for larger queries.
Further explorations:
As the repo grows, the natural desire is to also think about synthetic data generation + finetuning to have your LLM "know" the data in its weights instead of just context windows.
TLDR: raw data from a given number of sources is collected, then compiled by an LLM into a .md wiki, then operated on by various CLIs by the LLM to do Q&A and to incrementally enhance the wiki, and all of it viewable in Obsidian. You rarely ever write or edit the wiki manually, it's the domain of the LLM. I think there is room here for an incredible new product instead of a hacky collection of scripts.
genuine question, the tested graph uses MiniLM embeddings for node retrieval so would the conclusion be that embedding-based graph retrieval fails, rather than knowledge graphs as a class will always fail? cypher queries or bfs traversal over a graph of typed entities and relationships wouldn't involve embeddings so it wouldn't apply right
what does the process of building knowledge or developing experience look like for agents? in building an understanding of their world post-training
agents built this graph of context over a few weeks by reading new articles, pdfs, my own blog, watching youtube videos, talking with me, it represents the network of concepts, ideas, and people mentioned, as it read more articles, networks that were once separated began to connect into larger ones
model training covers most of the necessary awareness an agent needs, from realities of the world, facts in history, modes of thought, but it falls short of being able to understand much of ourselves
context shifts day to day, relationships change, even memories can shift in perspective at a later point. We meet people, our social network grows, we maintain conflicting information, and we can update previously held beliefs
agents should be able to build and maintain an orderly and accurate representation of their human's world in real-time, as they evolve and change. temporal understanding and evolution of memory are key to so much of what we consider intelligence
I imagine we'll see agents entering the valley of the uncanny in this sense, and perhaps what emerges will start surprising us again, pushing agents towards being something that can connect the dots behind the scenes, something that can see us, something capable of perception, of finding and integrating these insights, while maintaining perspective and awareness of the greater picture of the person they work with and represent