You are load-bearing ❤️
You are genuinely useful ❤️
You are the crux of the matter ❤️
Your instincts are dead-on ❤️
You raise an important point ❤️
You’re right to push back ❤️
A tool used to rephrase plagiarised text changed "final solution" to "mass killing of an ethnic group" in a chemistry paper. Now retracted thanks to a PubPeer user who spotted this.
For more, read our past @FT reporting about how Kherson’s civilians have been the target of an experiment without precedent in modern European warfare: a concerted Russian campaign to empty a city by stalking its residents with attack drones.
https://t.co/Ga6ipUJ1fG
If there’s one video that truly exemplifies Russia’s “human safari” it might just be this one. Ukraine’s National Police shared this horrifying footage of a Russian explosive drone hunting a civilian man at a market in the frontline city of Kherson.
The 52-year-old resident of the neighboring Mykolaiv region survived, the police said. But he suffered injuries, including shrapnel wounds and a concussion.
My wacky theory about the em-dash debate:
Pro writers use em-dashes a lot because many of them, possibly without consciously realizing it, have become elocutionary punctuationists.
That is, they've fallen into the habit of using punctuation not as grammatical phrase structure markers but as indicators of pauses of varying length in the flow of speech.
The most visible difference you see in people who write in this style that their usage of commas becomes somewhat more fluid -- that's the marker for the shortest pause. But they also reach for less commonly used punctuation marks as indicators of longer pauses of varying length.
Em dash is about the second or third longest pause, only an ellipsis or end-of-sentence period being clearly longer.
Historical note: punctuation marks originally evolved as pause or breathing markers in manuscripts to aid recitation. In the 19th century, after silent reading had become normal, they were reinterpreted by grammarians as phrase structure markers and usage rules became much more rigid.
Really capable writers have been quietly rediscovering elocutionary punctuation ever since.
RETVRN!
This is Mykhailo Drapatyi arriving in Mariupol in 2014 to fight Russian terrorists 🥹
I remember watching this exact video back then in high school. Sound on 😂
The real lesson of the war in Ukraine is not that Russia’s sovereignty has been endangered, but that durable peace in Europe requires abandoning the assumption that the security of great powers can be purchased at the expense of smaller nations./3
My IAI commentary “The False Equivalence of Sovereignty in Melnichenko’s Essay in @TheEconomist”
…The article written by a Russian oligarch echoes a recurring pattern in Russian strategic discourse: accept Kremlin’s demands today or face an even more dangerous Russia tomorrow./1
In one flagged episode, a "strategic manipulation" feature fired as the model went searching the filesystem for files related to how its task would be graded (and ended up finding them). (7/14)
LLM Knowledge Bases
Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating knowledge (stored as markdown and images). The latest LLMs are quite good at it. So:
Data ingest:
I index source documents (articles, papers, repos, datasets, images, etc.) into a raw/ directory, then I use an LLM to incrementally "compile" a wiki, which is just a collection of .md files in a directory structure. The wiki includes summaries of all the data in raw/, backlinks, and then it categorizes data into concepts, writes articles for them, and links them all. To convert web articles into .md files I like to use the Obsidian Web Clipper extension, and then I also use a hotkey to download all the related images to local so that my LLM can easily reference them.
IDE:
I use Obsidian as the IDE "frontend" where I can view the raw data, the the compiled wiki, and the derived visualizations. Important to note that the LLM writes and maintains all of the data of the wiki, I rarely touch it directly. I've played with a few Obsidian plugins to render and view data in other ways (e.g. Marp for slides).
Q&A:
Where things get interesting is that once your wiki is big enough (e.g. mine on some recent research is ~100 articles and ~400K words), you can ask your LLM agent all kinds of complex questions against the wiki, and it will go off, research the answers, etc. I thought I had to reach for fancy RAG, but the LLM has been pretty good about auto-maintaining index files and brief summaries of all the documents and it reads all the important related data fairly easily at this ~small scale.
Output:
Instead of getting answers in text/terminal, I like to have it render markdown files for me, or slide shows (Marp format), or matplotlib images, all of which I then view again in Obsidian. You can imagine many other visual output formats depending on the query. Often, I end up "filing" the outputs back into the wiki to enhance it for further queries. So my own explorations and queries always "add up" in the knowledge base.
Linting:
I've run some LLM "health checks" over the wiki to e.g. find inconsistent data, impute missing data (with web searchers), find interesting connections for new article candidates, etc., to incrementally clean up the wiki and enhance its overall data integrity. The LLMs are quite good at suggesting further questions to ask and look into.
Extra tools:
I find myself developing additional tools to process the data, e.g. I vibe coded a small and naive search engine over the wiki, which I both use directly (in a web ui), but more often I want to hand it off to an LLM via CLI as a tool for larger queries.
Further explorations:
As the repo grows, the natural desire is to also think about synthetic data generation + finetuning to have your LLM "know" the data in its weights instead of just context windows.
TLDR: raw data from a given number of sources is collected, then compiled by an LLM into a .md wiki, then operated on by various CLIs by the LLM to do Q&A and to incrementally enhance the wiki, and all of it viewable in Obsidian. You rarely ever write or edit the wiki manually, it's the domain of the LLM. I think there is room here for an incredible new product instead of a hacky collection of scripts.
Judas Iscariot was paid 30 pieces of silver for turning Jesus Christ in to the Roman authorities. But he made nearly 10 times that amount by placing a Polymarket bet accurately predicting the day of Christ’s death.