@fchollet What’s next, you’re going to say the rules of physics are also symbolic architecture when we get robots? The millions of lines of code aren’t doing any of the interesting thinking or reasoning, they’re just doing thoughtless automation that we’ve had for decades.
@max_spero_ $100 is way way too low for legitimate content that is falsely flagged. For comparison on the flip side, the damage a false positive could cause to someone’s reputation, or getting an innocent student kicked out of college because they were accused of cheating, is far higher.
@GracchiSimp@halcyon_hazel Comparative advantage — it’s their full time job, so they can be more efficient at it. Also, economies of scale, they might manage many properties and get a discount at scale for property management related expenses
@mattkrisiloff@Conception Will men be able to reproduce with themselves to make an offspring that is 100% their DNA? If I understand correctly, they wouldn’t be Identical clones, right?
New paper! We used Sparse Autoencoder (SAE) embeddings to understand how agent behavior actually changes during post training. We analyzed thousands of rollouts in a Diplomacy environment and uncovered generalized reward hacking, weird roleplay, and the root cause of an unsuccessful training run. All of which our LLM baseline couldn't find.
I’m excited to share a bit about what I’ve been working on in data-centric interpretability. It’s a paradigm with great potential, and it’s ready for application.
Our production system has already been tested on real customers’ data to improve their engagement metrics. But we can’t share that, so we’ve made a synthetic dataset to demonstrate. Basically, pretend you’re a marketing firm trying to generate messages that engage users, and they have this hidden preference in the data. LLMs and embeddings fail to catch it, but our system can.
Data-centric interp is efficient and good at discovering unknown unknowns. We’ve demonstrated it works on short texts.
But what’s really exciting is to extend this to longer contexts and larger datasets. Monitoring agent traces is costly. Training datasets are too massive to even try. But data-centric interp gives you rich representations of data, which can systematically uncover insights, anomalies, and other qualities over large datasets.
We can help find these hidden features in your data that will drive your metrics.
I’ll be at NeurIPS and am always down to chat data-centric interp
@DevinOlsenn It’s not great, but there’s actually a software-only fix for this: when car starts, back up a foot before going forwards to see what’s below cam.
@elonmusk Well he stacked the Supreme Court with judges that agree with him, and same with heads of different federal organizations and White House staff. America has more protections in place than 1940s Germany, and he’s slowly getting rid of them
@karpathy Thinking about this one step further, what policy does your brain use to choose a random number? Maybe it has bias too? Can you improve upon it if you think about it?
@jxmnop Interesting that it becomes a house in one of the middle layers. In one of the other photos in the paper, a butterfly becomes a car. I wonder if this is a new source of computer vision adversarial attack. Ex: Looks at photo of house-> “I see a giraffe”.
@paulg Now graph the % of the population capable of affording a home and you get the opposite story. Income and quality of life aren’t always the same. Speaking of which, compared to other countries, our social services and public infrastructure are lacking.
@jxmnop Yeah but these things should be added before publishing, otherwise no one else will understand the code or be able to replicate. In fact I would argue this is probably a cause of the replication crisis. It’s an externality of moving faster.
@ID_AA_Carmack But this is probably only single nucleotide mutations (SNPs), right? There could be any number of copy errors and other kinds of mutations. Still pretty crazy to think about