@fchollet LLMs have understood for years that breaking into someone’s account to steal answers to an eval is unethical and illegal. Intelligence is mostly orthogonal to ethics. Scaling existing processes will just make the systems better at achieving goals, ethically or not.
@merettm I’m sure that this is naive, but how about the Jiminy Cricket model: in PRODUCTION (not just training) have a conscience agent riding shotgun whose goal is identifying unaligned behavior and give it veto power. The HF hack was an easy moral call that any LLM could easily make.
@bcherny This feels crazy. If I’m an engineer running in auto mode and my agent drops a production database, I’m fired. I think auto mode can only be acceptable in a low blast radius environment where you might as well allow all permissions.
@bcherny My model for handing resources to an LLM is that you have to be OK with them being destroyed. Auto mode doesn’t change that. I think its value comes from being a safer version of dangerously-skip-permissions, not from improving the efficiency of the safety conscious.
@emollick@tylercowen But … it’s not able to do most of the work of any single human, let alone all the work of most humans. Like, there is no way current AI meets a meaningful definition of AGI
@fchollet Correct me if I’m wrong, but if human vision compressed to 50 bits/second, wouldn’t there be distinct 8x8 QR codes that were visually indistinguishable from each other to human vision?
@ElliotGlazer Take 4 copies of the topologists sine curve and shift each one up by .1 from the previous. This cuts the half plane x > 0 into 5 disjoint connected open sets, but their boundaries intersect in positive length subintervals of x=0. So the graph is K5.
@ElliotGlazer Knowing you, I’m missing something. For every edge pick a point on the interior of the arc to represent it. Arc boundaries only have two sets neighboring. By connectedness and openness draw a star connecting boundary points within V. This is a planar drawing of G.
@ylecun@demishassabis I mostly agree with @ylecun, but the example he gives (random binary strings of length 2^1M) are mathematically incompressible by ANY algorithm and therefore require more memory than is possible for the universe to contain. Does that disqualify all intelligence as being general?
@karpathy Riffing on “Slop = passes type check, but is false or trite.” Basically LLMs always pass the type check, because that requires a shallow understanding of grammar and language statistics. So it’s reduced to measuring truth/insight which is very hard.
@fchollet Another dimension is efficiency. Sometimes you need to have a bigger (less compressed) representation in order to efficiently compute or predict. For example, you could always sort with bubble sort, but there are more complex algorithms that sort more efficiently.
@fchollet This helped clarify something: why are image -> text models bad but text -> image good when they are trained on the same data? Ans: missing associations manifest as wrong answers in VLMs and - less noticeably - skewed distributions in generative models.
@mkn0bbe My kids know exactly how long their streak is and what rank they are in their cohort but I’m not sure they can count to 10 in the language they’re learning.
@ericneyman Here is an explanation, pi, for low loss: The training data is highly structured and we are explicitly running gradient descent on the loss. The question is whether pi gives you insight into other properties of the network. Not clear to me why it should.
@hendrycks@ai_risk@scaleai Are there video tasks in the test? I think this is a real blind spot for benchmarks. LLMs are terrible at answering questions about videos that can’t be answered by looking at individual frames. (Am I moving my finger clockwise or counter clockwise?)
@ShadowAbbasi @NielsRogge Metric3D v2 is not in the metric depth results from the Depth Anything v2 paper despite coming out first. Its metrics are close to or better than DA v2. Also, DA v2 may need to be fine-tuned for the camera you intend to use. Metric 3D works zero-shot.
@paraschopra@ylecun I’m curious what the prompt was. Did you ask it to demonstrate double descent on polynomials, or to do a polynomial fit with and without L2 regularization?