@kalomaze@willcb this has spiritual analogues in non-vision stuff and in theory countless task-specific ways it can be addressed, but arc-agi does a solid job of making it a bit more salient and less gameable than it’d be with another backdrop
@kalomaze@willcb if I get what you're saying it just folds into why the benchmark/failure is informative, the model /can/ "see" well enough to deconstruct, reformat etc. the input itself but it can't id the pattern even where ~any human that could "see" as well could (bc the "see" is "cheap")
@willcb am old enough to remember unironic versions of "bro this benchmark sucks bc our model hasn't been trained on this exact sort of visual processing task so how's it fair" flying around to hundreds of likes
@damekdavis or w/e form of this question you’d find most informative to answer. icl-based optimization seems pretty powerful for pushing for an extreme case of something that isn’t necessarily easily sgd-learnable, but idk enough to contextualize any of these results
@damekdavis do you know how you’d approach b1 with an ~arbitrary compute budget (say, something like what you’d estimate the alphaevolve authors had)? without frontier llms (let’s say involved in the program optimization loop itself, still allowed otherwise)? with them?
depends on the balance between complexity of the problem and how important it is to align or make a decision on a timer. settings in which need for consensus outweighs need for nuance are more meeting-heavy and vice versa
some people also benefit from the “aura buff” audio/video provide in technical problem-solving just due to mood or “infectious energy” reasons, though maybe this just folds into the “need to align” above
I'm not well-versed enough in [strong] longtermism to know whether this piece gives it a fair shake, but many of the opinions here make more sense in the context of a framework that projects x as meaningless in comparison to y and (by proxy) allows for high uncertainty tolerance wrt reasons for deciding what to do in the present
useful reference. my glaring issue with being anti-progress is that (delay -> ruin x lives) is low-uncertainty, easy to model, and ~guaranteed whereas (delay -> save y lives) is uncertain, yet-impossible to model, and low-probability. seems the strategy is to hopefully figure it out while x lives are being ruined
@eshear what maximizes a user’s engagement (on here, anyway) likely doesn’t have much overlap with what the user would indicate wanting to see more of
you’d need some micropayments type thing to fix the incentives
@effectfully@ludwigABAP incentives will rotate around this, who knows when and how much
llms also make it easier for trained eye to spot garbage without spending too much time (and speed the “get exposed” cycle in other ways)
by “your” I meant the team, should’ve specified
question is why bullposting from monad team members themselves (excluding you, which is why I’m asking you this, and insofar as I’ve seen it) is ~indistinguishable from bullposting from projects that have nowhere near the same technical value props
seems like a drag-down/undersell to me, but maybe unavoidable if you want to bootstrap an eco? wouldn’t know
[think about some field a lot] -> [become attached to your conclusions (not nec emotionally attached, just unbalanced reinforcement of thought patterns)] -> [become (in said field) a comically slow updater who (unintentionally?) tries to cram everything observed into the rigid framing of prior views]
intelligence can work against you in these situations if you don't pilot it very carefully, gives you the tools to retroactively craft coherent (to you) arguments that fit your framing. once in this phase, much easier to do this than to consider alternatives on the same plane
not really "age effect" though correlated with how long one has spent in the field and how rapidly the field is changing
this makes some sense if you're talking about the marginal net-positive impact of one person rather than the HFT industry as a whole (which can be argued to benefit society quite a bit: tighter spreads, more liquidity, faster/more accurate price discovery, less risk for retail traders etc)
but this "marginal" argument also leads to strange conclusions like "being a doctor in a major hospital system is negligibly net-positive" + the jury is still out as to whether the present ai boom (closed-source, centralized harbingers in particular) will do much besides hyper-accelerate technocapitalism
a call to ethics/meaning definitely misses the mark as an openai recruitment strategy in 2025, "marketing" your most interesting/difficult engineering problems is all you need to do if the comp is there
@ZhongRuiqi it’s just that pushing the upper end of capabilities is much slower (and spikier) than pushing efficiency and catching up, so the resource advantages are often not visible
they’re visible in the sense that you’d expect more {chatgpt moment, o1 moment} from openai than deepseek