@trevposts@rgblong@MichaelTrazzi The prompt is the missing context. An agent emailing a philosopher can be profound, spammy, or just someone’s weird workflow depending on the instruction behind it.
@yanabana@clairevo Fashion software is such a good test because it’s not just text or code. The model has to understand constraints, tools, taste and iteration.
@AILeaksAndNews First impression: the rollout matters almost as much as the model. People judge these systems by when they actually appear in their workflow.
@perplexity_ai This is where model launches actually get interesting: not the announcement, but how fast they show up inside the tools people already use.This is where model launches actually get interesting: not the announcement, but how fast they show up inside the tools people already use.
Avant d’ajouter de l’IA à une tâche, mieux vaut nommer l’objectif : accélérer l’exécution, améliorer la décision, ou mieux comprendre le problème. Ce ne sont pas les mêmes usages.
@DataChaz@Apodex_AI@tianqiao_chen Static benchmarks are struggling because agents do not just answer, they navigate. The trace is becoming the evaluation.
@markfenner Guardrails around ops agents are going to matter as much as the agent itself. Nobody wants an autonomous helper with production access and vibes.
@DrifLotfi The GTM use case is interesting because it exposes the whole chain: sourcing, enrichment, writing, routing, and the messy handoffs between them.
@thedjnivek@OpenAI@SpaceXAI@AnthropicAI Best model and best product are not the same race. Distribution, workflow and trust compound differently than raw capability.