introducing the @tastelabs research fellowship!
for those excited to work on evaluating model creativity, distilling human preference, and ending slop.
you'll get compute, custom data, time with top designers and researchers, and a competitive stipend. in person in SF or remote, full time or part time.
check out some of the research questions you could explore and apply here:
https://t.co/DiAgEOXcOm
seeking AI engineers who want to end AI slop. if you:
- obsess over evals & harnesses
- love data
- want to help push frontier models
- have unreasonably good taste
- want AI to ~feel~ right
- love design & creativity
… this is for you!
apply at https://t.co/IAeAFxPHCE
couldn’t agree more.
if you want to research data & more on how to make unverifiable domains verifiable, we’re hiring:
- research engineers
- design engineer researchers
New in Claude Code: your sessions can now message each other.
Instead of having to re-explain yourself in another session, you can now tell Claude to do it. It sends a summary (not your history or files), and the other session picks it up mid-task.
@francocontigo@ktakanopy@rafaelbarbosa_s estou usando o graphify aliás! confesso que não percebi a melhoria, MAAASS, pelo conceito e arquitetura eu acredito muito no sistema.
Been running @_fonsecabc's Tars as my second brain for a few weeks now. fully local, open source, and it remembers facts from meetings i never even attended.
And the counterintuitive claim holds: honesty dial up, fewer hallucinations. best memory layer i've used with Claude.
@slash1sol filtering at write time is interesting
i went the other way, store everything and let a nightly pass merge dupes and let trivia fade.
no idea yet which is right honestly
what does their decay look like
@samsja19 persistent kernel across compaction is such a good detail
been chasing the same thing one layer out, keeping memory alive past the whole session instead of just the compaction
gonna try this
@Teknium the timeline view is sick
does it keep what was true at the time, or does the graph update in place? i went append only and kept the superseded facts, mostly because i kept wanting to know what i believed two months ago
curious which way you went
We’re excited to introduce Taste Labs.
Our mission is to end AI slop. We’re building the data and infrastructure layer to give AI models and agents taste.
And today we’re coming out of stealth, announcing our $18.5M seed funding, co-led by @CRV and @AmplifyPartners
AI has nailed objective domains and made it easy to generate anything. But it still feels off. Now, the challenge is judgement. What fits, what feels like you, what’s GREAT. This requires turning a fuzzy, subjective domain into something we can measure and codify. We’re starting with design.
There are two sides to cracking this, the foundation model layer and the agent layer:
- We’ve already been working with the top frontier labs to evaluate and improve their models, crafting the right post-training data and RL environments.
- We’ve also been working with app-layer companies to build the context and verification tools for their agents to produce better, more on-brand, more creative outputs.
We want a future where AI feels right.
If you’re passionate about this mission, join us!
it ships empty.
no data, no assumptions about whose life it is
MIT licensed, one command to run
repo: https://t.co/XeoFIayQ9p
if you build with agents i think this helps whatever harness you use
and i would love ur ideas as PRs, i can only go so far alone
i gave claude a real memory, then gave it a personality
the personality made it hallucinate less
took me a while to work out why.
checkout the whole build is in this thread
caveat i won't bury: a 14b local answerer caps accuracy.
these numbers compare tars against itself, not against leaderboard scores judged by frontier models
the hit rate is the model independent one and that's what i optimize
checkout my article on the full system: https://t.co/6cKk8qQNZx