Building Albus, to help you parent with intention. Previous: early hire @relyanceai, @medallia, @pwc, @oracle. Like π ππ½ββοΈ π¨π½βπΌ β°οΈ
@lil_dill Great talk and completely agree! Software should feel delightful and have a soul.
I just redid my entire website because of this thought. How can I communicate the story with deligh and soul behind it.
https://t.co/MweLjjoHs9
Say hello to the new Albus website! Very happy with how this came together. For weeks, I was trying to figure out how to tell the story. It clicked 3 days ago, and with Asta + Opus 5.5, it's reality now.
https://t.co/LiQT2V5ns2
Say hello to the new Albus website! Very happy with how this came together. For weeks, I was trying to figure out how to tell the story. It clicked 3 days ago, and with Asta + Opus 5.5, it's reality now.
https://t.co/LiQT2V5ns2
@synthwavedd Opus 5.5 is truly incredible. Though Astra is still great at design aesthetics IMO. Consistently outperforms in translating feelings into compelling UX.
@synthwavedd Opus 5.5 is truly incredible. Though Astra is still great at design aesthetics IMO. Consistently outperforms in translating feelings into compelling UX.
Opus 5.5 and GPT 6 both dropped today, and I am really starting to feel how powerful these models are becoming. The tell for me is that my own restraint is becoming the key design skill with these newer models.
A year back, my default operating model was to make my prompt files "heavy". Long specific explanations, guardrails, few-shot examples, the works. The rationale was reasonable at the time: the model can cook, but it is also a confident liar, so keep the handcuffs on, and you get some semblance of predictability in the output.
Now, with this new class of models, the handcuffs are actually constraining them and stripping away their superpower. They make the models obedient and docile instead of letting them cook.
So my advice now is to audit your key prompts and rewrite them from scratch. Keep the product ethos in there, but kill the examples, kill the guardrails you can, and remove the conflicting rules. And use good old deterministic enforcement wherever your code can do the job.
I am in the middle of doing this to the whole Albus report architecture. The pipeline running today uses 25 separate prompts to produce one report. My new one uses 8, with 45% less prompt text.
@wminshew oh man, I feel this. also seeing your child the first thing in the morning is so much fun! their smile, the energy when they get up. wish that could last forever.
Opus 5.5 and GPT 6 both dropped today, and I am really starting to feel how powerful these models are becoming. The tell for me is that my own restraint is becoming the key design skill with these newer models.
A year back, my default operating model was to make my prompt files "heavy". Long specific explanations, guardrails, few-shot examples, the works. The rationale was reasonable at the time: the model can cook, but it is also a confident liar, so keep the handcuffs on, and you get some semblance of predictability in the output.
Now, with this new class of models, the handcuffs are actually constraining them and stripping away their superpower. They make the models obedient and docile instead of letting them cook.
So my advice now is to audit your key prompts and rewrite them from scratch. Keep the product ethos in there, but kill the examples, kill the guardrails you can, and remove the conflicting rules. And use good old deterministic enforcement wherever your code can do the job.
I am in the middle of doing this to the whole Albus report architecture. The pipeline running today uses 25 separate prompts to produce one report. My new one uses 8, with 45% less prompt text.
I've been thinking about what a "regression" means in products where the core workflow is predicated on LLM reasoning and judgment (like Albus). It's not usually something failing outright. I am finding in real time that, more and more, it just shows up as quality slowly eroding.
Part of it is that probabilistic reasoning has so many levers, and they're hard to control.
- Model drift is real. Something that worked fine 2 months back starts degrading silently. Opus 5 comes to mind; I had to yank it away from some of the prime report writing in my product because it just kept getting worse.
- You switch models, and suddenly your old prompts are somewhat useless.
- Bad evals.
- Surgical fixes and rules that I add over time end up jeopardizing the output. This is the one I'm actually grappling with right now. As I try to improve my reports, we tweak prompts, find bugs, tweak again, and some of these band-aid fixes pile up until the model gets overly compliant and loses its ability to be creative and actually reason.
I'm really indexing on radically simplifying my architecture. Genuinely curious how other folks are dealing with this?
This is a really good read. I strongly agree with a lot of things you said. One of the big things I follow which@richroll said at some point, is "Mood follows action." The story you tell yourself, how you think, drives real-world action, and there is a lot of power in realizing that.
It's funny that I was thinking about the stories our mind creates from a different angle: The voice in your head knows you and is saying something, and it's worth listening.
https://t.co/e3ZPnwgBUB
I have a problem-discovery skill with Sol that I absolutely love to get to the heart of the problem.
I also have a /spec loop that takes a problem brief + competitive research + metrics + red teaming with a growth subagent / codex adversarial review and then Fable to give me a complete spec.
@iam_preethi This one was great! It basically converts the seat into an almost bed for the kiddo. The other thing that worked was to buy a seat for the toddler. That helped a ton!
Inflatable Travel Foot Rest... https://t.co/HmAZP1AOk5
@iam_preethi Stickers, berries, chips (my son like holding a chip in each hand), busy board, getting one of those inflatable things that make the seat a bed, lollipops (messy but useful!).