Been reworking our services site with GPT-6 Astra in Codex... pretty damn pleased with this. Still had to show it what good looks like and get specific with the design. Even asked it to put this showcase together, music and all.
@melvindvivas Don't spam Astra. Just use it for scoping, focused high risk reviews and parts of the codebase that need elevated judgement. Luna xhigh is fine for most day to day execution
@dev_K95@Jaytel Xhigh is best in my experience- marginal cost savings changing out to lower reasoning with degradation in quality, so I keep it on xhigh.
@Theluckyjha I think the coding knowledge still matters though... especially when it gives you something that looks right but isn't. Knowing what to question is a pretty useful skill, even if you're writing less of the code yourself.
@boristane Slack as the place to start a job makes sense... if you can also steer it and review the changes there. Having to keep jumping back into another app rather defeats the point.
@arb5z@kloss_xyz The separate review makes sense... especially when the plan changes halfway through. Easy for the next agent to keep working from the original brief while the code has gone somewhere else.
@charnpreet89@zcoderrr@v_tilneac Getting a clean bill of health, then finding vulnerabilities with the next model... doesn't inspire much confidence in that first review. Curious whether it actually tested anything or just read through the code?
@FelipeFr1702 41% since Tibos last reset.... leaning on Luna a lot though which tbh is pretty good for most things, and step up to Astra light for some hefty work.
@lopezunwired If this ships, not having to keep uploading statements would save some hassle... the memory needs updating too though. A salary or rent figure from six months ago could quietly throw off every answer after it.
@SoFi The plain-English setup is appealing... though it's easy to keep tweaking until the historical chart looks great. Trying the same rules on a period you haven't used to build them could tell a different story.
@0xZenad Getting them to stop stepping on each other would make a big difference... especially once you count the time spent reviewing and fixing their changes. Two with separate jobs could get you further than five needing constant untangling.
@allen_lattimer Annual bills are awkward with these... a car repair or insurance payment can look like overspending when it's just landed that month. Looking back a little further could stop the budget treating those as money you can cut.
@NickAbraham12 Having this updated weekly makes sense... could be worth running a version where a couple of clients pay late too. The profit forecast might still look good, but seeing whether the cash arrives before payroll would be handy.
Using AI for UI design... 'make it better' keeps getting me stuff I don't want.
For this checkout concept: show the full price without scrolling, put Continue beside it, keep the rest as it is.
Then check it on mobile too. Easy to miss when the desktop looks good.
@N01ennn Yeah, an agent saying it's finished isn't enough to hand its answer to the next one... especially if they're all building on the same assumption. Checking the result before it gets passed on saves the others doing a whole lot of work on something that's wrong.
@UniofOxford@OxUniMaths I think this carries over to using AI for learning too... an explanation can make perfect sense while you're reading it, then you change one assumption and you're lost. Working through an example together, then trying another without it, shows whether you've understood it.
@wholemars Yeah, you still have to spell out what working means... same task, same inputs, same result on both apps. Annoying having to be that specific when you've already given it the working version, but it needs to check the actual behaviour on the phone too.