Ramadan almost ends but our relationship with the Qur’an shouldn’t.
Let’s build habit that lasts beyond Ramadan.
Start small. Just 3 ayahs a day.
Try my new app Steady Quran for iOS:
https://t.co/06dNL3rWCZ
“Steady” is an uncommon English word for people in Indonesia. I also worry it might be confused with “Study Quran”.
On the other hand, everyone knows the word “life” (I hope!)
https://t.co/OE5ccOTq2T
Changed my app name from “Steady Quran” to “Quran Life”. It’s the exact same app, just with a better name (I hope!)
The thing with the old name was that I had problem myself pronouncing it, and it makes it hard to tell people around me about it.
@cline This might finally push me to do multi-agent properly, I have been struggling to find a decent way to manage them that doesn’t overwhelm my brain.
Might be useful to someone: to configure Google Analytics without IDFA in your mobile apps thru Firebase, its quick setup guide will mention using `FirebaseAnalyticsWithoutAdId`. That library doesn't exist anymore. Use `FirebaseAnalyticsCore` instead.
Interesting and new perspective to me. I had seen anxiety as something that appears and disappears from within me, not something that I can control/increase/decrease. I love it.
The Prophet ﷺ came across ‘Abdullah ibn Mas’ud whilst he was sad, and said to him:
“Do not increase your anxiety. That which is decreed occurs and that which is provided shall come to you.”
📚 Shu’ab al Iman 1188
Interesting part about dead code and commented out blocks specifically, because in my experience Claude Code noticed them and understood enough not to take them into account. I agree those ideally should be removed anyway, though, I just have commitment issues 😌
𝗟𝗟𝗠𝘀 𝗔𝗿𝗲 𝗡𝗼𝘁 𝗥𝗲𝗮𝗱𝗶𝗻𝗴 𝗬𝗼𝘂𝗿 𝗖𝗼𝗱𝗲
We keep calling LLMs "AI coding assistants." But writing code and understanding code are not the same thing. Researchers from Virginia Tech and Carnegie Mellon University just ran 750,000 debugging experiments across 10 models to determine how well LLMs actually understand code.
The results show that you should not blindly trust your AI coding assistant when debugging.
Here is what they found:
𝟭. 𝗔 𝗿𝗲𝗻𝗮𝗺𝗲𝗱 𝘃𝗮𝗿𝗶𝗮𝗯𝗹𝗲 𝗯𝗿𝗲𝗮𝗸𝘀 𝘁𝗵𝗲 𝗱𝗲𝗯𝘂𝗴𝗴𝗲𝗿
Researchers created a bug, confirmed that the LLM found it, then made changes that don't touch the bug at all, such as renaming a variable or adding a comment. In 78% of cases, the model could no longer find the same bug. The bug was still there. The variable names and comments changed, and that was enough.
𝟮. 𝗗𝗲𝗮𝗱 𝗰𝗼𝗱𝗲 𝗶𝘀 𝗮 𝘁𝗿𝗮𝗽
Adding code that never runs reduced bug-detection accuracy to 20.38%. Models treated dead code as live, and flagged it as the source of the bug. But the bug was in another line. So, LLMs cannot reliably distinguish "this runs" from "this never runs."
𝟯. 𝗠𝗼𝗱𝗲𝗹𝘀 𝗿𝗲𝗮𝗱 𝘁𝗼𝗽-𝘁𝗼-𝗯𝗼𝘁𝘁𝗼𝗺, 𝗻𝗼𝘁 𝗹𝗼𝗴𝗶𝗰𝗮𝗹𝗹𝘆
56% of correctly found bugs were in the first quarter of the file. Only 6% were in the last quarter. The further down the code, the less attention the model pays to it. If the bug lives in the bottom half of your file, the model is already less likely to find it.
𝟰. 𝗙𝘂𝗻𝗰𝘁𝗶𝗼𝗻 𝗿𝗲𝗼𝗿𝗱𝗲𝗿𝗶𝗻𝗴 𝗮𝗹𝗼𝗻𝗲 𝗰𝘂𝘁 𝗮𝗰𝗰𝘂𝗿𝗮𝗰𝘆 𝗯𝘆 𝟴𝟯%
Changing the order of functions in a Java file caused an 83% drop in debugging accuracy. The code still remained the same. Where the code physically sits in the file matters more to the model than what the code does. So, obviously, this is a sign of pattern recognition, not real code understanding.
𝟱. 𝗡𝗲𝘄𝗲𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝗵𝗮𝗿𝗱𝗹𝘆 𝗺𝗼𝘃𝗲 𝘁𝗵𝗲 𝗻𝗲𝗲𝗱𝗹𝗲
Claude improved ~1% between 3.7 and 4.5 Sonnet on this task. Gemini improved by ~1.8%. Every model release comes with a new benchmark leaderboard and new headlines. But the ability to reason about code under realistic conditions is improving slowly.
𝟲. 𝗧𝗵𝗲𝘀𝗲 𝘄𝗲𝗿𝗲 𝗯𝗲𝘀𝘁-𝗰𝗮𝘀𝗲 𝗰𝗼𝗻𝗱𝗶𝘁𝗶𝗼𝗻𝘀
The study used single-file programs with ~250 lines, and each had a clear description of what the code should do. The authors say this was intentional. They wanted the best-case conditions. Real production code is multi-file, cross-module, and poorly documented. It will perform worse for sure.
Here are three things worth changing based on the research:
🔹 𝗣𝗮𝘀𝘀 𝗲𝘅𝗲𝗰𝘂𝘁𝗶𝗼𝗻 𝗰𝗼𝗻𝘁𝗲𝘅𝘁, 𝗻𝗼𝘁 𝗷𝘂𝘀𝘁 𝗰𝗼𝗱𝗲. When asking an LLM to debug, include test output, stack traces, and failure messages alongside the source. Without runtime details, the model is guessing based on the code.
🔹 𝗗𝗼𝗻'𝘁 𝘁𝗿𝘂𝘀𝘁 𝗶𝘁 𝗼𝗻 𝗱𝗲𝗲𝗽-𝗳𝗶𝗹𝗲 𝗯𝘂𝗴𝘀. If the suspect code is in the bottom third of a long file, the model will have trouble finding it. Consider splitting the context or feeding the relevant function directly.
🔹 𝗖𝗹𝗲𝗮𝗻 𝘂𝗽 𝗱𝗲𝗮𝗱 𝗰𝗼𝗱𝗲 𝗯𝗲𝗳𝗼𝗿𝗲 𝘂𝘀𝗶𝗻𝗴 𝗔𝗜 𝗱𝗲𝗯𝘂𝗴𝗴𝗶𝗻𝗴 𝘁𝗼𝗼𝗹𝘀. Commented-out blocks and unreachable branches will mislead the model. It cannot filter them out.
We rate AI coding tools on HumanEval. That tests whether a model can write a function from a description, but this says nothing about finding a bug in code it didn't write.
Those are different problems. We're using the wrong benchmark.
Fuck it, we just made Fizzy completely free.
The open source installable version was always free, but the SaaS version was pay. No more. Basecamp and HEY's largess will subsidize Fizzy for all.
So go grab your account at https://t.co/pGxSDeGKNK. It's Kanban the way it should be, not the way it has been. Fresh, fun, light, fast, and perfect for working with agents, too.
An official CLI is coming soon as well. Stay tuned for that.
The native iOS app should be out once Apple approves it (it's in approval right now...). Android is already out, you can get it on the Play store.
(and BTW if you were a paying customer, you will no longer be charged moving forward)
One of the most surprising "habits" I see Claude Code does is making all these tiny python tools to help it do a task. Converting one text format to another, image manipulation, dowloading resources, etc. It's pretty cool and small tools are easier to verify.
I've been mostly shitting on AI these days (also sick of the "built 15 apps in 15 minutes" crowd), but Claud just did something wild for me, a project that would otherwise take MONTHS.
Had it help me rip out our entire search engine (we're talking millions and millions of records) and migrate it from SQL "full-text search" to a small embedded, in-process Lucene port (Lucene is what powers Elasticsearch under the hood).
Our app has thousands of tenants with millions of tickets - the search went from 7-8 seconds down to fucking milliseconds.
The rewrite was the easy part though. The real work was all the one-off CLI tooling for index rebuilding, compaction, deduplication, gradual deploy... Literally dozens of tools. Also prep-work, tech-stack deciding, risk analysis, benchmarking, planning...
The reindexing job itself ran for 39 HOURS STRAIGHT, reporting nice progress graphs and auto-fixing errors as it went. It just finished and eberything checks out
not gonna turn into one of those AI evangelists, but claude saved me weeks of the most tedious infrastructure grind imaginable. And I've been trying to approach this project for literally YEARS.
@SebastianRoehl Yup, same here. Are you a visual learner? For me seeing ideas made real on the screen seem to help my brain see connections and possibilities, something that doesn't happen if I just sit around and think about things.
@Shpigford It’s the edge cases that kill. This is also why WordPress is still very popular today. All the accumulated knowledge and bug fixes over the years makes it more battle-tested and not something any new CMS can match, AI-assisted or not.
Now that analytics appear, I still see no crash reported both in Crashlytics and Apple Developer for these. So when an app is killed due to a memory spike, there is no reporting. You gotta check manually by seeing logs in the device and profiling.
I used SVG images (purely decorative) for my iOS app only to find that various manipulations on it—like sizing and recoloring—caused a memory spike up to 2GB (!). Converted to PNG and now it’s at 50~MB. This was maybe what caused the crashes on older devices…
@jasonfried AI helps me a lot to ship faster, but it’s still far from a magical button that will let most people to create their own software. There are still a lot of pain involved. I am more excited about AI helping people who are interested in making software (my kids!) to get in faster.