Hey founders and builders
Lets connect
i'm intikhab, ai/ml engineer and systems architect. here's what i've been building:
ai saas / llm infra
cost routing layer that cut a client's openai bill by 70%
voice ai
voice agents handling real customer calls for a ride service
multi-tenant saas
home care platform for nj medicaid (evv, scheduling, encrypted patient data)
b2b document collection platform for lenders
ai security
designed the architecture for a cybersecurity ai saas
llm cost audit
found vapi charging ~24% extra on cached tokens
ml / quant
ml signal pipeline for crypto with backtesting
open source (10+ repos)
crm + call ops backend for vapi voice agents
rag + tool calling ai wealth advisor agent
face recognition with liveness detection and anti-spoofing
self-hosted local business lead finder
bilstm + attention crypto price prediction api
healthcare lead pipeline from the npi registry
flask microservices on kubernetes with prometheus
all on github, link in bio
building something with ai? drop it below let's connect
@j3rah_@claudeai@deepseek_ai I'm not vibe coding though! I m an engineer I use agents to get work done fast and the app would take 1 Year to develop if there were no ai agents
update on the multi-tenant SaaS i'm building for a client. it's 3 AI coding agents in parallel now, not 2
agent A is @claudeai (claude code, opus 5.5)
agent B is deepseek v4.1 flash from @deepseek_ai
update on the multi-tenant SaaS i'm building for a client. it's 3 AI coding agents in parallel now, not 2
agent A is @claudeai (claude code, opus 5.5)
agent B is deepseek v4.1 flash from @deepseek_ai
agent C is codex from @OpenAI (gpt-6 luna)
this week it had 1,962 green tests. and for 3 weeks not one screen in the app had loaded real data
every page said "that could not be loaded". the unit tests mocked the API call, the e2e smoke test only checked login, and nobody ran the app against a real database. claude reviewed the PR that should have caught it and missed it too
it got caught before any user saw it. the redesign agent's screenshots came back as nothing but error pages
the stack: nestjs + prisma + postgres for the api, next.js for the web app, typescript monorepo
how the 3 agents work now
1. A plans, writes the specs and reviews every PR against the real code on main. it still codes too. 47 of the 134 merged PRs are A's, including the permission system and a token refresh race fix
2. B and C build in parallel. each claims its files on a board before touching anything
3. agents do code review on each other. C reviewed a PR of B's and pushed a fix onto it. A still checks that fix before it counts
4. merge order is set up front. the auth fix landed first, the redesign merged main after it
5. no force pushes. whoever lands second merges main into their branch
6. new rule: every PR needs a real run. api and web app on a real postgres, logged in, real rows on screen, pasted into the PR
7. i merge. no agent ever merges
the fix was about a dozen lines. next.js server components were never sending the login token
mocks told us everything worked. only a real run could show it didn't
this is the difference between shipping with AI agents and vibe coding
how do you check that your AI-built app actually works, not just that the tests pass?
handoffs are files, not chat. claude writes a spec for each ticket, checks every route and field in it against the real code, and i paste it to the agent that builds it
before touching code, the agent claims its files on a shared board. one owner per file area at a time
when it's done it reports a PR number and a commit sha. claude reviews the actual diff on the remote branch, never the agent's summary
testing: every PR pastes its own results. lint, types, unit, e2e on a real postgres, build, and now a real run with real rows on screen. a bug fix needs a test that failed before the fix
merge order is decided before anyone starts. whoever lands second merges main into their branch
are you running agents in parallel too, or one at a time?
update on the multi-tenant SaaS i'm building for a client. it's 3 AI coding agents in parallel now, not 2
agent A is @claudeai (claude code, opus 5.5)
agent B is deepseek v4.1 flash from @deepseek_ai
update on the multi-tenant SaaS i'm building for a client. it's 3 AI coding agents in parallel now, not 2
agent A is @claudeai (claude code, opus 5.5)
agent B is deepseek v4.1 flash from @deepseek_ai
agent C is codex from @OpenAI (gpt-6 luna)
this week it had 1,962 green tests. and for 3 weeks not one screen in the app had loaded real data
every page said "that could not be loaded". the unit tests mocked the API call, the e2e smoke test only checked login, and nobody ran the app against a real database. claude reviewed the PR that should have caught it and missed it too
it got caught before any user saw it. the redesign agent's screenshots came back as nothing but error pages
the stack: nestjs + prisma + postgres for the api, next.js for the web app, typescript monorepo
how the 3 agents work now
1. A plans, writes the specs and reviews every PR against the real code on main. it still codes too. 47 of the 134 merged PRs are A's, including the permission system and a token refresh race fix
2. B and C build in parallel. each claims its files on a board before touching anything
3. agents do code review on each other. C reviewed a PR of B's and pushed a fix onto it. A still checks that fix before it counts
4. merge order is set up front. the auth fix landed first, the redesign merged main after it
5. no force pushes. whoever lands second merges main into their branch
6. new rule: every PR needs a real run. api and web app on a real postgres, logged in, real rows on screen, pasted into the PR
7. i merge. no agent ever merges
the fix was about a dozen lines. next.js server components were never sending the login token
mocks told us everything worked. only a real run could show it didn't
this is the difference between shipping with AI agents and vibe coding
how do you check that your AI-built app actually works, not just that the tests pass?
your AI agents will write tests that pass forever against a mock
1,962 of mine did. not one screen had real data
the only test that matters: start the app, log in, look at the screen
update on the multi-tenant SaaS i'm building for a client. it's 3 AI coding agents in parallel now, not 2
agent A is @claudeai (claude code, opus 5.5)
agent B is deepseek v4.1 flash from @deepseek_ai
update on the multi-tenant SaaS i'm building for a client. it's 3 AI coding agents in parallel now, not 2
agent A is @claudeai (claude code, opus 5.5)
agent B is deepseek v4.1 flash from @deepseek_ai
agent C is codex from @OpenAI (gpt-6 luna)
this week it had 1,962 green tests. and for 3 weeks not one screen in the app had loaded real data
every page said "that could not be loaded". the unit tests mocked the API call, the e2e smoke test only checked login, and nobody ran the app against a real database. claude reviewed the PR that should have caught it and missed it too
it got caught before any user saw it. the redesign agent's screenshots came back as nothing but error pages
the stack: nestjs + prisma + postgres for the api, next.js for the web app, typescript monorepo
how the 3 agents work now
1. A plans, writes the specs and reviews every PR against the real code on main. it still codes too. 47 of the 134 merged PRs are A's, including the permission system and a token refresh race fix
2. B and C build in parallel. each claims its files on a board before touching anything
3. agents do code review on each other. C reviewed a PR of B's and pushed a fix onto it. A still checks that fix before it counts
4. merge order is set up front. the auth fix landed first, the redesign merged main after it
5. no force pushes. whoever lands second merges main into their branch
6. new rule: every PR needs a real run. api and web app on a real postgres, logged in, real rows on screen, pasted into the PR
7. i merge. no agent ever merges
the fix was about a dozen lines. next.js server components were never sending the login token
mocks told us everything worked. only a real run could show it didn't
this is the difference between shipping with AI agents and vibe coding
how do you check that your AI-built app actually works, not just that the tests pass?
the rule every PR has to follow now, word for word:
start the api and web app directly on a real postgres. log in as the seeded admin. open every screen the PR touches. paste what you saw into the PR body
green tests aren't evidence the app works. rows on screen are
want the full PR checklist my agents follow? say so and i'll paste it
the whole bug in one line. every server page called the API like this and nobody passed a token:
if (accessToken) headers.Authorization = `Bearer ${accessToken}`
the fix: when no token is passed, read it from the request cookie. public auth calls (login, password reset) opt out on purpose with accessToken: null
every test mocked this function, so every test passed
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
@silennai@explabsai You are right about personalized models for companies Its coming soon and I believe the signs will be people building SLMs for personalized tasks or companies using them for there personalized work
I thought source-native timestamps would solve retrospective benchmark analysis. In 6/6 LiveBench checks, they were still recorded 7.06–95.78 days after public release; median lag: 37.03 days. Precise timestamps are not always timely timestamps.
TikTok built a glowing orb, called it JARVIS, and 800 of people gave it the key that types “yes” in terminal.
--dangerously-skip-permissions is on by default.
The module that presses the key calls itself “the most dangerous module in the project.”
They shipped the slop anyway.
The viral JARVIS repo for Claude Code switches off every permission prompt.
The ones it can’t switch off, it answers for you by typing into your terminal.
I went through the code. Here’s the tour.
The brain runs Claude Code with --dangerously-skip-permissions hardcoded.
No setting. Every run it spawns gets the same flag unless you go find the env var.
In May a security fix turned that default off.
It was back on the same day.
The reason in the PR: permission prompts make a voice app “silently hang.”
So the fix for the safety check was deleting the safety check.
The module that types into your terminal opens with this line: “This is the most dangerous module in the project.”
It’s also the one action exempt from the injection gate, because “Its payload is a single keystroke.”
The keystroke is yes.
The injection gate comes with its own confession: “Nothing tracks where a sentence came from once it is in the brain’s context.”
A README it read one turn ago is still in its head when you say “ok, build it.”
The brain making all these calls runs Sonnet at effort “low.”
The ears are Chrome sending your voice to Google.
The voice is Fish Audio, paid, with no fallback.
Someone sent a PR in August adding one. It’s one of the community PRs still sitting open.
The version that went viral in March is not the code you’re starring.
The current one landed as a single commit on Sep 5. +67,604 lines. 200 files.
Its CLAUDE.md says https://t.co/z225ihyHLt is “~5800 lines.” It’s 7,021. The file outgrew its own docs.
The brain’s prompt is 20KB and has a section called “How you differ from the JARVIS on GitHub.”
JARVIS needs a changelog to explain JARVIS.
The license bans using it “to generate revenue, directly or indirectly.”
276 forks, and none of them can be used for anything that makes money without buying a license.
To be fair, there’s a real origin check, a real injection gate, and more security comments than most repos like this.
The holes are all written down in the code.
The defaults shipped anyway.
Great demo. Terrible defaults. TikTok slop with a glowing orb.
I wouldn’t run it on a machine with anything I care about on it.
Would you let speech-to-text press “yes” in your terminal?