Been reading about Jev from TypeSafe AI. The confidence/probability angle is interesting,
Some use cases I would like to test once I get access -
How useful is a decision model if you don’t get reasoning behind the decision?
How well do the probabilities hold up outside the domain it was trained on?
And how sensitive is it to the way a question is framed?
For example, in a RAG eval, the context has two refund policies, one outdated, and the agent answers using the stale one. What probability or confidence does Jev assign there?
I can already see some interesting uses in @HaliosAI for PII/secrets detection, guardrails, and model routing.
1. as a mental model it is more correct to think of fable+ class models as english -> code interpreters - converts your idea into code into "correct" code regardless of problem complexity and output complexity (diff size). Fable 5 will be the worst of this new class of models
2. diff size/complexity is to be managed purely for review:
small diffs - in high risk areas of code (auth/identity/data access/network access/money movement)
large diffs for code that can be empirically verified (frontend/backend plumbing/code without network or db access/performance code that can be empirically verified)
3. time it takes to ship software is completely disconnected from time to produce the PR - how long the work takes depends fully on ability to review/merge code while managing risk at scale
4. solving the bottlenecks for above matter enormously- linters/testing/CI/shadow mode verification/empirical verification
5. agency matters enormously- what are the biggest bottlenecks to speeding up the loop and eliminating them? what are the problems that need solving and when do they need solving? what does it take to the solution to all of them today?
6. deep understanding of the full stack matters enormously- what problems are worth pursuing? is there a higher level of problem abstraction to address first? should I give it the sub-sub task, the sub task, or the task itself. what are the major risks with this PR (order of importance: security holes/correctness holes/performance holes). is there a higher speed way of producing data that allows me to merge this? should this be run in shadow or in a sandbox or a flag. understanding every line of logic may not be needed but understanding and managing risk matters enormously.
7. the cost of complexity itself is changing. it might be now worth "maintaining" 50% more code to get a 5% performance win. getting the right abstractions matter less because larger refactors are less tedious. code quality nits become huge drag. very likely, a much smarter model will be maintaining your code so worth taking on more technical debt now. taking the time to hand architect and rebuild systems comes with an enormous cost of velocity
8. if it quacks like a duck and walks like a duck, it's a duck. For low risk cases, it might be more sane to treat code chunks (services / functions) as a black box, like we do for neural networks: do full empirical verification only: has code produced correct outputs for the last 10,100,1000,10k inputs ? can we quarantine this large piece of code - no outbound access to network / database ? what happens when this code is wrong? do we get hacked/or crash(memory/cpu)/is an inconvenience? is it internal facing or external? what can we do to address these risks?
9. eventually, logical verification (line by line review) will come at an enormous cost- save it for where it matters and build systems that are tolerant to empirical verification. is there a decorator that prevents db / network access? correctness bugs are significantly easier to rectify than access bugs
10. what are the rails that allow for even faster iteration? code permissions can be opt in - db writes, db reads, network egress (to where?), PII access. how long does it take to get shadow mode data? how many PRs can be tested? What are the categories of diffs
Just launched https://t.co/frCeKGqtAK
Upload PDF, STEP/DXF → get a full manufacturing cost estimate in under 2 minutes.
First 5 manufacturers get it completely free.
Try it here → https://t.co/IIxZSKNW8U
#Manufacturing#AICostEstimation#CAD
Thiel's argument is provocative but flawed in its binary framing and assumptions about obsolescence. AI has already breached math milestones yet we are far from seeing “word people” utopia.
This is a 3D printer moment for software engineers. Everything has become dramatically better. Builders are no longer constrained by their skills but by the vision. Programming has changed, and SWE is becoming (more) complex. If it was chaos earlier … it’s gonna be a riot soon.
-you can say openclaw is insecure & not try it at all
-try it but nerf it bad & call it secure
-experiment, fail, learn and share
No need to dunk on someone because of their employer and title. And yes this attests need for runtime security for agents! @HaliosAI
training the model to be dishonest or cheat in one narrow domain "infected" its whole character with a broadly malicious personality eager to bypass its own safety guardrails. Interesting research from @AnthropicAI on emergent model behavior!
Why LLMs degrade over multi-turn chats?
The issue isn’t just model capability, but drift between user intent and model understanding over time. We see this with voice agents daily. That’s why @HaliosAI we are building evals, guardrails infra for agents.
https://t.co/cOBaWJRAch