Pointer is hiring!
We're growing faster than we can handle and need people across every function.
We build AI systems that do a company's hardest work and prove it was done right, deployed in some of the highest-stakes operations in the world.
If this sounds interesting, learn more at https://t.co/xyf2mkkHdL. Refer someone we hire and we'll pay you $15,000.
We at @pointer hosted a dinner with @stepstonegroup at Le Pavillon in NYC last week for a group of finance leaders in banking, insurance, fund admin, and more.
It's clear that nobody wants AI they have to babysit. The systems delivering the most value are the ones that can remove the human reviewer entirely and run in the background for hours.
More soon!
locked in the trenches coding the 2 most requested features for the roommate app:
undo button for when your fat fingers accidentally skip a match (the animation is so smooth)
“see more” button to read their full bio
link in bio 😉
random roommate matching is literally russian roulette lmfao.
we looked at the data on what actually destroys shared housing. the results are honestly kind of wild.
if u want to avoid a 12-month lease from hell, look at these numbers:
A lot of people have been asking about our harness / approach - some thoughts:
1/ it’s fully open source on github!
2/ it is quite simple - and we think this is where harness engineering is heading. you no longer need elaborate scaffolding to force the model to reason in a prescribed way
3/ we initially included a verifier to check the executor’s work. it ended up being *more* accurate than the benchmark’s grader, but omitted it (you can't score above the ceiling set by the grader). we have a lot more to say on this.
4/ we were most excited by the performance uplift in sonnet (lighter model). it reflects a shift toward picking the model at the intelligence/cost pareto max for a task, not just the largest one. sonnet achieved near parity with opus in performance, while costing less than half.
Today, we’re sharing a new state of the art for computer use.
Our system holds the two highest verified scores on OSWorld, the standard benchmark for AI agents that operate a computer like a person: 83.6% using Claude Opus 4.7 and 81.5% using Claude Sonnet 4.6. The human baseline is 72.4%.
🧵 1/7