Three different people sent me the same link this week, which is usually a decent signal that something is either genuinely interesting or being very well marketed. In this case I think it’s a bit of both.
The thing is called Jev, from a startup called TypeSafe AI. It went generally available on Sunday. In its first 24 hours on Vercel’s AI Gateway it was picked up by about 13% of their paying teams – faster than any model they’ve ever carried. Worth ten minutes of your time to understand what it is, and about the same again to understand what it isn’t. If you’re looking to get your hands dirty – you’re out of luck – as of Tuesday 22nd – TypeSafe AI isn’t taking on any more tire kickers.
What it actually does
Every AI model most of us have used in the last four years is, fundamentally, a writer. You ask it something, it writes you something back. That’s the interaction model and we’ve stopped noticing it.
Jev refuses to write anything at all. You hand it some raw material – an email, a support ticket, a log entry, a document – and you ask it a question where you’ve already listed the possible answers. Is this billing, technical or sales ? On a scale of routine to urgent, how bad is this ? Is this person angry, yes or no ?
It picks one. That’s the whole output. No explanation, no “certainly, I’d be happy to help”, no prose. Just the answer, plus a number telling you how confident it is.
“Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.”
Diogo Almeida, TypeSafe AI
It’s the difference between asking a colleague to write you a report and asking them to just point at the right folder. And if you’ve ever actually looked at what software spends its time doing – pointing at the right folder is most of the job.
Why anyone is excited
Two reasons and they’re both unglamorous.
It’s fast. Answers come back in well under half a second. A chat model asked the same question takes seconds, because it’s busy composing sentences one word at a time before it gets round to the bit you wanted. Jev produces the whole answer in one go. Ask it five questions about the same document and it answers all five in roughly the time it takes to answer one – which turns out to matter a lot, and I’ll come back to that.
It’s cheap. Roughly four cents per million words in, and the output is free. That’s not a discount, it’s a different order of magnitude.
Put those together and a whole category of thing that was previously too expensive to bother with becomes trivial. Every decision point buried inside a piece of software – which queue does this go in, does a human need to see this, should we escalate, does this look dodgy – can now have a bit of judgement applied to it for essentially nothing. The model is named after William Stanley Jevons, the Victorian economist who noticed that when something gets cheaper we don’t save money, we just use far more of it (Jevon’s Paradox). They are not being subtle about the thesis.
The bit where I stop being impressed
TypeSafe’s homepage says Jev is 193 times faster and 444 times cheaper than the alternatives. Those are lab numbers – a single isolated decision, measured against a single call to a big expensive model, on test cases the company designed itself. To their credit they more or less admit this in their own materials.
The real world figures published so far are a lot more modest. A production document pipeline saw about six times faster. An independent test came in around five. And in the most honest datapoint of the lot, one of TypeSafe’s own employees swapped a single step of a live pipeline over to Jev and measured a 15.9% speed improvement and a 30% cost reduction.
Not 193 times. Fifteen point nine percent – because replacing one link in a chain doesn’t make the chain much shorter. Anyone who has ever done performance optimization knows this feeling.
For the record – 15.9% is a good result, and the honest range of three to thirty times faster is an excellent one. I’d have more confidence in the product if the marketing had more confidence in the actual numbers.
The other claim doing heavy lifting is that Jev “can’t hallucinate”. This is technically true and practically misleading. Because the answers are listed up front, it can’t invent a fourth option when you gave it three. What it absolutely can do is pick the wrong one of the three and tell you it’s 90% sure. That isn’t hallucination in the technical sense but it’ll ruin your afternoon in exactly the same way.
And that confidence score – the thing that’s supposed to make this trustworthy – is the part I’d test hardest. Someone ran an independent check on it. On standard public test sets, the confidence numbers were honest and well behaved. On a set of invented support tickets the model couldn’t possibly have seen before, they drifted badly, and in different directions depending on the type of question – overconfident when picking from a list, underconfident on yes/no. Useful, but you’d want to calibrate it against your own data before you build anything load bearing on top of it.
What I’d actually do with it
Assuming you’re building something, a few things worth knowing :
Break the hard question into several small ones. This is the genuinely interesting finding. On a single difficult judgement call, a decent small chat model beats Jev comfortably. Decompose the same judgement into five simple typed questions and Jev wins – it went from 62% to 95% on one phishing test while the chat model it was up against actually got slightly worse. That’s an unusual property and it changes how you’d design around it.
Don’t trust the confidence number out of the box. Test it on your data, per question type.
It’s a component, not a system. It replaces the decision, not the pipeline. See the 15.9%.
Price your own workload. One developer found it 10-20x more expensive than the alternatives for his particular shape of traffic. TypeSafe also say straight out that they can’t prove the price isn’t subsidised, which I appreciate them saying and would still plan around.
Today versus where this goes
Today it is good at sitting inside software and making small, fast, cheap, repetitive decisions that previously needed either a rigid rule or an expensive model. Sorting, routing, triage, flagging, screening, guarding an agent before it does something silly. It is not good at anything requiring explanation, nuance, or a question you didn’t anticipate. It can’t tell you why. And it runs only on TypeSafe’s servers – no published research, no model you can run yourself, no option to put it inside your own building. In my day job that last one is the end of the conversation, not the start of it, and I suspect that’s true for a lot of regulated industries.
Where this goes is the more interesting question. I don’t think the bet is that Jev gets smarter. The bet is that “AI model that returns a decision instead of an essay” becomes a standard shape that everybody offers – it doesn’t look especially hard for the big labs to copy, and open source versions are apparently already turning up within days of launch. If that happens, the lasting contribution here isn’t this particular model, it’s the idea.
This is directionally very encouraging – we’re burning tokens at an alarming rate and using a warehouse sized brain to answer trivial classification questions – remember early ChatBots trying to auto-complete math questions and failing – now they just reach for a Python CLI. Horses for courses.
Jev went GA on 21 September 2026. Everything above was accurate the day after, which in this field gives it a shelf life of about a fortnight.
We wanted to end the trip on a big day and Wilson Mountain was our final choice after considering a few alternatives. Wilson is the highest point around Sedona and it does not let you get there cheaply – a long, steep hike, done on tired legs after four days of canyon walking. We started early and cloud kept the sun off our backs in the morning but made up for it in the afternoon.
The views from the top were amazing and worth every bit of it.
Two hours south to Sedona via the route 66 town of Williams for dinner and vintage Americana.
My first time in Sedona – fairly largish tourist town full of crystal shops and Mexican restaurants and surrounded by some stunning landscape. A nice VRBO house up in the hills on the edge of town gave us stunning sunrise and sunset views.
Devils Arch from the Dry Creek trailhead was a shorter hike, made a bit longer by a wrong turn – which, in fairness, is becoming a tradition on these trips.
Probably the pick of the canyon days. Hermit sits west of the village and away from the corridor trails, and the difference is immediate – a far less touristed path, and an amazing landscape as you get down into the canyon. The shuttle bus takes about 20 minutes (a little quicker on the way back). The Hermit trail landscape is very different to SK and BA – even though you’re in a smaller side canyon – it feels like your much deeper down.
Santa Maria Springs is the turnaround, with its little stone rest house. If you’ve done South Kaibab and Bright Angel and want something quieter, this is the one.
Bright Angel is the other half of the corridor and a completely different day to South Kaibab. It works its way down through a side canyon, so you get shade at the right times and water at the rest houses – which after the day before felt like an outrageous luxury. There were warning that the water would be completely shut off at 11am but that was delayed until after we’d finished. Lucky for me – I sweat a lot and need to drink a lot – without the opportunity to refill I would’ve been in trouble. I think 5L of water is probably the minimum you need for a hike like this in the heat.
Havasupai Gardens is a little oasis – cottonwoods, running water, and a genuine change in temperature when you walk into it. This is also the day that gives you a good sense of doing the rim to rim, which was as close as we were going to get this time.
Ten of us were supposed to walk the Grand Canyon rim to rim in the first week of September – down the North Kaibab from the North Rim, two nights at Phantom Ranch, then out to the South Rim. The cabin was booked, the Trans-Canyon shuttle was booked, and we’d spent the the last few months training with regular walks at Umstead and single day ascent / descent of Mount Mitchell.
Then, on the afternoon of Saturday 30th August – five days before we flew – nearly an inch of rain fell in thirty minutes onto the Dragon Bravo burn scar above Bright Angel Canyon. The Dragon Bravo fire stripped the Kaibab Plateau back in July 2025 and a Forest Service assessment had already flagged that watershed as high risk for post-fire flooding and debris flows. It did precisely what the assessment said it would. Around 50 people were airlifted out of Phantom Ranch, 2 people killed, and one still missing. Nearly all the footbridges over Bright Angel Creek were destroyed.
Closed until further notice : Phantom Ranch, Bright Angel Campground, all fourteen miles of the North Kaibab, and the Colorado to boating.
Then the second punch. The flood took out the trans-canyon waterline – which is what supplies the South Rim – so every in-park lodge closed to overnight guests from the 31st with no reopening date. We lost the hike and our accomodation in the same week.
So : Plan B. Cancel everything, keep the flights, and rebuild the week out of the trails that were still open. What we ended up with was five days of day hikes off the South Rim and out of Sedona – not the trip we planned, but an epic trip in its own right.
Day 1 – getting there
RDU to Vegas, then the four hour drive to the South Rim. Just enough daylight left for a walk along the rim at sunset – which is still the best possible introduction to the place, however many photographs of it you’ve seen.
Days 2 to 6
Three days in the canyon, then two in Sedona, stats in the table and full routes in the Strava links in the posts:
The canyon doesn’t owe you the itinerary you booked. We had eight days’ notice and the closures were absolute. Everyone rolled with it without a single moan, which is the real reason the week worked.
Day hiking off the South Rim is underrated. Same geology, same light, and you sleep in a bed at night. A different trip to the rim to rim, but not a lesser one.
South Kaibab down, Bright Angel down, Hermit down – three trails, three completely different experiences of the same hole in the ground. Doing them on consecutive days makes the contrast obvious in a way one big traverse never would.
We probably hiked double what we expected to hike doing the R2R.
Book refundable where you can. Cancelling a cabin, two shuttles and a lodge in one afternoon is not how you want to spend a Sunday.
The rim to rim is still on the list. Phantom Ranch, the North Kaibab and those footbridges are going to take a while – maybe next year.
Finally – a huge thanks to Bryan and Iain for organizing, then re-organizing the trip and the whole gang for making the trip memorable.
First real day of hiking with Kip, Iain and David. Even if we couldn’t get down to the river, we wanted to hit the major corridor trails on the South Rim – starting with South Kaibab. South Kaibab shows you the canyon all at once. There’s no easing in – no switchbacks up a side canyon, no gradual reveal – you drop straight onto a ridge line and the whole thing is just there, on both sides of you, the entire way down.
Skeleton Point is the turnaround and it’s the first place you properly see the Colorado – we walked a little further down to find some shade for lunch and were treated to views down to the river below, and condors or hawks flying above.
This was our biggest training hike in preparation for the Grand Canyon R2R.
Mt Mitchell is the highest point east of the Mississippi River and the closest thing we have to a real mountain in NC . It rises to 6,684 feet, so gave us some good elevation as well.
We started at the Black Mountain Campground off of South Toe River Road which makes it a 3,694 ft climb and a 11.8 mile round trip. Weather was hot and humid – underfoot was rocky, rooty and a little slippy in places. It’s a very decent slog and I can’t believe this is the first time I’ve hiked our biggest local mountain. Views from the top were awesome – coudl see all the way to Beech and Sugar.
The trail on AllTrails is here, and my Strava activity is here.
As it turned out, the rim to rim didn’t happen – a flash flood closed Phantom Ranch and the North Kaibab a few days before we flew. What we did instead is here.
The next big feature is integrated the PullBook App with the Half Decent scale so I can completely automate the “pull a shot” workflow in the app but due to a shipping delay from Hong Kong all I have is a rough implementation and a simulator as shown below :
Status report – I gave the web site a bit of a makeover and improved the automation but nothing major. I also fixed a few minor bugs in the app and ended up in a bit of a merge hell. I have a branch open for the scales integration on one machine and working on bugs on another machine and at one point I just wanted the two Claudes to connect to each other and figure out how to move forward – relaying GitHub instructions through me was just slowing the whole process down and I don’t think I was adding much to the discussion. Anyway – everything got synched and committed and a new Apple TestFlight build is on its way.
I’m still manually onboarding beta participants – we’re talking small numbers so it’s entirely manageable but would love to automate this but away as well. My long term goal is to automate everything and reduce my role to design authority, product strategy, release planning / feature prioritization and any task that still needs a pair of hands to complete.
What this experiment has show me so far is :
I think my PM background and distance from the underlying tech is an advantage – there’s a strict division of labor between me (the what) and Claude (the how). I occasionally have to wade in with an opinion (usually in the form of a question) if I think the implementation may not be optimal. I’m equally open-minded and generally interested in Claude’s opinion of the what.
Everything is falling into place for self-improving / self-adapting software. Take the feedback, prioritize it, implement it, test it, ship it. Rinse and repeat. An external API changes or some other catastrophic bug – find it, fix it, ship the fix. All implemented with Claude in a local loop. There needs to be name for this.
At some point I need to invest a bit of time in automating the App store distribution – right now that’s very manual and requires XCode – my suspicion is that I will have to pay GitHub (for the OS/X images) or Apple – XCode Cloud and that would violate one of my requirements for this project that I don’t spend any money (beyond Anthropic tokens).
Reminder, if you are an Espresso aficionado, you might find the app very useful, more information here :
On these pages, I’m keeping notes of my experimentation with new AI based software development tools. My goal is to understand for myself whether tools like Claude and Codex are ready for developing commercial software. Low code and zero-code tools have been around for a decade or so – this is not the same. Vibe coding has been around for a few years as well – this isn’t the same.
This new approach has various names – agentic coding, spec driven development, personally I think “intent driven development” captures it best, though I’m confident that name won’t catch on. With intent driven development – I’m responsible for the what, the AI is responsible for the how.
As the cost and time to develop software shrinks, the focus will shift to other critical elements of the product lifecycle. I believe the next areas of differentiation for builder / product organizations will be:
Choosing hard problems to solve and solving them in ways that delight users. The easy problems can now be easily solved with a few well authored prompts and a vibe coding platform. Vibe Coding is incredible for solving small local problems – it’s the new Excel (and I mean that with a huge amount of respect)
Accelerating all the other toil related to building and product – pricing, training, marketing, enablement
Convincing enough people that your solution to the problem is the best. As time to market collapses, you have to assume competition and cheap imitation will be rife – how do you stand out ?
And so it is with my little pet project. I’ve invested 6.5 hours in the development of the application – included in that is some overhead in getting the development environment setup (GitHub automation, local Xcode), design, development and building regression test harnesses. For fun, I did a little costing analysis on the main code repo using scc – clearly COCOMO hasn’t really kept up with even pre-AI development tools but it’s a good way to reason about the magnitude in productivity – 8 months compressed to a week – even if the COCOMO estimate is out by a factor of 10 – the point still stands.
But that’s not really the point of this post. In my humble opinion, software development for low-stakes, greenfield problems are largely solved at this point and development time and cost is collapsing – it’s hard to debate this any longer and I personally don’t require any convincing. I don’t see an imminent plateau either – even if the models are throttled – agentic development can still scale in other ways. The way we develop software – heavily augmented with AI and human as the designer / orchestrator – that’s the way – there’s no going back.
So on to the next bottleneck – that’s pretty much everything upstream and downstream of software development. Software development is the tip of the spear in terms of agentic augmentation and performance improvement – everything else is next.
I spent a few frustrating hours this week downstream – fighting with the Apple TestFlight review process – clearly the user-centric design and detail to attention that Apple products are known for is pretty much absent from their developer tooling. I won’t belabor the point here but it’s classic “toil” – work that has to happen but brings you no joy at all. But it does highlight the point about the whole product lifecycle – everything has to get leaner and faster – it’s not going to be acceptable to wait three days for an App Store approval if the app only took a day or two to develop. You can’t spend a month on a marketing plan for a new release if the release will be out in two weeks. Everything has to speed up.
As part of the early access stage for my little test app – I needed an onboarding process. I could have used the Apple TestFlight defaults but want to own the onboarding and capture some user information in the process so I pretty much single-shotted a simple website (hosted on GitHub pages) and setup a Google form to capture registrations – not as automated as I’d hoped but mostly a one time cost.
Now I have the simple web presence – I have a home for the app’s change log so I got Claude to create an Action to create release notes every time I push a new release. The only twist here is that the Action calls out to Anthropic to turn PR text into human readable release notes – so far It seems to work well.
How we think about planning and project management will have to change. Even for a small one man project – normally there would be some planning involved but what I’ve learned this week is that if you have time to write a GitHub issue or a Jira ticket, then you probably have time to “do the thing” you were going to write the ticket about. What kind of planning is eve required in world of constant feature delivery – all you really need is direction / themes and prioritization.
The next thing I’m thinking about is how to take this experiment further – I have the first batch of beta users (you can sign up here if there are still slots available) and I expect to get some feedback – I’m hoping I can largely cut myself out of the loop. My goal is for Claude (or whatever) to present the work planned for the next release (new features, bug fixes, tech debt) based on customer feedback and product analytics – let me review it then just get on with the implementation, testing and delivery. Self improving software/ products / systems with human defined policy and guardrails – that’s where were heading. Sure someone still needs to inject strategy and innovation into the loop but product maintenance and improvement can be largely automated.
Three different people sent me the same link this week, which is usually a decent signal that something is either genuinely interesting or being very well marketed. In this case I think it’s a bit of both.
We wanted to end the trip on a big day and Wilson Mountain was our final choice after considering a few alternatives. Wilson is the highest point around Sedona and it does not let you get there cheaply – a long, steep hike, done on tired legs after four days of canyon walking. We started…
Two hours south to Sedona via the route 66 town of Williams for dinner and vintage Americana. My first time in Sedona – fairly largish tourist town full of crystal shops and Mexican restaurants and surrounded by some stunning landscape. A nice VRBO house up in the hills on the edge of town gave us…