Skip to content

Build 2b90e1d ·

← The Log
Deep dive · · 12 min read

A model that can't write a sentence

TypeSafe's Jev only answers in probabilities. One morning on real work: a 96% that measured the wrong thing, a 13 of 13 that held, and one job I kept from it.

One Jev call: a made-up restaurant lead goes in, three probabilities come out (0.88, 0.90, 0.93). Below it, a confidence line at 0.7: thirteen of thirteen answers above it right, the answers below it about half wrong.

TypeSafe’s Jev has never written a sentence, and by design it can’t. You hand it some state, a block of text or JSON, plus a list of typed questions, and it hands back one number per question.

That’s the whole interface. I spent one morning finding out what those numbers were worth on real work: where they were worth less than they looked, where they held, and one job I decided not to give it at all.

What TypeSafe Jev is, and the endpoint that 404s

TypeSafe AI launched Jev on September 15 as its first “System One” model. Their pitch is “unstructured state in, typed probabilistic decisions out.” They also say it “can’t hallucinate.” That’s a vendor claim, and it’s true in a narrow way: a model that produces no text can’t make text up. It can still be confidently wrong, which is most of what this post is about.

There are three question types, named on Cloudflare’s docs page for the model:

  • noul is yes/no, and you get back P(yes).
  • choice picks one of the options you supply, with a probability for each.
  • score places the state on an ordered scale you define.

All the questions in a request are answered in one parallel pass, so a third question doesn’t cost you a third round trip.

It’s on OpenRouter with no waitlist, as ~typesafe/jev-latest. When I tested, that alias resolved to typesafe/jev-1.13-20260917.

The first gotcha: the endpoint that works is POST https://openrouter.ai/api/alpha/decisions. The /api/v1/alpha/decisions form, which is what you’d write if you’ve used any other OpenRouter route, returns a 404. The model doesn’t appear in /api/v1/models either, so if you look for it there you’ll conclude it doesn’t exist.

Here’s a working request with all three question types:

# /api/alpha/decisions, NOT /api/v1/alpha/decisions (that one 404s)
curl -s https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "~typesafe/jev-latest",
    "state": { "message": "Hi, we own a small restaurant and need a new website before next month." },
    "questions": {
      "is_real_lead": { "type": "noul",
        "instructions": "Is this a genuine prospective client asking about paid web work?" },
      "intent": { "type": "choice",
        "instructions": "What is the sender primarily asking for?",
        "criteria": { "new_site": "A new website or rebuild", "fix_existing": "A fix to an existing site", "other": "Something else" } },
      "urgency": { "type": "score",
        "instructions": "How soon does this lead need a reply?",
        "criteria": ["This week is fine", "Within a day or two", "Today"] }
    }
  }'

choice criteria is an object of options; score criteria is an ordered array. state can be a string or JSON.

That was my first test, a made-up contact-form lead. Back came a real lead at 0.88, new_site at 90%, and urgency at a score of 1.07, with 93% of the probability on “within a day or two”. The response reported 655 input tokens and a cost of 0.00002751 USD.

Pricing, per the vendor: 0.042 USD per million input tokens, and output is free. TypeSafe’s phrase is “too cheap to meter.” Calls took 300–450 ms each when I timed them.

It made a tidy demo, and my plan was lead triage. Then I checked the premise before writing any code: my contact form didn’t have enough volume to test on. There was nothing to measure, so I dropped the idea before building anything.

The 96% that was mostly free

So I pointed it at a question I actually have. I attribute my AI coding sessions to clients, for invoicing and for project history. Which client was each of the 230 or so sessions from the last 60 days for? The workflow those sessions come from is in How I Actually Use Claude Code.

A plain folder lookup places about a third of them: if a session ran inside a client’s repo, the path says whose it is. That gave me 78 sessions with a known answer.

On those 78, Jev got 75 right: 96%. A keyword matcher I wrote as a baseline got 68.

Then I looked at the inputs. 73 of the 78 session summaries contained the project path. The model wasn’t inferring anything. It was reading the answer off the page.

The 78 folder-placed sessionsCount
Jev right75 (96%)
Keyword matcher right68 (87%)
Summary contained the project path73

That 96% measured the path, not Jev. It’s the same family of mistake as the watchdog in 14 runs, 0 alerts: a check that looks like it’s testing one thing while it’s testing another.

The hard 152, and the 13 of 13 that held

The other 152 sessions ran from my home folder or my tooling folder. No client path to lean on. A folder lookup can’t place them, which makes them the only sessions where a model has to earn anything.

There’s no answer key for these, so I made a small one. I sampled 20 at random and had Claude grade Jev’s answers against the session transcripts. One had too little text to grade, leaving 19. That grading is self-graded, not a verified key, and 19 is a small sample.

In the first round, the 10 answers at a reported confidence of 0.7 or higher had no clear errors. Below 0.7, half were clearly wrong, and those low-confidence answers changed between runs. Same input, different client.

The misses were name collisions. Two unrelated people and businesses shared a first name. One session got pinned on the wrong client because an email address in it contained an agency’s domain name, and the model took the domain as the signal. Reasonable, and wrong.

I reworked the questions and added a noul, “is this about personal life rather than business?”, with an override: if the client pick is one of the two names that collide with personal contacts and that answer is above 0.5, the session goes to nobody. On the second round, same 19 sessions, 13 answers cleared 0.7 and all 13 were right. Below the line, 4 of 6 were wrong.

Two caveats. The fix was scored on the same sessions its misses came from. And the override barely mattered: it fired 3 times across all 230 sessions, and on the actual misses the personal question scored only 0.39 to 0.47, under its own threshold. Overall accuracy moved from 14 of 19 to 15 of 19.

What moved was how much of it I could trust. Above the line, the answers held. Below it, I was paying a fraction of a cent to flip a coin.

Confidence calibration: the number wasn’t what I assumed

I’d been treating the reported confidence as a probability of being right. It isn’t quite that.

Stanislav Yurin published an independent analysis on September 18, across 738,164 live choice answers. His data shows that for choice questions, Jev’s reported confidence is exactly a rescaled top probability:

confidence = (N · p_max − 1) / (N − 1)

N is the number of options, and p_max is the probability of the top one. Hold p_max steady, add options nobody picks, and confidence goes up by itself. My client list is a choice question. Every client I added would have quietly loosened the gate.

So I switched the gate to the top probability directly: 0.72. At 19 options that’s equivalent to the 0.7 I’d tested, since (19 × 0.72 − 1) / 18 = 0.704. Same line today, and it stays put when the list grows.

Aman Kumar’s independent test, about 16,000 calls, also published September 18, landed where my 19 did: “confident answers were right 90 to 100% of the time; the middle is a coin flip.” His sample is several hundred times larger than mine, and that’s the most reassuring thing about mine.

I nearly didn’t find either write-up. Six days after launch, Perplexity told me Jev had no public adoption. Exa, run right after, found LangChain, the Cloudflare docs, both of those tests and a Java SDK. That gap turned into a small skill that picks a search engine by the kind of question before I trust any one answer.

What shipped: 76% automatic, the rest to me

The session sorter that runs now has three layers:

  1. Folder lookup. Exact and free. If the path names the client, done.
  2. Jev, above the gate. Top probability 0.72 or better.
  3. Me. Everything else is flagged for review.

That assigns 76% of sessions automatically. On the same 19 graded sessions, it assigned 14 with none wrong and sent 5 to me. Jev’s guess was right on 4 of those 5; it just wasn’t confident, and not confident is exactly when I want to look.

One filter runs before any of it. Each event, a tool call or a prompt, owns the time until the next event on one merged timeline, and any gap over 30 minutes counts as idle, so two sessions running at once never count the same minute twice. Automated sessions, such as email daemons and a commit-review hook, get dropped first. Otherwise a bot reply at 2 AM would bridge an idle gap and add hours that never happened.

The news briefing, and an answer key that stopped in July

The second job was a daily dev-news briefing built from Cloudflare, Astro, Vue and Nuxt feeds. It had been picking articles with a GPT prompt. Would Jev pick better?

For an answer key I used my own starred articles from March through July: 420 articles, 138 starred. Also self-graded, since the stars are mine. Compared at the same number of picks:

PickerPrecisionRecall
Old GPT prompt49%59%
Jev57%68%
Jev, tagging against scanned repos54%64%

Scored on March through July only. For scale, 138 of 420 is 33%, the precision of picking at random.

The answer key had its own trap. I stopped starring after July, so my first score, run across every month, counted every good August pick as a false positive. Jev was losing points for surfacing “D1 enforces free tier daily query limits” and “Nuxt 3 EOL”, both of which I’d have starred on sight. An answer key that ends in July can’t grade August.

The bigger change is the third row. The briefing used to match articles against a project list I kept by hand, and the list was stale. Now it scans every repo, reads the real stack from package.json and the wrangler bindings (D1, R2, KV, Workers AI and so on), and Jev tags each article by the features it touches. Every article lists the repos it affects, split into “ACTIVE THIS WEEK” and “WARM”.

On the old stars it scores slightly lower. I kept it anyway. It reads the repos on every run, so there’s no hand-kept list to go stale, and the old stars were graded against the old list, so of course they prefer it. The WARM list does something the old briefing never did: it lights a fire under my ass to revisit a project.

Jev on Workers AI, with two gotchas

Jev is also on Workers AI as a third-party model, typesafe/jev. This is the example from my r/CloudFlare post:

const res = await env.AI.run('typesafe/jev', {
  state: 'Help! My payouts have been failing for 3 days.',
  questions: {
    is_urgent: { type: 'noul', instructions: 'Does this convey urgency?' },
    department: {
      type: 'choice',
      instructions: 'Which team should handle this?',
      criteria: { billing: 'Payments, refunds', technical: 'Bugs, outages', sales: 'Pricing, upgrades' },
    },
  },
})

It came back billing at 88% and urgent at 95%.

The first gotcha is billing, and this is what I observed, not what’s documented. The model page only says “View pricing in the Cloudflare dashboard.” In practice it bills against AI Gateway credits or BYOK, not regular Workers AI usage. With no credits loaded, you get:

2021: Insufficient balance; add money to your gateway or use BYOK

When I topped up, the minimum was 10 USD plus a 0.50 USD fee.

The second is the response shape. Through the REST API the answers sit at result.result.answers, next to "state": "Completed", which isn’t the shape on the docs page.

Latency through the REST API, not from inside a Worker, was about 1.2 s cold, then 340–480 ms. Moving the session sorter and the briefing into Workers is the plan, not something I’ve done, and I haven’t measured latency from inside a Worker yet.

The job I didn’t give it

Claude, working with me that morning, suggested a hook that would run Jev on every prompt I type. My first question was whether it ran locally.

It didn’t. The hook would run on my machine, but every prompt would go out to OpenRouter and TypeSafe. Routing it through Cloudflare instead doesn’t help: TypeSafe still receives the text, and AI Gateway logging is “enabled by default for each gateway,” request and response data included. My prompts carry client work. The thing the hook would improve rarely misses anyway. I dropped it.

The rule I landed on: third-party AI gets public data, or data approved for a specific purpose. The session summaries went out for one purpose, attributing work to clients, and I decided that was worth it. Every prompt I type, for no particular purpose, stays home.

The sessions this sorts are counted on the dashboard.

What do you gate on when a model hands you a number? If you’re running Jev on something, I’d like to hear what: the Reddit thread is open, and my email is on the front page.