🔍 Read the full analysis: A Playbook For Jev: 24 Ways To Apply Decision Models In AI on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer published a playbook describing 24 ways to use Jev, a decision model for answering typed questions about text or JSON. He says three uses are live in his publishing operation, 12 meet his fit test, seven need measurement, and two are poor fits. The published material reports results from his own operation; independent verification and details for all 24 applications are not provided here.
Thorsten Meyer published a 24-use-case playbook for Jev, a system that returns typed answers to decision questions so software can route or act on them. Meyer says three applications are already live in his publishing operation, while 12 more qualify as strong fits under his test. The playbook describes an approach for teams considering whether repeated judgments can be automated while uncertain cases are routed for review.
Meyer describes Jev as a tool that does not write, summarize or extract information. Instead, a caller sends text or JSON state and typed questions, then receives answers that code can use. The supported answer types include a yes-or-no probability, a choice among options with probabilities and confidence, and a score on ordered levels. Meyer says one request can carry multiple questions, takes about 0.3 to 0.9 seconds, and costs about $0.04 per million input tokens. These performance and pricing figures are his claims in the playbook.
The suggested operating pattern is to act on high-confidence answers and send uncertain cases to a person or another system. Meyer reports that, in a 31-topic classification measurement, Jev agreed with a frontier large language model 97% to 99% of the time when confidence was at least 0.8, compared with 42% agreement below 0.5. The comparison is with that model, not a reported human-reviewed accuracy study. Meyer says the application’s code sets the action rule; Jev supplies an answer rather than deciding what happens next.
In the three uses Meyer says are running, a relevance gate assessed story and site pairings, a language check flagged non-English articles, and a topic classifier served as a fallback when a primary model erred. He reports that one relevance effort judged about 10,000 pairings in three days, with 22% clearly on topic. For the language check, he says a scan of 78,889 articles cost $2.01, found 1,576 non-English articles and fixed 1,553. His classifier reportedly agreed with a frontier model 89% overall. Those figures describe his operation and are not independently confirmed in the material provided.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Small Decisions Add Up
The playbook proposes using a low-cost model to check routine items and routing ambiguous decisions for review based on confidence. For publishers, examples include language checks, disclosure detection and comment moderation. Meyer proposes sending uncertain disclosure cases for human review rather than allowing automatic publication. In this design, the model provides an answer and application code determines whether to act on it.
Meyer’s four-part test requires high volume, a narrow question, errors that are inexpensive or routed to a stronger reviewer, and evidence that an existing heuristic fails. He advises replaying 300 to 500 past decisions, reviewing disagreements and checking performance by confidence band before putting a use into production. He recommends enabling a feature flag gradually, beginning with a 5% to 10% canary. These are deployment recommendations from the author; the material does not establish that every proposed use has passed such evaluation.
The Playbook’s Fit Test
The material groups the 24 proposed applications into four categories: three live, 12 strong fits, seven that need measurement, and two poor fits. Meyer says a use should qualify only if all four conditions are met. He says a failing heuristic should be measured rather than assumed. If a simple keyword rule works, his advice is to keep it until evidence shows a problem.
The excerpt supplies six publishing examples, not the full list of 24. It labels disclosure detection and comment moderation strong fits. Thin-source detection, product relevance in roundups and headline-quality checks need measurement first. Same-event deduplication is a poor fit in Meyer’s account: he says a canary found no duplicates, so he found no demonstrated problem for the system to address. The material says the remaining examples span commerce, software, business operations and the home, but their individual details are not included here.
For the thin-source detector, Meyer says 88% of processed news items begin from a bare headline. The proposed question is whether a source contains enough verifiable facts to support a factual report. This describes his workflow and the motivation for the test; it is not a general estimate of how news organizations source stories.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, in the playbook
Evidence Still Needed
The results are presented by Meyer, whose publishing operation is the setting for the reported live uses. The supplied material does not include an independent audit, detailed evaluation data, or the underlying comparison set for the reported agreement rates. It also does not specify the frontier model used, the timing of each measurement, or how disagreements were adjudicated. Those details would help readers assess how the figures apply to other workloads.
The excerpt does not describe all 24 use cases, so the stated totals cannot be examined use by use from the material available. For the seven marked “measure first,” it is also not clear what error rates would qualify as a failing heuristic in each setting. Meyer’s proposed replay and canary steps provide a process, but no results are supplied for those pending applications.
Measure Before Wider Rollout
Meyer’s proposed next step for a new application is to replay several hundred real past decisions, compare results across confidence bands and review a sample of disagreements. He says a use should be wired in only where the high-confidence band reaches 95% accuracy, followed by a feature flag and a limited canary. The playbook does not give dates for these evaluations or announce a rollout schedule for the cases still awaiting measurement.
Further evidence could include the remaining case-by-case details, the evaluation methods behind the reported results and measurements from teams beyond Meyer’s operation. The available material describes an author’s applied playbook, with three reported production examples and a test for deciding where further trials may be conducted.
Key Questions
What did Thorsten Meyer publish?
He published a playbook mapping 24 possible Jev applications across publishing, commerce, software, business operations and the home.
What does Jev do?
According to Meyer, it answers typed questions about supplied text or JSON with probabilities, choices or scores. The calling software sets the action taken on each answer.
How many uses does Meyer say are live?
Meyer says three are live in his publishing operation. He classifies 12 as strong fits, seven as needing measurement and two as poor fits.
Are the reported performance figures independently verified?
The supplied material presents them as measurements from Meyer’s work. It does not provide an independent audit or enough evaluation detail to verify how the results generalize.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
