🔍 Read the full analysis: 24 Use Cases For Jev In AI-Assisted Decision-Making on ThorstenMeyerAI.com
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer’s Sept. 29 article maps 24 potential uses for Jev across publishing, commerce, software, business operations and the home. He says three are live in his publishing operation, 12 are strong fits, seven need measurement and two are poor fits; the article’s detailed examples and metrics come from his own operation and testing.
Meyer says Jev accepts a state, such as text or JSON, alongside typed questions, then returns answers that code can use without parsing a prose response. Its answer types include yes-or-no probabilities, choices with probabilities and confidence, and scores on ordered levels. He says a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. These performance and price figures are claims in his article; no independent benchmark is provided there.
The three live applications are a relevance gate matching stories to a site, a language check and a fallback topic classifier. Meyer reports that a scan of 78,889 articles cost $2.01; the language check found 1,576 non-English articles and he says 1,553 were fixed. The relevance gate assessed about 10,000 story-site pairings in three days, with 22% judged clearly on-topic. For the classifier fallback, he reports 89% agreement with a frontier large language model, rising to 97% to 99% for answers with confidence of at least 0.8. The article does not provide independent validation of these results.
Among publishing examples, Meyer marks disclosure detection and comment moderation as strong fits. He says disclosure checks could catch paraphrases missed by a regular expression, while uncertain cases should go to human review rather than automatic publication. He labels a thin-source detector, product matching in roundups and headline quality checks as needing measurement first. Same-event deduplication is a poor fit in his assessment: he says his canary test found no duplicates, leaving no demonstrated problem for the tool to address.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Confidence Can Cut Review Work
The proposal is aimed at work involving many small decisions, where clear cases can be handled automatically and uncertain ones can be routed to a person or a more capable system. Meyer says Jev’s reported cost and speed could make these checks practical at scale. Whether that applies to other users would depend on task-specific accuracy and the cost of errors.
Meyer’s fit test asks teams to confirm four conditions: high volume, a narrow question, low-cost errors or a route for uncertain cases, and a heuristic that is measurably failing. He says this requirement helps teams determine whether an existing rule is working. In his deduplication example, a canary test found no duplicates, so he classifies the use as a poor fit.
Meyer recommends that teams replay 300 to 500 past decisions, examine disagreements, and connect the system to a workflow only if the high-confidence band reaches 95%. He also recommends a separate feature flag, initially off, followed by a 5% to 10% canary. These are practices he proposes; the article does not report that every listed use has passed such an evaluation.
How Meyer Defines Jev’s Fit
Meyer describes Jev as a decision aid rather than a writing, summarization or extraction system. The application supplies the material and questions; its own code determines what happens next. In his examples, Jev might classify a comment or assess whether a disclosure is present, while rules outside the model decide whether to approve, hide, publish or send a case for review.
The article says confidence is central to this design. Meyer cites an internal measurement on a 31-topic classification in which Jev agreed with a frontier model 97% to 99% of the time at confidence of 0.8 or higher, compared with 42% below 0.5. The comparison is specific to that measurement and does not establish the same accuracy for other tasks. Meyer’s broader list spans publishing and content, commerce, software, business operations and home use, but the supplied article excerpt details only part of the list.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Accuracy Beyond Meyer’s Tests
The article presents Meyer’s measurements and production figures, but does not describe an independent audit, the full evaluation method, or how results vary across data sets and use cases. It is unclear whether other operators would see similar accuracy, cost and speed. The reported agreement with a frontier model is also a comparison to that model, not a direct measure of truth.
The supplied source material ends partway through the commerce and customer operations section. It does not include the full descriptions of all 24 use cases, so the remaining cases and their specific evidence cannot be assessed here. Meyer’s overall tally is stated, but the available detail does not establish how each non-publishing case was tested.
Measure Before Wider Deployment
Meyer’s recommended next step for teams considering a use is to replay 300 to 500 real past decisions, compare results overall and by confidence band, and review 20 disagreements to determine which answer was right. He says deployment should follow only where high-confidence performance reaches 95%, with an off-by-default flag and a small canary before wider rollout.
For the seven cases he labels “measure first,” the next milestone is evidence that the existing heuristic is failing and that the new check improves the workflow. The article does not give a date for further results or a rollout beyond Meyer’s current publishing operation.
Key Questions
What did Thorsten Meyer publish?
He published a map of 24 possible uses for Jev, a tool that returns confidence-scored answers to narrow questions for software workflows.
How many use cases does Meyer say are ready or running?
Meyer says three are live in his publishing operation and 12 more meet his four-condition fit test. He classifies seven as needing measurement and two as poor fits.
What examples does the article describe in detail?
The source details three live publishing applications: story relevance checks, language checks and fallback topic classification. It also discusses disclosure detection, comment moderation, thin-source detection, product matching, headline quality and event deduplication.
What evidence supports the reported performance?
The figures are Meyer’s own reported measurements, including a 31-topic classification comparison and results from his publishing operation. The article does not provide independent validation or enough methodological detail to determine how well the results generalize.
When does Meyer say Jev is a poor fit?
His test requires high volume, narrow questions, inexpensive or safely routed errors, and a visibly failing existing heuristic. He says a canary found no duplicate events in his workflow, so he classifies deduplication as a poor fit for now.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
