🔍 Read the full analysis: How To Explore AI Decision Models With Jev: 24 Ideas on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A Sept. 29 article by Thorsten Meyer maps 24 possible uses for Jev, a tool that returns typed answers to narrow questions for software to act on. Meyer says three publishing uses are live, 12 ideas meet his fit criteria, seven need measurement and two are poor fits; the supplied material details only the publishing examples.
Thorsten Meyer published a map of 24 potential uses for Jev on Sept. 29, reporting that three checks are running in his publishing operation and have processed about 90,000 decisions. The article classifies 12 ideas as strong fits, seven as needing measurement and two as poor fits, and sets out criteria for assessing narrow automated judgments in other workflows.
Meyer describes Jev as a system that accepts text or JSON plus typed questions and returns answers that software can use directly. Its response types include a yes-or-no probability, a choice among options with probabilities and confidence, or a score on ordered levels. Meyer says a call with multiple questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
In a measurement on a 31-topic classification task, Meyer reports that Jev agreed with a frontier large language model 97% to 99% of the time when confidence was at least 0.8, compared with 42% below 0.5. He presents this as his own measurement, not a general benchmark. His proposed pattern is to let code act on clear cases and route uncertain answers elsewhere.
The three live examples concern content relevance, language detection and fallback topic classification. Meyer says a one-night scan of 78,889 articles cost $2.01 and found 1,576 non-English items, of which 1,553 were fixed. For relevance, about 10,000 story-site pairings were judged in three days, with 22% deemed clearly on-topic. The classifier fallback reportedly agreed with a frontier model 89% overall and 97% to 99% at confidence of 0.8 or higher.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Fit
The article describes criteria for choosing among a specialized model, a simple rule or human review for routine decisions. Meyer proposes confidence-based handoffs to apply checks across more items while routing uncertain cases through an existing workflow. The reported figures come from his operation and do not establish performance in other settings.
The examples involve different consequences. A language flag may lead to an in-place rewrite, while a missed disclosure about free products or affiliate links could raise a compliance concern. Meyer recommends human review of missed disclosures rather than automatic publication. In his account, software rules and review procedures determine how each answer is used.
Meyer’s Four-Part Fit Test
Meyer says a suitable Jev task has high volume, asks a narrow question, has errors that are cheap or can be routed to a smarter system, and replaces a heuristic that has been shown to fail. If a keyword rule already works, he advises keeping it until evidence shows a problem. The test is intended to prevent teams from adding a model simply because a use case sounds plausible.
Before wiring a candidate into production, Meyer proposes replaying 300 to 500 real past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements. He says to proceed only where the high-confidence band reaches 95%. His rollout suggestion is a separate feature flag that is off by default, followed by a canary on 5% to 10% of units.
In publishing, he labels disclosure detection and comment moderation strong fits. Thin-source detection, product matching in roundups and headline quality need measurement first. Same-event deduplication is a poor fit in his account: a canary found no duplicates to address, leaving the claimed heuristic failure unproven. The source material also says the map covers commerce, software, business operations and home uses, but does not provide those individual examples.
““Use Jev only when all four conditions hold: High volume. Narrow question. Cheap errors. A heuristic fails visibly.””
— Thorsten Meyer
What the Report Does Not Establish
The source presents measurements from Meyer’s operation but does not provide the underlying dataset, evaluation method or independent replication for the reported agreement rates. It is not clear how the 31-topic test was sampled, how disagreements were judged, or whether results would hold for other content and workflows. The reported scan and cost figures are also not accompanied by a detailed accounting method.
The available source text stops partway through the commerce section. Although it gives totals for 24 ideas and their fit categories, it does not show the full set of ideas across commerce, software, business operations and home. The criteria behind the two poor-fit labels beyond the deduplication example are not available here.
Measure Before Wider Rollout
Meyer’s proposed next step for teams considering a use is to test it on 300 to 500 past decisions, examine errors across confidence bands and review a sample of disagreements. If the high-confidence results meet his 95% threshold, he recommends a separate flag and a limited 5% to 10% canary before broader deployment. The article does not announce a product launch or a rollout date for new Jev applications.
For teams applying the framework, the remaining step is to establish whether their current rule fails and what an error costs. Meyer’s account calls for evidence of a failure, a narrow question and a path for uncertain cases before production use.
Key Questions
What is Jev?
Meyer describes Jev as a tool that takes text or JSON with typed questions and returns structured answers, such as probabilities, classifications or scores, for software to use.
How many Jev applications does Meyer say are ready or running?
He says 15 of 24 are ready to build or already running: three are live and 12 are strong fits under his test. Seven need measurement first, and two are poor fits.
What does Meyer recommend checking before deploying Jev?
He recommends testing 300 to 500 past decisions, reviewing disagreements and comparing accuracy by confidence band. His suggested deployment threshold is 95% for the high-confidence band, followed by a limited canary rollout.
Do the reported accuracy figures apply to every Jev use?
No. The 97% to 99% agreement figure is from Meyer’s reported measurement on a 31-topic classification task at confidence of 0.8 or higher. The source does not establish that the result generalizes to other tasks or operators.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
