
A few years ago, I opened a post about Betterer with a line you've probably heard before:
If we can't lint it, then we can't enforce it
Today, I would go one step further: if we can't enforce it, then it isn't a best practice, it's a suggestion!
Best practices used to spread through code review: someone left a comment, and the author remembered it. But agents write most of the code I review now, and an agent doesn't remember last week's comment. It writes the same pattern again, and I leave the same comment again.
At that point, repeating best practices on every single review wastes everyone's time. You could fix the pattern everywhere so the next agent picks up the right one from the surrounding code, but even then, there's no guarantee it will.
Our monorepo has an AI-REVIEW.md full of best practices extracted from past pull request reviews. One entry looks like this:
// Bad - manual useState + useCallback
const [isDismissed, setIsDismissed] = useState(false);
const dismiss = useCallback(() => setIsDismissed(true), []);
// Good - useBooleanState provides memoized handlers
const [isDismissed, toggle] = useBooleanState(false);
const dismiss = toggle.on;
An agent may or may not read that document, and an AI reviewer may flag the pattern on one pull request and miss it on the next. I could write an AST rule for that one, exceptions included. Then I'd need one for every other entry, and some, like "Name test cases to reflect their purpose", can't be expressed as syntax at all.
Jev is a TypeSafe model that doesn't generate text: give it some state and a typed question, and it returns a typed answer with a calibrated probability.
Think Hotdog, Not Hotdog from Silicon Valley, except you write the question and get back the probability that the answer is yes. A lint rule is the same app: violation, not violation.

oxlint-plugin-jev, by Robert Soriano, turns each rule into a question, a target (a function, a call, a JSX element, or a file), and a cutoff. We already run Oxlint, so adding the plugin was the easy part.
Let's write the rule. As the plugin's README puts it: "The wording of the question is the rule." Here's ours, pinned to jev-1.13.0:
{
"id": "prefer-use-boolean-state",
"target": "file",
"question": "Does this file contain a violation of our best practice #41: ...",
"cutoff": 0.9,
"location": {
"question": "Select the variable declaration line for the React useState call ...",
"cutoff": 0.75
}
}
The question (abridged) reads like the best practice itself, with every exception spelled out:
Does this file contain a violation of our best practice #41: prefer useBooleanState from @/hooks/useBooleanState over React useState initialized with a boolean plus a memoized setter-only callback?
A qualifying state is a React useState boolean paired with useCallback whose entire body unconditionally calls that state's setter with the literal true or false.
Resolve React import aliases. Exclude a candidate state if its callbacks perform additional work or conditional logic, it can also be null or undefined, it uses a lazy initializer or functional updater, or its setter is passed elsewhere.
If no candidate clearly qualifies, answer no.
The high cutoff and the last sentence are there on purpose, we would rather miss a violation than have a check that cries wolf and gets ignored after a week. And since the wording is the rule, it has tests: labeled fixtures, including near-misses, that run against the real model whenever the question changes.
Unfortunately, with target: "file", the plugin reported every finding on line 1, which isn't super helpful on a 400-line component.
So I added an optional location question. The plugin sends the file with numbered lines, and asks Jev to select one:
L3| const [open, setOpen] = useState(false);
L4| const close = useCallback(() => setOpen(false), []);
Each line is an option in a multiple-choice question, plus unknown. If we ask the model for a line number, it could make one up, but if it has to pick one from the list, the answer is always a real line, and it comes with a probability. An uncertain answer keeps the file-level diagnostic. It was released as part of 0.1.3.
Oxlint's github formatter already prints workflow commands that GitHub turns into annotations, but its message is the model, the score, and the whole question. A small wrapper replaces it with a short explanation and a link to the guideline:

We say suspects on purpose, Jev only gives us a probability, so the annotation shouldn't sound more sure than the model is.
Over our first 175 runs, Jev cost us a total of $0.35, about a fifth of a cent per run! Jev only bills input tokens, at $0.042 per million.
Over the same runs, the workflow used about 167 runner minutes, roughly $0.67 at our runner's per-minute rate. Running GitHub Actions ended up being more expensive than Jev itself!
Agents iterate until the checks pass. The annotation gives agents the line to fix and a link to the guideline, so I don't have to leave the same comment again.
With Jev, we can write a rule like this one in English, including the exceptions we already documented.
What's the first entry in your best practices document you'd turn into a rule?