Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Question gallery

Ready-made custom questions for conventions teams ask for and linters cannot check, each measured before it was shipped: asked of the code of real projects, with every finding labeled right or wrong by reading that code.

jevgate rules add swallowed-errors resource-leak   # writes .jevgate/questions/<name>.toml
jevgate check --rule custom --dry-run               # what they would ask and cost, offline
jevgate check --fail-on custom=report               # ask them without failing the gate

jevgate rules add writes each question’s file into .jevgate/questions/, where it becomes a file the project owns: adapt its guidance and paths to the code, and commit it. It says how often the question was right and whether it fails the gate: like any custom question, a review question fails the gate on its reviews and a note never does, so try a review question with --fail-on custom=report first. jevgate rules add is offline and writes the wording the installed version measured; --force restores it over an edited file. Each file is also shown below, and kept in the repository’s gallery/ directory.

QuestionAsked of eachLevelProjectsUnitsFindingsRight
todo-without-ownercommentreview at 0.9579733131 (100%)
swallowed-errorsfunctionreview at 0.8061,5501714 (82%)
resource-leakfunctionreview at 0.8061,2121412 (86%)
thin-handlersrequest handlerreview at 0.80126451312 (92%)
n-plus-onefunctionnote at 0.80111,86475 (71%)

How they were measured

Each question was asked of 6 to 12 projects where it applies, without the built-in questions, with jev-1.13.0 in September 2026. Every finding at the question’s threshold was labeled from the code: right, wrong, or debatable when competent maintainers would disagree. A debatable finding counts as not right. The numbers are for the files as shipped: replayed from the cached answers, the files ask exactly what the measured runs asked.

Any level but note fails the gate once a question is added, so a question ships at review when at least 80% of its findings were right over at least 10 findings, and as a note, whose findings are listed but never fail the gate, from 60%. That is a lower bar than JevGate’s own rules meet before they fail the gate by default, at least 80% right over 20 findings on projects never used to tune them, and three of the four review questions draw most of their right findings from one project each, as their sections say. A question’s first wording was revised at most once, and a project whose findings informed the wording or threshold is counted as tuned; the sections say which. Of the 29 projects, 24 are open source and 5 are the maintainer’s own. Four of them (microblog, laravel-realworld, nest-realworld and bakerydemo) were asked last, of the files as shipped: they added one right and one debatable n-plus-one finding, none in 79 more handlers, and two wrong findings that took a sixth question, global-state, out of the gallery.

The counts are small. They say how often a question is right when it fires, not how much it finds: no one labeled the units it cleared.

Cost

Each unit carries the question’s own text: its question, background and guidance, about 160 to 300 tokens for these. A unit whose source a built-in question already sends, such as a function beside function simplification, adds only that. Asked alone, a function question took 540 to 750 new input tokens per function and the comment question about 450 per comment (by dry run on mdbook and flysystem), so a first run over 1,000 functions costs about $0.03. Reruns, and units that did not change, are answered from the cache.

todo-without-owner

Is a comment a TODO, FIXME, HACK or XXX that names neither a person or team to do the work nor an issue that tracks it? ruff’s TD002 and TD003 check one format of owner in Python; this question reads free-form owners and links, such as @maria or see #123, in every language JevGate parses.

On 7 projects (973 comments), all 31 findings were right: gson’s // TODO: strip wildcards?, shiori’s // FIXME: This only works in local filesystem, fd’s // TODO: support writing raw bytes on unix?, refined-github’s // TODO: Add support for PRs by detecting deferred-content wrappers. The threshold is 0.95, chosen on the first four projects: at 0.80 the question also flagged gson’s // OK: will assume everything is accessible, which holds no marker, and 26 of refined-github’s dated TODOs, such as // TODO [2027-01-01]: Drop after legacy PR files view is removed, which a lint rule there fails once the date passes. They are tracked by that project’s convention, so they were labeled debatable; 0.95 dropped all 27 and 4 right findings. On the three projects not used to choose it (fd, Online Boutique, refined-github), 14 of 14 were right.

At 0.95 a comment is clear only at 0.05 or below, so most comments stay undecided (562 of 973): they never fail the gate, and --verbose lists them. At 0.90, 91 stay undecided and 35 of 43 findings were right, the other 8 being refined-github’s dated TODOs and the comment without a marker. If your project dates its TODOs, add dated TODOs to what guidance calls tracked. Five of the projects were asked only about the files that hold a TODO or FIXME marker.

# todo-without-owner: a TODO or FIXME that names neither an owner nor an issue.
# From JevGate's question gallery, which gives how often it was right on real code:
# https://tech-byte-frontier.github.io/jevgate/question-gallery.html#todo-without-owner
question = "Is this a TODO, FIXME, HACK or XXX comment that names neither a person or team to do the work nor an issue, ticket or link that tracks it?"
background = "An anonymous TODO outlives everyone who remembers it. The convention: each one names an owner or links the issue that tracks it, so it can be scheduled or deleted."
guidance = "Tracked: a person or team, such as `TODO(maria)`, `@maria` or `FIXME(payments):`, or an issue, ticket or link, such as `#123`, `JIRA-88`, `gh-45` or a URL. Not tracked: a TODO with only a description, however detailed. Not a TODO: comments without one of these markers, and the word used in another sense, such as a to-do list feature."
unit = "comment"
threshold = 0.95
level = "review"
next_step = "Name who will do it or link the issue that tracks it, or do it now and delete the comment."

swallowed-errors

Does a function catch or receive an error that means something went wrong, and drop it without logging, returning, rethrowing or reporting it? A linter sees an empty except or _ = err; it cannot tell a failure hidden from an expected case handled on purpose, such as a division by zero that returns 0 or a missing optional file that gives defaults.

On 6 projects (1,550 functions), 14 of 17 findings were right. shiori’s UserConfig::Scan ignores json.Unmarshal’s error and returns nil, so a corrupt stored config loads as defaults; GenerateEbook drops AddImage’s error, so a cover that fails to load leaves an ebook without one; Online Boutique’s getProductByID returns on an error with neither a response nor a log. Wrong: shiori’s importHandler, where every failure is printed and the ignored fmt.Scanln error only keeps the default answer. Debatable: Online Boutique’s placeOrderHandler, which turns unparsable numbers into zeros that validation then rejects, and a webhook sender that returns only whether any webhook succeeded. Nine of the right findings are one of the maintainer’s projects, which catches Exception and continues in its data fetchers.

The first wording also flagged expected cases handled on purpose: 7 of its 17 findings on shiori, django-debug-toolbar and one of the maintainer’s projects were wrong. The guidance now names those cases, and those three projects count as tuned: 4 of 5 right there, 10 of 12 on the others.

# swallowed-errors: an error caught, or returned by a call, and then dropped.
# From JevGate's question gallery, which gives how often it was right on real code:
# https://tech-byte-frontier.github.io/jevgate/question-gallery.html#swallowed-errors
question = "Does this function catch or receive an error that means something went wrong, and then drop it without logging, returning, rethrowing or reporting it?"
background = "An error dropped in silence turns a failure into a wrong result nobody can trace. The convention: every error is handled, passed on or logged, and an error ignored on purpose says why in a comment."
guidance = "A violation: a failure such as an I/O, network, database, parse or decode error, or any exception caught broadly, that is dropped: an empty catch or except block, `except Exception: pass`, `.catch(() => {})`, a Go error assigned to `_` or checked and then ignored, a Rust `let _ =` or `.ok()` that discards it, or a fallback value that hides it without a log. Not a violation: an exception used for an expected case that the catch handles as intended, such as a division by zero returning 0, a lookup that finds nothing returning None or an empty result, or a missing optional file giving defaults; an error that is logged, rethrown, returned, wrapped, reported or shown to the user; a comment beside the ignore that says why it is safe; best-effort work whose failure changes nothing, such as cleanup or reading the terminal width; a function that handles no error."
unit = "function"
threshold = 0.8
level = "review"
next_step = "Log the error, return it, or rethrow it; if ignoring it is right, say why in a comment."

resource-leak

Does a function open a file, connection, cursor, stream or lock that it does not close on every path, errors included, and does not hand to its caller?

On 6 projects (1,212 functions), 12 of 14 findings were right, 9 of them in javavulnlab, an intentionally vulnerable Java application whose servlets never close their JDBC connections. The others: pgweb’s Tunnel::handleConnection, which never closes the remote connection and leaves the local one open when the dial fails; Online Boutique’s email client, which creates a gRPC channel per call and never closes it, and its chatBotHandler, which reads a response body without closing it. Wrong: Online Boutique’s two initTracing functions, which hand the collector connection to the trace exporter for the life of the process. sqlite-utils, websocket and chi had no finding.

# resource-leak: a file, connection, stream or lock not released on every path.
# From JevGate's question gallery, which gives how often it was right on real code:
# https://tech-byte-frontier.github.io/jevgate/question-gallery.html#resource-leak
question = "Does this function open a file, connection, cursor, stream or lock that it does not close or release on every path, including when an error is raised, and does not hand to its caller?"
background = "A resource released only on the happy path leaks when an error is raised midway: file handles, sockets and pooled connections run out under load, and a lock left held stops every other caller."
guidance = "Count: `open()` without `with` or a `finally` that closes it; a connection, cursor, response body, stream or lock acquired, used and released later in the same function with a `return` or a call that can raise in between; Go without `defer x.Close()` or `defer mu.Unlock()`; Java or C# without try-with-resources or `using`. Fine: `with`, `using`, `defer`, try-with-resources and `finally`; Rust values closed when they are dropped; resources returned, stored in a field or passed on for the caller to close; resources a framework opens and closes around a request."
unit = "function"
threshold = 0.8
level = "review"
next_step = "Release it on every path: with, using, defer, try-with-resources or finally."

thin-handlers

Does a request handler do business work itself, such as calculations, rules or several data changes, instead of reading the request, calling a service and building the response? It is asked of the functions in files its paths match: controllers, handlers, routes and views under common names. Change paths to where your handlers live.

On 12 projects (645 handlers), 12 of 13 findings were right, 11 of them in lobsters, whose Rails controllers hold the login rules, moderation records and karma changes (LoginController::login, StoriesController::destroy), and one in linkace, whose single sign-on callback links accounts and sets defaults for new users. Debatable: linkace’s saveAppSettings, mostly input copied onto settings. The Django, Wagtail, ASP.NET, Express, Symfony, Laravel and NestJS projects had none.

# thin-handlers: a request handler that does business work instead of delegating it.
# From JevGate's question gallery, which gives how often it was right on real code:
# https://tech-byte-frontier.github.io/jevgate/question-gallery.html#thin-handlers
question = "Does this request handler do business work itself, such as calculations, rules or several data changes, instead of reading the request, calling a service or model and building the response?"
background = "Business rules inside a handler cannot be reused from jobs, scripts or other endpoints, and are tested only through HTTP. The convention: controllers stay thin and delegate to services, models or domain functions."
guidance = "A handler is a controller action, route handler, view function or endpoint. Fine: reading and validating input, one or two calls to services, models or repositories, authorization checks, pagination, choosing a status, redirect or template, and building the response. Count: prices, limits, eligibility or state transitions computed in the handler; loops that change several records; queries built step by step; several services orchestrated with business decisions between them; more than a few lines of logic that another entry point would need too. A function that is not a request handler is not a violation."
unit = "function"
paths = ["**/controllers/**", "**/*Controller.*", "**/*_controller.*", "**/*.controller.*", "**/handlers/**", "**/routes/**", "**/views.py", "**/views/**/*.py"]
threshold = 0.8
level = "review"
next_step = "Move the business logic into a service or model method the handler calls."

n-plus-one

Does a function run a database query or a network call once per item of a loop, where one query or call could handle all the items?

On 11 projects (1,864 functions), 5 of 7 findings were right: linkace’s HTML and CSV exports, which query each link’s tags (and lists) over all of a user’s links; lobsters’ MessagesController::batch_delete, a query and a save per selected message; spring-realworld’s createNew and laravel-realworld’s ArticleController::store, a lookup and an insert per tag of an article. Debatable: linkace’s getOldTaxonomyItems, one lookup per item of a form shown again after a validation error, a handful at most, and bakerydemo’s random-data command, one insert per item, where bulk_create would skip the save() its models may rely on. It is a note: 71% right, below the 80% a question needs here to fail the gate.

# n-plus-one: a query or network call once per item of a loop, where one call would do.
# From JevGate's question gallery, which gives how often it was right on real code:
# https://tech-byte-frontier.github.io/jevgate/question-gallery.html#n-plus-one
question = "Does this function run a database query or a network call once per item inside a loop, where one query or call could handle all the items?"
background = "One query per item (the N+1 problem) is fast with ten rows and slow with ten thousand. It is the most common performance bug in web applications, and it stays invisible until it meets production data."
guidance = "Count: a query, ORM lookup, lazily loaded association, `fetch`, HTTP client call or RPC inside a `for`, `while`, `each`, `map`, `forEach` or comprehension, once per element, when the elements could be loaded or sent together (an `IN` query, a join, eager loading, a bulk insert or a batch endpoint). Fine: a loop over a handful of known items; pages of a paginated API fetched in order; retries; calls that must run one at a time to be correct; work that already batches; a loop that only builds one query sent after it; calls on objects already in memory."
unit = "function"
threshold = 0.8
level = "note"
next_step = "Load or send the items together: one query with IN or a join, eager loading, or a bulk call."

Measured and left out

Ten more questions were measured the same way and are not shipped: under 60% right, or too few findings to measure.

QuestionAskedFindingsRightWhy it is left out
global-stateDoes this function both read and change a global, module-level or static variable?6 on 113Two Laravel model factories that count a static timestamp offset down on purpose, in seed data, and a main that only assigns a start time. The three right: functions that replace a module’s database engine through global, and a loop that changes six module globals.
log-and-rethrowDoes this function log an error and then also return or rethrow it?7 on 7 projects4A gRPC handler that logs and passes the error to its callback, which answers the client; two logging decorators whose job is to log what passes through (debatable).
flaky-testCan this test pass or fail from one run to the next with no code change?11 on 66Tests asserting that many random draws are not all equal, which fail with a negligible chance, and timing margins of seconds.
test-name-mismatchDoes this test’s name promise what its assertions do not check?17 on 67Go test names that name the unit under test, type-level tests, a test whose check is the race detector; 0 of 5 right on projects it was not tuned on.
debug-outputDoes this function print debugging output left over from investigating a bug?13 on 75Call traces of a service whose only logging is Console.WriteLine. Linters also catch most stray prints (no-console, ruff’s T201, clippy’s dbg_macro).
untranslated-textDoes this function put user-facing text into the interface as a literal instead of through the translation function?18 on 24Console-command output and example placeholders such as R$ and L, kg read as interface text. Its first wording was right 25 times in 42 on four other projects, with 15 debatable demo screens.
money-in-floatDoes this function store or compute money as a binary float?50 on 96Percentages, quantities, fee rates, simulations and display code read as money.
stale-commentDoes a comment in this function say something the code does not do?4 on 61Comments that describe what a callee or an override does.
undocumented-returnDoes this public function return a special value for failure or not found without saying so?1 on 51One finding in 1,565 functions of five libraries: too few to measure.
hidden-side-effectsDoes this function’s name promise a lookup, check or conversion while it writes or sends?0 on 6No finding in 916 functions once HTTP handlers named get were excluded; the one right finding of its first wording was lost.