Accuracy
How often JevGate’s findings are right, from findings labeled by reading the code on projects JevGate was never tuned on. These numbers decide what fails the check by default: a rule’s reviews or considers fail it once at least 80% of them were right on those projects, over at least 20 labels (the mature level in configuration).
By rule and level
With the default rules, a check fails only on function-simplification reviews, right 87% (20 of 23); once selected, also on agent-context considers, right 92% (22 of 24). In all, 661 findings were labeled on the unseen projects and 1,525 on the tuned ones.
| Rule | Level | Right on unseen projects | Right on tuned projects | Fails the check by default |
|---|---|---|---|---|
maintainability/file-organization | review | 2 of 5 | 50% (15 of 30) | no |
| consider | 59% (17 of 29) | 44% (19 of 43) | no | |
maintainability/function-simplification | review | 87% (20 of 23) | 83% (57 of 69) | yes |
| consider | 67% (85 of 126) | 75% (147 of 197) | no | |
maintainability/shared-logic | review | 54% (46 of 85) | 75% (181 of 240) | no |
| consider | 59% (76 of 129) | 64% (121 of 190) | no | |
maintainability/hardcoded-values | review | 1 of 8 | 54% (15 of 28) | no |
| consider | 17% (5 of 29) | 56% (32 of 57) | no | |
security/injection | review | 3 of 4 | 84% (81 of 96) | no |
| consider | 5 of 13 | 57% (27 of 47) | no | |
security/sensitive-data | review | 42% (10 of 24) | 64% (41 of 64) | no |
| consider | 0 of 5 | 12 of 15 | no | |
security/unsafe-settings | review | 2 of 4 | 74% (53 of 72) | no |
| consider | 0 of 4 | 76% (16 of 21) | no | |
security/access-control | review | none labeled | 2 of 5 | no |
| consider | none labeled | 5 of 14 | no | |
security/workflows | review | 1 of 1 | 0 of 1 | no |
tests/value | review | 3 of 5 | 5 of 11 | no |
| consider | 1 of 2 | 74% (20 of 27) | no | |
tests/redundancy | review | 1 of 1 | 4 of 4 | no |
| consider | 60% (27 of 45) | 79% (27 of 34) | no | |
tests/laws | not measured | not measured | no | |
documentation/agent-context | consider | 92% (22 of 24) | 94% (64 of 68) | once selected |
documentation/large-docs | consider | 1 of 1 | 1 of 5 | no |
documentation/staleness | consider | 2 of 2 | 14 of 15 | no |
documentation/duplication | consider | 15% (3 of 20) | 23% (5 of 22) | no |
documentation/comments | consider | 54% (39 of 72) | 65% (97 of 150) | no |
Below 20 labels a cell gives the counts without a percentage: a few more labels could move such a share by many points. This is the table this release of JevGate uses; jevgate rules prints its unseen shares, and each finding in the reports says how often its rule and level were right. Each rule’s page gives what it looks at and findings it got wrong.
The table is the ten supported languages’. A finding in a preview language is weighed by that language’s own counts, from 37 projects chosen for those languages and labeled the same way, and never fails the check by default.
How it is measured
- The corpus. Open-source projects of many kinds, from web frameworks and command-line tools to intentionally vulnerable applications, plus the maintainer’s own applications, each pinned at a commit. JevGate runs every rule on each of them.
- Labels. Each review and consider is labeled by reading the code it points at, and the code around it, by the maintainer or by a coding agent following a written labeling guide. Notes are optional by design and are not labeled. A label is kept by the finding’s fingerprint (its rule, file and unit), so it carries over to later versions while the finding stands.
- What counts as right. Right: the claim is true of the code, and acting on it is an improvement a competent maintainer of that kind of project would accept; for a consider, true and worth a look is enough. Wrong: the claim is false (the value is bound, the copies do different work, the test checks real behavior), or acting on it would be wrong or pointless there (an idiom the framework requires, generated code, a design documented beside it). Debatable: competent maintainers would disagree. The shares count a debatable label as not right; counted as right of right and wrong, leaving debatable labels out, they would be higher.
- Unseen and tuned projects. The unseen projects are 11 held out from the start and 14 added later, none used to tune the rules (listed below). The tuned projects are the other labeled projects, without Bend 2 code, which the table leaves out. Only unseen numbers decide what fails the check.
- Examples. The wrong findings on the rule pages come from open-source projects used for tuning only: explaining why an unseen project’s findings were wrong would be tuning on it.
A third set: 27 public projects
After this release’s table was measured, function simplification ran alone on 27 public projects JevGate had never run, three per supported language except Bend 2 (from xh, requests and axios to Dapper, Puma and HikariCP), chosen by language, size and price before any was run. The test asked whether considers on long functions could be a third mature level, with the rule fixed before any finding was read: 80 or more lines (or 50 or more with a split answer’s top level of at least 0.55) had to be right at least 75% of the time on these projects and 80% on all unseen ones. Seven labeling agents labeled 311 findings with one brief and the labeling guide, seeing neither the band nor the probabilities.
- Function-simplification reviews held: right 80 of 93 times (86%) on functions of 50 lines or more, against 20 of 23 in the table; the 21 reviews on shorter functions were not labeled. This supports failing the check on them by default.
- Considers did not. On functions of 80 lines or more they were right 81 of 132 times (61%), and 68% pooled with the earlier unseen projects; outside both bands, 6 of 30 sampled considers were right. The whole consider level on these projects was right about 42% of the time, against the table’s 67%, which leans on the maintainer’s own repositories (78% right there, 65% on public projects): read a function-simplification consider’s share as an upper bound.
These labels are not in the table: the reviews and considers labeled were chosen by the length of their functions, not drawn from all findings.
Why tuned numbers are higher
Each release changed questions and composition until wrong findings on the tuned projects went away; the changelog records each change with its numbers. A change that removes one project’s wrong findings need not carry over to code nobody looked at, which is why the tuned column overstates what a new project sees. The tuned projects also hold 8 of the 9 intentionally vulnerable applications, where security findings are right far more often: on the tuned projects, injection reviews were right 76 of 83 times in those applications and 5 of 13 times in the others.
Limits
- Precision only. The labels say how often a reported finding is right, not what JevGate misses.
- A snapshot. The table is measured on one release’s findings, joined with labels made on earlier ones: 0.25.0’s findings with every rule and tests, replayed from the answer cache, with the shared-logic threshold of 0.28 applied. Each release’s changelog says what moved.
- Other rules beside it. 0.25.0 asked each rule about a function in a request of its own, as a run of this release does when function simplification is the only rule that judges functions, the default. With hardcoded values or a security rule selected, a function’s questions share one request, and a function-simplification finding near a threshold can differ: on 28 labeled projects, its reviews were right 50 times in 55 that way and 49 in 56 apart.
- Not entirely unseen. A few changes before 0.22 came from findings on these projects: in 0.20.0, flysystem’s copies in deprecated code (8 wrong shared-logic findings); in 0.21.0, Online Boutique’s Go modules (9 of 10 copies found between them were wrong) and its connections without TLS (7 reviews), the React Native template’s i18next escaping, and the follow-ups for hardcoded-value and injection considers, which the fresh projects’ first labels pointed to. Since then, changes are fitted on the tuned projects and only checked on these.
- Whose projects. 9 of the 25 unseen projects are the maintainer’s own applications, and 23 of the 24 unseen labels of agent-context considers come from them.
Measure your own
When you accept findings with jevgate baseline, jevgate baseline mark intended|later|wrong PATH:LINE records whether each was right (meant that way, or to fix later) or wrong, and jevgate baseline stats gives each rule’s rate of wrong findings among those marked: the same measure, on your own code. A wrong finding reported with the wrong finding template is how the rules improve.
The unseen projects
- Held out (11): starlette, koa, chi, fd, flysystem, javapoet, DVJA, Symfony demo, gorilla/websocket, tenacity and mdBook.
- Fresh (14): Online Boutique, Refined GitHub, a React Native template, Uniswap v2 core, jaffle-shop, and 9 of the maintainer’s own applications, which are private.