Scoring
How the score is computed, what it ranks first, and when it is authoritative.
Rust Doctor scores five dimensions independently, then takes a weighted
average from 0 to 100. The model is core-v4, and every report names it under
audit.score.model.
#A density, not a count
A dimension carries a rate rather than a tally: the distinct sites its findings cover, weighed by severity and by the tier of their rule, discounted by the rate at which the rule was measured wrong, over the amount of code the workspace holds.
site density D = Σ (distinct sites × severity weight × tier weight × (1 − noise rate)) ÷ denominator
dimension score = round(100 · exp(−D / λ))One diagnostic is one distinct site, whatever it counts in occurrences. A
near-duplicate family names its other members through related and weighs one,
because it is one place to go and look.
| Severity | Weight | Tier | Weight |
|---|---|---|---|
| Error | 2 | P0 | 8 |
| Warning | 1 | P1 | 4 |
| Info | 0 | P2 | 2 |
P3 | 1 |
An Info finding is shown and costs nothing: a print in a binary target is
the program's output, not a defect. A severity the catalog could not resolve
weighs nothing and drops the authoritative flag instead, because it is the
absence of a measurement rather than a small penalty.
The noise rate is the one the corpus adjudicated for the rule, read
through a Beta(1,1) prior, so a rule measured wrong on a third of its sites
costs two thirds of what it reported. A rule the corpus never adjudicated is
charged whole. The rate is published on every rule the scan ran, under
policy.rules[].corpus_noise_basis_points.
#What the density is divided by
| Producer | Denominator |
|---|---|
clippy::*, rust_doctor::source::*, rust_doctor::structure::* | production kilolines |
rust_doctor::cargo::*, rust_doctor::repo::* | 1 |
Per-site rules divide by the production lines the scan measured, published as
audit.production_lines and counted in thousands, so duplicating a workspace
changes nothing. Workspace-scoped rules divide by one, because a missing
lockfile is not less serious in a large repository.
A workspace under two kilolines is charged against two, so a 120-line crate is not scored on a denominator that makes three findings fatal.
#λ, the calibration
λ is the density that costs a dimension 63 points. A dimension sitting exactly at its λ scores 37, twice its λ scores 14, and half of it scores 61.
| Dimension | λ |
|---|---|
security | 4.0 |
reliability | 10.0 |
maintainability | 3.0 |
performance | 4.0 |
dependencies | 6.0 |
Security is the tightest once the tier weight is applied: one P1 security
finding per two kilolines already costs the dimension a third. Reliability is
the loosest: the correctness lints fire densely on code nobody would call
broken. The table is calibrated against a
measurement of eighteen public repositories and frozen against it, so a λ
moved without a new measurement fails the build rather than republishing every
recorded score under a model that never produced it.
#Dimensions and weights
| Dimension | Weight | Categories that feed it | Rules |
|---|---|---|---|
| Security | 2.0 | security | 6 |
| Reliability | 1.5 | correctness, reliability | 25 |
| Maintainability | 1.0 | maintainability | 16 |
| Performance | 1.0 | performance | 10 |
| Dependencies | 1.0 | dependencies | 5 |
Each dimension decays from 100 on its own density, and the overall score is
their weighted average over a total weight of 6.5. The worst rule tier a
dimension carries then caps it: P0 at 20, P1 at 50, P2 at 75. The worst
tier anywhere caps the overall score too, P0 at 40 and P1 at 65, so one
security finding is visible whatever the rest of the workspace scores.
#What it tells you to fix first
The report ends with up to three rules, and they are not the three most
frequent. Each place goes to the rule whose repair, on top of the ones already
named, gives the most points back through the ceilings: the score is
recomputed without the rule, caps included, so the rule holding the ceiling
comes first even when it is one site. audit.score.projected_after_top_three
is the value the three are worth together.
A rule the corpus found wrong more often than right is named apart under
audit.score.withheld_rule_ids rather than listed last, because naming it
would still be telling you to go and change something. The report says so in
the terminal too: a rule that is loud in your codebase and absent from the
list of what to fix has to be legible without recomputing anything.
The report prints the sample beside the rate wherever it names a rule,
(34% noise on 39 sites), because thirty-four percent measured on one site and
on thirty-nine are the same number and not the same claim. A rule with no
measurement is named as unmeasured rather than shown as a number.
#Authority
A score is published as authoritative only when the scan completed and every pass produced trustworthy diagnostics. If a pass fails, the report carries the failure at its stage and drops the flag rather than presenting a number that quietly means less.
A non-authoritative score prints no share link, and the report names no top three, since ranking what to fix first from an incomplete scan would be guessing.