Skip to content

Scoring

How the score is computed, what it ranks first, and when it is authoritative.

Rust Doctor scores five dimensions independently, then takes a weighted average from 0 to 100. The model is core-v4, and every report names it under audit.score.model.

#A density, not a count

A dimension carries a rate rather than a tally: the distinct sites its findings cover, weighed by severity and by the tier of their rule, discounted by the rate at which the rule was measured wrong, over the amount of code the workspace holds.

site density D  =  Σ (distinct sites × severity weight × tier weight × (1 − noise rate))  ÷  denominator
dimension score =  round(100 · exp(−D / λ))

One diagnostic is one distinct site, whatever it counts in occurrences. A near-duplicate family names its other members through related and weighs one, because it is one place to go and look.

SeverityWeightTierWeight
Error2P08
Warning1P14
Info0P22
P31

An Info finding is shown and costs nothing: a print in a binary target is the program's output, not a defect. A severity the catalog could not resolve weighs nothing and drops the authoritative flag instead, because it is the absence of a measurement rather than a small penalty.

The noise rate is the one the corpus adjudicated for the rule, read through a Beta(1,1) prior, so a rule measured wrong on a third of its sites costs two thirds of what it reported. A rule the corpus never adjudicated is charged whole. The rate is published on every rule the scan ran, under policy.rules[].corpus_noise_basis_points.

#What the density is divided by

ProducerDenominator
clippy::*, rust_doctor::source::*, rust_doctor::structure::*production kilolines
rust_doctor::cargo::*, rust_doctor::repo::*1

Per-site rules divide by the production lines the scan measured, published as audit.production_lines and counted in thousands, so duplicating a workspace changes nothing. Workspace-scoped rules divide by one, because a missing lockfile is not less serious in a large repository.

A workspace under two kilolines is charged against two, so a 120-line crate is not scored on a denominator that makes three findings fatal.

#λ, the calibration

λ is the density that costs a dimension 63 points. A dimension sitting exactly at its λ scores 37, twice its λ scores 14, and half of it scores 61.

Dimensionλ
security4.0
reliability10.0
maintainability3.0
performance4.0
dependencies6.0

Security is the tightest once the tier weight is applied: one P1 security finding per two kilolines already costs the dimension a third. Reliability is the loosest: the correctness lints fire densely on code nobody would call broken. The table is calibrated against a measurement of eighteen public repositories and frozen against it, so a λ moved without a new measurement fails the build rather than republishing every recorded score under a model that never produced it.

#Dimensions and weights

DimensionWeightCategories that feed itRules
Security2.0security6
Reliability1.5correctness, reliability25
Maintainability1.0maintainability16
Performance1.0performance10
Dependencies1.0dependencies5

Each dimension decays from 100 on its own density, and the overall score is their weighted average over a total weight of 6.5. The worst rule tier a dimension carries then caps it: P0 at 20, P1 at 50, P2 at 75. The worst tier anywhere caps the overall score too, P0 at 40 and P1 at 65, so one security finding is visible whatever the rest of the workspace scores.

#What it tells you to fix first

The report ends with up to three rules, and they are not the three most frequent. Each place goes to the rule whose repair, on top of the ones already named, gives the most points back through the ceilings: the score is recomputed without the rule, caps included, so the rule holding the ceiling comes first even when it is one site. audit.score.projected_after_top_three is the value the three are worth together.

A rule the corpus found wrong more often than right is named apart under audit.score.withheld_rule_ids rather than listed last, because naming it would still be telling you to go and change something. The report says so in the terminal too: a rule that is loud in your codebase and absent from the list of what to fix has to be legible without recomputing anything.

The report prints the sample beside the rate wherever it names a rule, (34% noise on 39 sites), because thirty-four percent measured on one site and on thirty-nine are the same number and not the same claim. A rule with no measurement is named as unmeasured rather than shown as a number.

#Authority

A score is published as authoritative only when the scan completed and every pass produced trustworthy diagnostics. If a pass fails, the report carries the failure at its stage and drops the flag rather than presenting a number that quietly means less.

A non-authoritative score prints no share link, and the report names no top three, since ranking what to fix first from an incomplete scan would be guessing.