Skip to content

Corrections log

What we got wrong

Every mistake we find in our own data or wording, with what it affected and what changed. Entries marked “before publishing” were caught before the figure was ever shown publicly — we log them anyway.

  1. · before publishing

    Withdrawn versions and archived repositories were labelled “verified once”

    What was wrong
    Our draft labelled both findings as checked by the registry at publish and never re-checked, and described the versions as withdrawn “since” publishing. The registry’s validator doesn’t check withdrawal or archive status at all, and npm, PyPI and GitHub don’t record when a version was deprecated or yanked or when a repository was archived.
    What changed
    Both are now labelled “never checked”, and the wording no longer claims when the change happened. No counts changed.
  2. · before publishing

    A finding looked bigger than the total

    What was wrong
    The draft table showed the repository-link finding as 15.0% (of entries that link a GitHub repository) next to a 13.9% total (of all entries). Different bases made a part look larger than the whole.
    What changed
    Every row now uses the same base as the total. Each finding’s rate within its own subset is shown underneath as context, labelled with its denominator.
  3. · before publishing

    Withdrawal reasons were mislabelled by a one-off pattern match

    What was wrong
    A quick regular expression produced “237 renamed”: the pattern for “moved” also matched inside “removed”. It also counted 4 false “security” reasons (2 matched the word only inside a URL, 2 were an upsell), left about half the messages unlabelled, and never matched plain “deprecated”.
    What changed
    Replaced by a tested classifier that strips URLs before matching and applies rules in a fixed order. Corrected: 108 entries from 89 publishers point to a renamed or merged package; 14 entries from 12 publishers carry a concrete security notice. Hand-checked 2 random messages per label against the raw npm/PyPI records: 23 of 24 exact (the one difference came from our own message-length cap, since raised). The figures 237, 18 and 8 must not be reused.
  4. · before publishing

    GitHub rate limits recorded as missing data

    What was wrong
    During the full catalogue run, GitHub’s hourly quota ran out and 611 of 33,300 records were written with GitHub fields marked as missing — our limit, recorded as if it were the server’s.
    What changed
    The job now waits for the quota to reset, and on resume re-fetches any rate-limited record instead of counting it as done. All 611 were re-fetched.
  5. · before publishing

    Repository 404s were silently dropped

    What was wrong
    In the first 380-server sample, a coding error overwrote the “repository not found” result, so all 62 repository 404s were reported as “no data” instead of as findings.
    What changed
    Fixed and covered by an automated test; the sample was discarded and re-run. The full-catalogue figure (14.95% of entries with a GitHub link) sits inside the re-run sample’s 95% interval (12.9–20.4%).