Corrections log
What we got wrong
Every mistake we find in our own data or wording, with what it affected and what changed. Entries marked “before publishing” were caught before the figure was ever shown publicly — we log them anyway.
· before publishing
Withdrawn versions and archived repositories were labelled “verified once”
- What was wrong
- Our draft labelled both findings as checked by the registry at publish and never re-checked, and described the versions as withdrawn “since” publishing. The registry’s validator doesn’t check withdrawal or archive status at all, and npm, PyPI and GitHub don’t record when a version was deprecated or yanked or when a repository was archived.
- What changed
- Both are now labelled “never checked”, and the wording no longer claims when the change happened. No counts changed.
· before publishing
A finding looked bigger than the total
- What was wrong
- The draft table showed the repository-link finding as 15.0% (of entries that link a GitHub repository) next to a 13.9% total (of all entries). Different bases made a part look larger than the whole.
- What changed
- Every row now uses the same base as the total. Each finding’s rate within its own subset is shown underneath as context, labelled with its denominator.
· before publishing
Withdrawal reasons were mislabelled by a one-off pattern match
- What was wrong
- A quick regular expression produced “237 renamed”: the pattern for “moved” also matched inside “removed”. It also counted 4 false “security” reasons (2 matched the word only inside a URL, 2 were an upsell), left about half the messages unlabelled, and never matched plain “deprecated”.
- What changed
- Replaced by a tested classifier that strips URLs before matching and applies rules in a fixed order. Corrected: 108 entries from 89 publishers point to a renamed or merged package; 14 entries from 12 publishers carry a concrete security notice. Hand-checked 2 random messages per label against the raw npm/PyPI records: 23 of 24 exact (the one difference came from our own message-length cap, since raised). The figures 237, 18 and 8 must not be reused.
· before publishing
GitHub rate limits recorded as missing data
- What was wrong
- During the full catalogue run, GitHub’s hourly quota ran out and 611 of 33,300 records were written with GitHub fields marked as missing — our limit, recorded as if it were the server’s.
- What changed
- The job now waits for the quota to reset, and on resume re-fetches any rate-limited record instead of counting it as done. All 611 were re-fetched.
· before publishing
Repository 404s were silently dropped
- What was wrong
- In the first 380-server sample, a coding error overwrote the “repository not found” result, so all 62 repository 404s were reported as “no data” instead of as findings.
- What changed
- Fixed and covered by an automated test; the sample was discarded and re-run. The full-catalogue figure (14.95% of entries with a GitHub link) sits inside the re-run sample’s 95% interval (12.9–20.4%).