npm Killed 101,885 of 101,886 Leaked Tokens. Postgres Killed 1,520 of 12,985.

Truffle Security published "GitHub Repos Exposed 543,699 Credentials. Nobody Revoked Them." on 29 September 2026. The scan covered 224,553,295 public GitHub repositories and 58,467,468,698 files — not a crawl of their own, but The Stack v3, the public code corpus assembled to train large language models, whose crawl closed on 7 August 2025.

Between 27 and 28 July 2026 they tested each candidate credential against the service that issued it. 543,699 came back live, drawn from 1,103,438 separate exposures. That is eleven months after the crawl closed and, for most of them, years after the commit. It is more than double the 221,303 working credentials the same team found in 7.6 petabytes of Hugging Face training data.

The volume is the least interesting part. The report's real contribution is that it isolates which control actually keeps a leaked credential from mattering — and it is not the one the industry has spent three years deploying.

Push protection works, and it is not the variable that matters

GitHub made secret-scanning alerts free for all public repositories on 28 February 2023, brought push protection to general availability for public repos in May 2023, and turned it on by default on 29 February 2024. Truffle split its findings by which control was in place on the day each credential leaked:

  • 245,959 (45.2%) predate free alerts.
  • 97,897 (18.0%) leaked while scanning was free and push protection was one setting away — and were committed anyway.
  • 199,843 (36.8%) landed after the block became the default, and were still answering to their providers more than two years later.

Truffle is fair about this: the block does work where it applies, roughly halving the rate at which recognised credential shapes reach public code. The problem is coverage. Push protection blocks only patterns precise enough to be safe to block on — more than 200 token types from over 180 partner providers. 51.8% of every live credential in the corpus is a shape a default-configured public repository will accept.

The three largest live categories divide exactly along that line. Google Cloud service account material leads at 69,041 and is covered. MongoDB connection strings at 51,067 are not, and neither are the 33,343 live Google API keys. Connection strings and private keys are classed as generic patterns: off by default, blocking only if an organisation deliberately enables them.

The Gemini key problem is a design decision, not an oversight

Google API keys sit on GitHub's pattern list explicitly marked as not push protected, and google_gemini_api_key appears as its own entry with the same answer. The reason is structural rather than negligent: a Gemini key is a billable credential attached to a model endpoint, and a Maps key carrying the identical AIzaSy prefix is designed to ship in a web page. One regex cannot separate them, so nothing is blocked, and the keys that matter ride in alongside the ones that do not.

Truffle found 31,374 live Gemini keys with a median leak date of February 2025. That entire population is younger than push protection — every one of them was committed while the default block was already on, to a credential class the block deliberately does not cover.

For anyone running AI infrastructure this is the practical takeaway of the whole report. A leaked Gemini key is a billable inference credential with no kill switch in front of it, and the denial-of-wallet economics we covered in the x47.c botnet's AI credit-burning campaign are exactly what 31,374 of them are exposed to.

The finding that reframes the problem: revocation, not prevention

Truffle went back to everything the scan matched — not just what verified — and kept only values that unambiguously identify themselves: a vendor prefix such as ghp_ or npm_, a connection string with a password, a PEM private key block, a service account JSON. That leaves 4,941,427 unambiguous credentials, of which 1,652,201 have a trustworthy liveness rate once Google API keys and private keys are excluded for methodological reasons the report explains carefully.

Sorted by whether the issuer runs an automated kill switch, the split is brutal:

  • npm tokens: 101,886 committed, 1 still live.
  • GitHub tokens: 73,048 committed, 260 still live (0.78% in 2025).
  • Hugging Face tokens: 30,437 committed, 15 still live.
  • Slack tokens: 8,903 committed, 198 live (2%).
  • AWS access keys: 82,411 committed, 6,819 live (8% overall).
  • Google Cloud service accounts: 126,963 committed, 69,041 live (54%).
  • SendGrid keys: 22,800 committed, 9,189 live (40%).
  • MySQL connection strings: 2,421 committed, 1,806 live (75%).
  • Postgres connection strings: 12,985 committed, 11,465 live (88%).

Three orders of magnitude separate npm from Postgres, and the dividing line is not credential sensitivity, commit recency, or whether push protection covers the pattern. It is whether somebody automatically revokes the token when it surfaces. GitHub runs a secret scanning partner program that forwards leaked tokens to the issuing provider — but the program does not require partners to revoke anything. What happens next is the provider's choice, which is why partner-provider credentials remain live for years, including an AWS key last touched in November 2009.

There is a discipline lesson buried in Truffle's own caveats worth copying. MongoDB is the second-largest live category and is deliberately excluded from every survival rate, because the detector only reports a URI it managed to connect to — a dead Atlas string never enters the dataset, so a 100% survival rate across 51,067 samples is a measurement artefact, not a policy finding. Splitting connection strings by engine also surfaced that Redis URLs survive at only 2.9%, because the ones people commit overwhelmingly point at localhost. Researchers who report the numbers they cannot trust are rarer than they should be.

The ages, and why the corpus makes them conservative

The median live credential had been sitting in a public default branch for 784 days. The 90th percentile is 6.3 years. The oldest — database credentials in an Erlang web server config, last touched 13 June 2009 — was still valid 16.1 years later. Behind it: an FTP login in a GPS logger's C source from September 2009, copied into 62 repositories; an AWS key in a Rails S3 config from November 2009. Truffle does not name the repositories, because the credentials still work.

Two properties of The Stack v3 make these figures conservative rather than inflated. The corpus keeps only the default branch as it stood at crawl time, so there is no commit history to dig through — every secret that was committed and later removed is invisible here. And dating relies on each file's last-modification timestamp, which means a file edited after the credential was added reads as newer than the leak. The errors run toward younger, not older.

The density trend runs the wrong way too. Live credentials per million files was 3.72 in 2014, 9.54 in 2022, 11.09 in 2023, and 11.62 in 2025 — the highest in the dataset, measured on the seven months before the crawl closed, two years into default-on push protection.

What to take from this

  • Grade your providers by revocation, not by scanning. The 0.001%-to-88% spread is a provider-behaviour ranking. When choosing where credentials for a new service live, ask whether the issuer automatically kills leaked tokens — npm, GitHub and Hugging Face effectively do; database engines and most SaaS vendors do not.
  • Turn on generic-pattern push protection explicitly. Connection strings and private keys — the two shapes with the worst survival rates — are off by default. That is one organisation-level setting standing between you and the 51.8%.
  • Treat AI API keys as unprotected by construction. Gemini keys are explicitly not push protected and cannot be, because the pattern is shared with keys meant to be public. Keep them out of code entirely; no platform control is coming.
  • Assume anything ever public is still live. The corpus is a training snapshot from August 2025 that was fully readable to anyone for the eleven months before verification. Push-protection era or not, rotation is the only control that terminates exposure — a point this site made from the other direction in the ToxSec work on hardcoded secrets in AI-generated code.
  • Public code is training data, and training data is a scanning surface. These credentials were found in a corpus assembled to train models, not by crawling GitHub. Anyone with a copy of The Stack v3 — it is public — had the same 543,699 keys.

The uncomfortable summary is that the industry deployed prevention and declared the problem handled. Prevention halved the inflow for the shapes it recognises, covers just under half of what is live, and does nothing whatsoever for the 543,699 credentials already sitting in public branches with a median age over two years. The control that actually ends exposure is the one nobody is contractually obliged to operate.

Sources: