npm supply-chain hardening · proof-gated removal · CLI · CI-native
ChainStrip
“Perfection is achieved when there is nothing left to take away.”— Antoine de Saint-Exupéry
New approach to npm supply-chain security: do not ship the vulnerable stuff. Delete it.
That's what ChainStrip does: it removes files and functions from your dependency tree and proves that the removal is safe.
If the package is vulnerable it could be either removed or flagged for update, along with the proof that your code actually reaches it. If the package is not considered vulnerable, it could be in the future, so removing it entirely or partially reduces the attack surface.
Your test suite may not be perfect and have gaps, so the unexercised code paths will be held pinned due to lack of evidence and ChainStrip will work with you to enhance your test suite.
not in what you ship — 56 of 75
- 55 not shipped package/version never made it into the hardened artifact
-
1
eliminated
GHSA-7f6v, a HIGH stored XSS: the sink
renderDocsHtml()pruned with its file; build + suite pass without it (elimination)
still in your build — 19 of 75
- 14 reached present and on a live path: patch these (2 CRIT · 5 HIGH · 5 MOD · 2 LOW)
- 5 unresolved advisory too coarse to map to a symbol; it stays on the books as a gap
start-ui-web (BearStudio's actively maintained React starter), commit 81d5424 of 2026-07-18 · run of 2026-07-23. Its stock pnpm install includes `@orpc/openapi` 1.13.4, unpatched for the XSS above. The elimination happens because start-ui-web never imports the vulnerable plugin. Your mileage may vary due to your dependency tree and your test suite.
or Book on calendar ↗ — pick a slot directly, skip the email round-trip.
What we can show in 15 minutes (or less): walkthrough of an implementation of ChainStrip for an OSS project, analysis of your lockfile without access to your codebase, or custom analysis of an OSS project which is close to yours in structure and frameworks used, if time permits.
Delivery is white glove: we integrate ChainStrip into your GitHub or GitLab CI/CD pipeline ourselves, free of charge.
The problem
A modern app installs hundreds of packages; the attack surface is everything they ship, not just what you call.
details
A modern JavaScript application installs tens or hundreds of npm packages and only uses a small fraction of code from them. But 100% of this code could be used by a malicious actor to mount an attack on your application if they find a way to execute it.
How does one protect themselves from this threat? Standard tools answer many
questions, but not this. Lockfiles pin versions, hoping that if the app
wasn’t compromised now it won’t be in the future. npm audit lists the CVEs
but neither provides any remedy nor tells you if you even need to do
anything. Reachability scanners come closer, they statically trace your
code and suggest which CVEs are reachable and which are not. But they don’t
prove it and don’t try to do anything about it. Tree shakers optimize the
client bundle, leaving server and development code aside, using them for
security hardening is gambling.
The attacks, meanwhile, keep coming, trying to force your endpoints to call something they are not supposed to call, others run scripts on install and sometimes even run code when you just import the infected modules without even running them. Everything below happened just this year:
-
Miasma & Hades wave Jun 2026
A week-long run of Shai-Hulud-descended waves — ~448 malicious npm and PyPI artifacts, zero CVEs. It poisoned 32
@redhat-cloud-servicespackages (~117K weekly downloads), a broad npm batch including@vapi-ai/server-sdk, then backdoored Microsoft GitHub repos — GitHub auto-disabled 73 of them in under two minutes. The payload ran from apreinstallscript (one wave hid it inbinding.gypto dodge install-script scanners), dropping an obfuscated credential stealer aimed at CI and developer machines.ChainStripChainStrip's release artifacts are unpacked into
node_modules, nevernpm install-ed — so nothing fires at install time: not a lifecycle script, and not a native build from abinding.gyp, the trick that wave used to slip past scanners that only grep for install-script fields. And the build consumes a hash-pinned, gated artifact, so the drifted patch releases that delivered every wave never resolve into it to begin with.The payload never runs: nothing executes at install time, and the poisoned releases never resolve into the build.
-
IronWorm May–Jun 2026
JFrog's "rustier cousin" of Shai-Hulud: ~37 packages in the Arweave/WeaveDB ecosystem, republished from one stolen maintainer account and cleaned within a day. A
preinstallhook launched a 976 KB native Rust binary — hidden in atools/setupdirectory, never imported — carrying an eBPF rootkit and Tor C2 to steal environment variables, cloud and AI-tool keys, and crypto-wallet seeds. It self-propagated by backdating commits and republishing through stolen credentials.ChainStripThe binary sits in a directory nothing imports, so additive extraction never pulls it into the artifact; the
preinstallhook that launches it is stripped; and pinning blocks the minor-version bumps it spread through.Blocked three ways — not shipped, not run at install, not pulled in by a pin.
-
node-ipc May 2026
Distinct from the 2022 protestware of the same name: on May 14, 2026 an attacker re-registered the lapsed domain behind the maintainer's email, took the account, and published three versions of node-ipc (~10M weekly downloads). An 80 KB obfuscated payload appended to
node-ipc.cjsfired on everyrequire("node-ipc")— no install script — harvesting 90+ categories of cloud, CI and SSH secrets; one version was hash-gated to a single targeted entry point.ChainStripChainStrip can't strip this one — code that runs on
requireof a package you use is part of your retained surface. But the poisoned versions never resolve into a pinned artifact, and evaluating the bump flags itretained-surface-changed: an 80 KB obfuscated blob appended tonode-ipc.cjsis a glaring diff on a module you call, forcing review instead of a blind upgrade.Not removable, still stopped: the bump lands in review instead of shipping blind.
-
axios Mar 2026
The maintainer account behind axios (~70M+ weekly downloads) was compromised and two patch versions published, live about three hours. The malware was not in axios at all: a phantom dependency,
plain-crypto-js, that axios never imports — itspostinstalldropped a cross-platform RAT, then deleted itself. Patch bumps over the safe versions meant^/~ranges resolved straight to the poisoned build.ChainStripA dependency axios never imports is outside its used closure, so extraction leaves it on the floor — and even installed, its
postinstallis stripped and the drifted version never resolves against a pinned artifact.The RAT's package was never in the used closure and never on the install path. There is nothing to ship.
A scanner scans the dependencies, returns the list of vulnerabilities and washes hands, it’s your business now. Reachability tool goes further, it runs the static analysis and guesses the outcome. ChainStrip goes much deeper: it actively removes dead code, proves that removal is safe using your build and test suite and stands guard on every update to prevent the dead code and vulnerabilities from sneaking back. It doesn’t do penetration testing, doesn’t replay the attack.
chainstrip scoreFind it
First, see what you are carrying: one command on your lockfile, no repo access, every package graded A–F, worst first.
details
You can try ChainStrip very light - just run it against your lockfile (npm,
pnpm, yarn) or node_modules folder. You'll get the list of CVEs in your
dependency tree and a score A-F for each package, ranked worst first.
Score is the sum of two variables: CVE severity and surface. Surface shows
what compromised package can do to your system: read files, open network
connections, install scripts or binaries, run arbitrary code through eval
and dynamic require, etc. This is not exactly the famous "blast radius"
because blast radius also requires reachability analysis and this is what
the rest of this page is devoted to. Each scorecard explains the score
naming each signal individually. Since score depends only on
package@version it can be cached across multiple repositories/services.
$ chainstrip score --dir ./node_modules 132 packages — A 96 · B 20 · C 12 · D 3 · F 1 38 with known CVEs · 11 with install scripts · 4 with native code · 4,554 files · 25.1MB as installed F axios@1.6.2 29 CVEs: 13 HIGH, 15 MODERATE, 1 LOW surface: network size: 81 files · 1.7MB D lodash@4.17.21 3 CVEs: 1 HIGH, 2 MODERATE surface: process.binding · fs-write size: 1,054 files · 1.3MB A chalk@5.3.0 no known CVEs surface: none size: 12 files · 42.7KB
chainstrip runRemove what you can
Then cut the code you don't use, usually most of it; your suite and build prove it. A CVE in that code goes with it.
details
Once ChainStrip has analyzed and rewritten your tree every CVE gets one of
four verdicts, sorted by "strength", a measure of how much evidence caused
each one. eliminated is the strongest, it means "total kill": the CVE was
found, but the static and dynamic analyses proved that it could be removed
from the code completely and it's gone. The proof is the green test suite
and production build. The case study elimination from the top of the page
is dissected below.
| verdict | meaning | strength |
|---|---|---|
eliminated |
vulnerable symbol lexically dead, removed from the artifact; build + full suite pass without it | issued only with proof of death |
reached |
vulnerable code present and on a live path: your app, its build, or the dependency's own internals | real exposure; patch it |
unresolved |
advisory too coarse to map to a symbol (e.g. “affects Next.js”) | stays an open item in every report |
not shipped |
affected package/version isn't in the artifact; most advisories live in dev- and build-only packages that never enter the production closure | free win from extraction |
Two eliminations, and the seven still held
The example at the top of the page is real (and you can read
the relevant case study). start-ui-web ships
@orpc/openapi 1.13.4; the fix for GHSA-7f6v, a HIGH stored XSS in oRPC's
opt-in reference-docs plugin, landed upstream only in 1.13.9. Since
start-ui-web never loads this plugin ChainStrip removed its code completely
and the vulnerable symbol from the advisory, renderDocsHtml(), is proven
to be never called. During the same elimination phase 7 known advisories
for better-auth module were identified and classified as "reached",
because the better-auth resolves modules dynamically, has to be shipped
unaltered and this could become a potential entry point. In the 3,028 npm
advisories with locatable fix commits about 2% can be classified like this
one.
A bigger case study comes with a similar verdict
(the dated claim is also there). Cal.diy, an OSS version of Cal.com SaaS,
carries 189 advisories, and ChainStrip completely eliminates one: glob's
HIGH advisory lives in the package's command-line entry file, a file
Cal.diy never imports and as such was removed with tests and build green.
7 advisories are reachable, but never executed, so ChainStrip dynamic
analysis engine can't substantiate the removal. ChainStrip has a "risky"
aggressive mode which overrides this behavior and our study of this very
case proved it unsound: tests and build stay green, but a production run
would crash if an error path used one of the reachable-but-unexecuted
functions.
Won't this break my app?
--trim-at-risk.The suite is no longer the ceiling
Every hold is one piece of missing evidence away from a removal, and the
ceiling moves from both sides: coverage --prompt names the exact suite
tests that would raise it, and ChainStrip gathers execution evidence from
three places the suite can't reach — probes it generates to make your app
run unproven code, an end-to-end pass against the served production build,
and browser-side coverage (all below). Held code graduates to
removed as the evidence lands.
"Eliminated" means removed from your artifact and proven by your gate; it does not mean "verified safe in production." The closest thing to that claim is an opt-in counter compiled into the shipped artifact that records, locally in your own infrastructure, whether production ever tried to call a removed function.
chainstrip updateKeep it out
Removed stays removed: every dependency change hits a one-second fingerprint check, and risky bumps land in review.
details
Dependency changes become classified diffs your CI can block
Run `chainstrip update` with no arguments and it becomes a PR gate: it classifies every dependency change in the diff — version bumps, brand-new dependencies, and new call-paths your code opens into existing ones — by building and validating a candidate artifact and diffing the retained surface, a diff small enough for a human to actually review. Blocking is policy-driven: an in-policy finding sets a non-zero exit so CI fails the check, and a reviewer waives it with a committed, fingerprint-bound approval that re-blocks the moment that dependency changes again. Compare that to the usual blind N-day version hold: instead of waiting and hoping, you learn in about an hour (measured on Cal.diy's ~230 runtime dependencies) whether the release added a capability or changed code you run — and that cost is paid only on PRs that change dependencies or usage; every other PR stays the one-second fingerprint check. Capability changes are caught from two directions. The new code is scanned for gained capability markers (network, spawn, filesystem writes, environment and secret reads), so a dormant payload is flagged without ever firing. And during validation, a runtime layer instruments the actual built-in surface and attributes every filesystem, network, DNS, and child-process call to the dependency that made it. Obfuscation is handled as its own signal: packing can hide what code does, but it can't hide that the code is packed, and a bump that puts packed code onto a path you run is a blockable finding in itself. The flag points a reviewer at hidden code you execute; it does not pronounce malware.
| classification | meaning | what you do |
|---|---|---|
no-usage-impact |
nothing your product executes changed | merge the bump |
retained-surface-changed |
code on your live call chains changed | review the retained-surface diff |
new-dependency |
the PR adds a package: its tier, capabilities, and known CVEs get flagged | review what it pulls in |
expanded-usage-vulnerable |
your code opens a new call-path into a dependency that carries live CVEs | block until patched or justified |
capability-escalation |
the new version gained a dangerous primitive (network, spawn, secret read) even though tests pass | block; a dormant payload still passes a suite |
validation-regression |
new version fails your suite against your usage | do not adopt; file or wait |
manual-review-required |
structure or dynamism changed beyond automatic comparison | a human decides, with the diff in hand |
coverage --prompt · witness · chainbotImprove your app
The ceiling on removal is your test coverage; closing each gap has a named tool: test, probe, or robot user.
details
ChainStrip executes your test suite and collects data on which dependency
code was executed. The executed code survives the trimming process, the
unexecuted code is either dead or not tested. If the static analysis shows
that this code is reachable then it's not tested and chainstrip coverage --callers will show you how you can improve the tests to better validate
this dependency. If the dependency also carries a CVE it will be floated to
the top for priority reasons.
$ chainstrip coverage --callers ⚠ live CVE + under-tested → write this test first 38% next-auth ⚠ 3×CRITICAL → Providers, AdminUser 5% kysely ⚠ HIGH → booking-query compilation 34% nodemailer ⚠ 6 CVEs → sendVerificationRequest.ts 31% js-yaml ⚠ 2×HIGH → parseFrontmatter() 53% dompurify ⚠ 13×MODERATE → markdownToSafeHTMLClient() 0% @urql/core no CVE → SalesforceGraphQLClient $ chainstrip coverage --prompt wrote a test-writing prompt → .chainstrip/coverage-prompt.md hand it to an LLM/agent to draft the gap-closing tests
Evidence beyond the unit suite
Another source of execution evidence comes from running the production build
under the end-to-end test suite which includes both server and browser
components. i18next.dir(), called from the app layout, is an example of
such evidence: neither the unit
suite nor the build process exposed the call to it, but running the e2e suite
moved it into the "must stay" category. The client execution recording uses
the source maps to correlate itself to the source and has provisions to avoid
overtrimming due to stale source maps.
From finding to work order
Here is how ChainStrip not only protects, but also serves you to write better
software: while regular dependency security stops here, with the AppSec team
filing a ticket and then waiting months for the right sprint, ChainStrip goes
further. chainstrip coverage --prompt writes a Markdown document describing
the tests that need to be written. You can integrate it into a fully
autonomous agentic loop, generating tests on full autopilot and produce a
ready-to-merge PR. You then can again fully automatically build this PR,
rerun ChainStrip on it, prove that the coverage improved and, if not, hand it
to humans to review and triage. Some things are not testable from
ChainStrip's perspective and they will be included in the document as well,
and humans can review and override. Once the workflow ends the result is a
set of new durable tests that join your test suite. On one of our case
studies a single session off this prompt resulted in
15 dependencies tested and 2 real bugs found in the dependency usage.
ChainStrip doesn't fix the bugs and doesn't generate a prompt to fix the
bugs.
Or ChainStrip probes it directly
Once you figured out what to do with the missing tests, you need to deal
with the rest of untested dependency code that does not deserve the test,
and coverage --prompt mentioned it as not worth writing durable tests
for. chainstrip witness covers this tail and it's mutually exclusive
with coverage --prompt, which means that after you generated the tests
you'll need to rerun the coverage analysis and then witness to generate
a shorter worklist with only the long tail remaining. witness generates
"witness probes", small scripts make your app execute the unproven
code. It runs them against the hardened build, grades them and keeps the
generated proofs (an LLM connection is required for ChainStrip to draft
the probes itself, or by default it can create a prompt for an air-gapped
LLM). Probes are kept in the ChainStrip working directory and not merged
into your repository. The probe is not a test. Probe only proves that code
could be run given a plausible input, it can keep the code from removal
but can never justify removal. A failed probe proves nothing. One case
study moved 41 dependencies from "kept whole"
into "minimum surface trimmed".
A robot user
The final source of information is a robot user, chainbot, an optional module included with ChainStrip. It is used as a DIY e2e testing framework to simulate a human interacting with the deployed product. Chainbot discovers routes, parses HTML, locates controls and then interacts with the app (handling errors, filling forms, clicking buttons) using an LLM only to fill "human like" values. It does not use an LLM for decisions about results. The set of discovered but not yet exercised routes and controls is referred to as a "frontier". The crawl is finished when we run out of frontier.
When functions are removed, stubs are left in place so if a crawl
encounters code that has been removed, it will record a failure, store it
in the output report, and the next ChainStrip execution will restore the
removed code. This one run touched 27 routes on Mattermost (run of 2026-08-12), interacted
with 93 controls, executed 4 live flows and encountered 0 marker events;
which serves as evidence that ChainStrip reduced the attack surface and
didn't cause any harm.
This execution cost 34k input tokens and 4.3k
output tokens (using Claude Sonnet 5). The test-writing worklist is formatted in
markdown similar to coverage --prompt, so if you choose to hook chainbot
into an agentic workflow there won't be any major modifications needed.
Certain UI features cannot be definitively parsed by the crawler and the
flow may stall. That type of flow will be marked as blocked along with an
explanation why.
If your project currently lacks an e2e test suite, you can set up chainbot as one through ChainStrip.
Better evidence → more code provably removed → smaller surface → a shorter worklist next run.
How it works
Five stages rebuild your tree from evidence, one validation-gated decision at a time.
details
Removal is possible because ChainStrip rebuilds your dependency tree from evidence of what your app runs, one validation-gated decision at a time.
-
Inventory dependencies and app usage
Scan project source for import sites; union with what other installed packages import from each shared dependency. The result is a usage fingerprint per package.
-
Extract dependency source into a controlled workspace
Your checkout is never modified. Validation runs in disposable copy-on-write clones; extraction emits vendored packages with licenses, notices, and per-file provenance preserved. Extraction is dual-engine: rollup by default (measured 48.9% fewer emitted bytes than esbuild on a six-dependency comparison set, 2026-06-10), with a per-dependency esbuild fallback for the cases where rollup would drop a named export.
-
Remove the provably dead, hold the rest
Each dependency lands in exactly one tier. Tier A — statically clean: bundle additively from your real import sites; tree shaking removes dead branches for free. Tier B — contained dynamism: subtractive, validation-gated file and function pruning. Tier C — frameworks, native modules, heavy dynamism: retained whole, lifecycle-stripped, hash-pinned, and update-gated. One disposition rules every tier: a function is removed only with lexical proof of death, and reachable-but-unexecuted code is held, because a missing test is not proof. Three precision passes keep the held set small without touching that bar: build-executed dependencies are held per function the toolchain actually runs rather than whole, a class method counts as reachable only where a caller names it (dynamic dispatch degrades to held), and identifiers resolve by exact lexical scope. On Cal.diy the precision work tripled proven-dead removals, from 5 dependencies to 16, with zero change to the guarantee.
-
Validate against the target project's test suite
Differential: baseline vs overlay, your full suite plus a production build run from scratch — no build cache, no cached replay, so a dependency that only builds because a stale cache remembers it can't slip through. Beyond pass/fail, an export-shape oracle proves every subpath still exports what your code reaches, the class of silent bundling breakage a green suite misses. A hardened dependency is accepted only at zero new failures; one that passes tests but breaks the build auto-demotes to a pruned-but-complete copy. A failed extraction demotes, it never blocks the run.
-
Publish reports and reusable hardened artifacts
Validated packages are packed into per-dependency tarballs keyed by content fingerprint. The store is a plain directory in ChainStrip's state (
.chainstripunder your repo by default; relocatable by config), so it travels through your normal CI cache; on Cal.diy it weighs about 124 MB. That state stays out of version control — the only ChainStrip files your repo tracks are the config and any committed waiver approvals. Builds consume artifacts; reports record every decision and the evidence behind it.
Evidence
The numbers below are from measured runs on our biggest testbed, Cal.diy (the Cal.com OSS release, a production monorepo), each recorded in a dated run report. The CVE verdicts, the proven-dead counts, and the surface-reduction bars all come from one full run of 2026-07-31; the bars cover the 152 dependencies whose extractions passed validation in it. The deploy side has its own measured strip: a burn-in of 2026-07-17 applied the artifact, ran the production build, then removed development-only build twins and non-runtime files (type declarations, source maps, docs; licenses always kept) — the served `node_modules` went from 3.3 GB to 2.0 GB with build, boot, and login e2e green, and zero strip-induced regressions. Your numbers depend on how much of each package you actually use.
$ chainstrip reachability
advisories 189 in the installed tree (OSV 2026-07-31)
not shipped 103 never entered the artifact
eliminated 1 glob HIGH: bin.mjs pruned, gates passed
reached 48 16 HIGH · 26 MODERATE · 6 LOW
unresolved 37 too coarse to map to a symbol
proven dead removals in 18 advisory-carrying deps;
glob's reached the advisory itself
validation full suite: differential pass
production build: from-scratch pass
$ chainstrip exceptions [no-test-evidence] cron-parser · 312 KB at stake No test imports this package, so test evidence is unavailable for it. do now add one test that imports it — the next sweep promotes it to full test evidence meanwhile production-build gate + update gate still cover it [build-gate] glob · overlay rejected Production build failed: require condition resolved as ESM (error persisted verbatim below). do now nothing — auto-demoted, stock copy pinned and update-gated meanwhile regression pinned; retried next sweep
Workflow fit
ChainStrip rides your existing CI and PRs; production builds consume the artifact store.
details
You do not run a full scan on every commit. The from-zero run behind the hero (start-ui-web, 67 dependencies) took 24 minutes on one machine; the one behind Cal.diy's numbers (232 runtime dependencies) took 70. After that, pruning runs when its inputs change; everything else is a fingerprint comparison.
$ chainstrip score --dir .
Cold, test-independent risk cards for any tree in minutes; there is nothing to configure. The first thing to run on a new repo.
$ chainstrip drift $ chainstrip pack --check
~1 second: compares usage and lockfile fingerprints against the stored artifact, then either passes or names what drifted.
$ chainstrip sweep --tier AB $ chainstrip reachability
Only changed dependencies re-enter the pipeline; unchanged fingerprints reuse stored tarballs. `reachability` re-classifies every CVE against the new artifact.
$ chainstrip apply --dest . $ chainstrip bundle --production
Release builds consume validated artifacts only — overlay into a checkout, or a hermetic node_modules tarball with embedded verifier. A malicious upstream release cannot reach production without passing the gate.
Adoption modes
Start with a cold score and stop wherever your risk tolerance is met.
details
Start with a cold score and stop wherever your risk tolerance is met; each mode is useful on its own.
Score & audit
Cold package risk cards plus the full inventory, tier classification, and attack-surface map. Nothing about your build changes; you just learn what you ship versus what you actually use.
$ chainstrip score --dir . $ chainstrip report
CI validation
Harden against your suite, give every CVE its verdict, and add the update gate to your PRs — one no-arg run blocks a risky bump, a new dependency, or a new call into a vulnerable one, on your policy. The coverage worklist drives the flywheel. Production still builds from stock node_modules.
$ chainstrip reachability --gaps $ chainstrip update
Enforcement
Release builds consume validated artifacts only. Lifecycle scripts stripped everywhere; unvendored deps hash-pinned. Native modules ship as the prebuilt binaries validation ran — built once on a machine matching the deploy platform, hash-verified, nothing compiles at install time. Upstream compromise has to pass your gate to ship.
$ chainstrip apply --dest . $ chainstrip bundle --production
How it ships
ChainStrip is self-hosted: a runnable bundle for x86 and arm64 that lives where your code lives; if Node runs there, ChainStrip runs there. Your source, lockfiles, and reports stay in your infrastructure. (The lockfile ask in the demo box above is for the demo itself; the product never sends anything out.) Onboarding is white glove: we integrate ChainStrip into your GitHub or GitLab CI/CD pipeline ourselves, free of charge. chainbot, the robot user, ships in the same bundle as an optional component and isn't priced separately. Licensing is commercial, with source available to customers. It isn't open source today for a plain organizational reason: a healthy OSS project takes maintenance resources we are putting into the tool instead. Pricing isn't public yet; ask in the demo.
Release announcements
Low volume: releases and measured run reports. We won't put you in a marketing sequence or resell your address.