George Bogdanov

Published

Updated

Updated

Beyond Dependabot: Agentic Supply Chain Security

Read time: ...

GC AI's codebase is nearing half a million lines of TypeScript, with ~350 direct npm dependencies resolved to ~3,400 packages. Last month alone, our small engineering team shipped over 1,000 pull requests.

At such a volume and pace, Dependabot didn’t provide much help in keeping our supply chain secure. It solved detection, but detection was never a bottleneck. The hard part was cutting through the noise of 300+ advisories and understanding what was actually important and actionable. What does the advisory severity really mean for our codebase? Is the bug reachable? What could the update break?

What helped us bring advisories down to zero after two months, is an agentic workflow we've built that automated what Dependabot left to humans: prioritizing and landing clearly argued, well-scoped, and tested PRs.

From 300 advisories to 20 merged PRs

1. Triage

First, instead of relying on the original advisory severity, our tool runs a custom assessment specific to the usage of the affected package in our codebase, calculated via the Impact x Access matrix:

Impact: Severe (code exec, auth bypass, cross-tenant data, secrets) vs Bounded (everything less, including all availability-only impact).

Access: Ungated (anonymous traffic or untrusted content we auto-process: uploads, email, fetched pages) / Gated (authenticated user) / Implausible / None.

Additionally, it assesses the cost - whether it’s a simple patch bump or a major risky update requiring compatibility fixes.

As a result, the tool outputs a clear priority, for example:

P0: RCE-class flaw in decompress, a transitive dep of our document parser, which processes every customer upload (severe + ungated)
P1: prototype pollution in an object-merging dep, reachable only via authenticated API request bodies (severe, but the attacker needs a paid seat)
P2: ReDoS in a date-parsing library on user-supplied strings (bounded: one request stalls)
P3: critical-labeled RCE in a CLI tool that only runs in CI on our own inputs (access: none)

The first few runs flagged the vulnerabilities that were actually urgent, while downgrading a pile of “critical” advisories in code paths our application never reached. Eventually, we addressed those, too.

2. Bundling

Once the priorities are determined, the tool composes PRs and can bundle multiple smaller fixes together or isolate the risky ones separately. For each advisory, it picks the cheapest safe fix:

  • a lockfile refresh

  • a bump of the parent package that pulls in the vulnerable dep (often clearing several advisories at once)

  • a pinned override recorded as debt to retire later

  • or, as a last resort, a major upgrade

Low-risk fixes batch into a single hygiene PR. One such PR cleared 24 advisories across 11 packages (axios, nanoid, js-yaml, and others) by touching exactly two files: the pnpm workspace overrides and the lockfile - without breaking anything.

Moreover, since the agent has already analyzed the touch surfaces of the affected dependencies, it also generates a checklist of UI things to be regression-tested in the PR.

The output is well-scoped PRs that a human can review quickly. For the same supply chain, Dependabot requested 100+ PRs, whereas our tool has cleared the entire backlog in about 20.

3. Breakage

In addition to the security impact, the tool also assesses the UX impact (for lack of a better term). It describes which parts of our product’s functionality this update affects and which specific parts need to be tested, if necessary.

4. Shepherding

The tool doesn’t stop at just opening PRs. It watches the tests, tries to fix what’s broken with follow-up commits. Also, it addresses review comments from people or other agents.

As a result, by the time humans look at those PRs, they are much closer to being shippable. Of the 20+ PRs that have gone through this workflow, not one was reverted for breaking something in production.

Implementation

It started as a single prompt, but as the workflow grew, we split it into dedicated prompts for specialized subagents. The main agent coordinated each run: dispatching subagents, validating their plan, and recording their results in a persistent state. Subagents gathered advisories, investigated how affected packages were used, assigned priorities, and prepared fixes in isolated worktrees. A follow-up subagent watched CI and repaired failures, while a reporting subagent summarized the outcomes and decisions that needed human attention. A scheduled GitHub Action launched Claude to run the workflow.

Another improvement was making the workflow stateful. Each run loaded a TOML state file persisted as a GitHub Actions artifact, recording prior assessments, attempted fixes, open PRs, human decisions, and dates for revisiting deferred work. That lets the agent pick up where it left off, avoid repeating failed approaches, and reassess earlier decisions when the evidence has changed.

In the next iteration, we simplified orchestration by letting Cursor Cloud Agents implement fixes and maintain the PRs:

That’s a clearer division of responsibility: the updater decides what needs fixing and why; Cursor owns getting each fix ready for review.

Security considerations

The agent reads advisory descriptions, repository content, dependencies’ code, and CI logs. It also has permission to change code and open PRs. So, what if it’s prompt-injected?

We explored whether moving orchestration into code could limit what a compromised agent could do. We prototyped a deterministic harness: agents would still investigate vulnerabilities and prepare fixes, while code would decide which tasks to dispatch, validate the results, and check who was authorized to give instructions.

However, this approach didn’t yield much net gain: to be usable, it still required the cloud agents to have permissions to open pull requests. So, it’s not the orchestration piece that is the weak spot here. Implementation agents could still be prompt-injected to produce malicious code, and still have to be reviewed, falling into the same threat model as our regular development workflow.

For this reason, although our codebase is private, we treat feature branches and PRs as a quarantine zone, much like untrusted contributions to an open-source project.

Credentials available to PR jobs are isolated and limited in scope, and we separate their execution and caches from more privileged parts of our development pipeline to contain the impact of a potentially malicious change. Implementation agents can prepare changes and open PRs, but their permissions do not allow them to modify workflow files or dispatch deployment workflows.

We keep revisiting the permissions and boundaries around each stage and treat CI security as part of the development process rather than a one-time setup.

Future

Here’s where we see this tool evolving:

  1. Keeping dependencies up to date even when there is no security advisory, including updates that require code changes beyond a package.json bump.

  2. Reviewing dependency code, diffs, and changelogs to assess what an update introduces and look for suspicious changes, rather than relying only on published advisories.

So far, we are enjoying the flexibility to customize our own tool to our needs. At the same time, we're excited to try new tools as they appear.

Back To Top

Back To Top

George Bogdanov

Back To Top

SOC 2

Type II Certified

SOC 3

Certified

GDPR

Compliant

Book a personalized demo call

The AI platform built for in-house legal teams. SOC 2 certified. Providers do not train on your data, and zero-data-retention agreements apply wherever feasible. See it for yourself.

What to expect:

A walkthrough of the GC AI platform, tailored to your team's use cases.

Answers to your questions about security, integrations, and onboarding.

A 14-day trial if the platform looks like a fit for your team.