Safeguard
Buyer's Guides

Evaluate a Security Scanner in Three Weeks, Not Six Months

A standard bake-off measures detection on code both tools have seen, weights forty rows equally, and tests week one of a week-fifty problem. Six questions with real variance, and how to run them against your incumbent too.

Aman Khan
AppSec Engineer
6 min read

You have an incumbent scanner, a renewal date, and a suspicion that a different tool would be better. The default response is a bake-off: two vendors, a proof of concept, six months, a scoring matrix with forty rows.

That process reliably produces a decision and unreliably produces a good one, because most of what it measures is not what determines the outcome after you switch. This is a shorter evaluation that tests the things that actually differ. Written for whoever owns the renewal, and it is deliberately not a pitch for any particular tool, including ours.

Why the standard bake-off misleads

It measures detection on code both tools have seen. Run two scanners against a public benchmark or your own main repository and you get two similar numbers, because the underlying advisory data is largely shared. Detection is the most visible axis and the least differentiated one.

The scoring matrix weights by row count. Forty rows means forty equal-ish weights, so twenty low-stakes features outvote the one thing that will decide whether the tool survives contact with your engineers.

The vendor tunes the proof of concept. Naturally. A solutions engineer configures for the repositories you chose, which are the ones you already understand. Nobody configures for the awkward service written in 2019 that nobody owns.

It measures the wrong time window. A proof of concept measures week one. Everything that matters about a scanner is a week-fifty property: whether people still read the output, whether the backlog grows, whether the gate is still enabled.

Six questions that actually differentiate

Ask these of any candidate, including the incumbent, and weight them above everything else.

1. What is the false positive rate on our code? Not on a benchmark. Take 100 findings at random from a real repository, have an engineer who owns that code adjudicate each one, and compute the number. Do this for both tools on the same 100 areas.

This is the single most predictive measurement available, because precision determines whether engineers keep reading the output. A tool at 70 percent precision gets used. One at 30 percent gets ignored by March, whatever it detects.

2. What does it do about the finding? A ticket, a fix suggestion, or a tested pull request are three different products. The gap between "here is a vulnerability" and "here is a change that resolves it and passes your tests" is where most of the labour lives, and it is the axis with real variance between vendors right now.

Test it honestly: take ten real findings, let the tool propose fixes, and count how many merge without human modification. That number is usually much lower than the demo implies, for everyone. What matters is the comparison, not the absolute.

3. How much of the backlog does it dismiss, and can it explain why? Every vendor claims large noise reduction. The question is whether you can audit a dismissal. Ask to see the reasoning for one specific finding it filtered, and whether that reasoning survives you asking a follow-up.

A filter you cannot audit is a filter you cannot defend to an auditor or a customer, and eventually somebody asks.

4. What happens to our awkward repository? Pick the worst one: the polyglot service, the monorepo, the thing with a vendored dependency tree and a bespoke build. Do not let the vendor choose the test corpus.

Coverage failures cluster in exactly these places, and they are invisible in a proof of concept scoped to your two best repositories.

5. What is the total latency in the pull request path? Measure wall clock from pull request opened to check complete, on a real change, at your repository size. Over roughly ten minutes, developers context switch, and a slow check gets bypassed regardless of quality.

6. What does leaving cost? Ask it about your incumbent first, because you are living the answer. Can you export findings, suppressions, and policy? If triage decisions across three years are locked in a vendor's database, the migration cost includes re-triaging everything, and that cost belongs in the comparison rather than being discovered afterwards.

The evaluation that fits in three weeks

  • Week one. Both tools on the same three repositories: your largest, your most awkward, and one representative service. Configuration by you, not by the vendor.
  • Week two. The 100-finding adjudication for precision. The ten-finding fix test. Time the pull request path. This is the week that produces your actual decision.
  • Week three. Run the gate in warn mode on real traffic, and ask four engineers what they think. They will be blunter than the scoring matrix and more predictive.

Three weeks, and you will know more than a six-month bake-off tells you, because you measured the properties that vary rather than the ones that are easy to tabulate.

What to do about the incumbent

Run it through the same evaluation. Two things happen, and both are useful.

Sometimes it wins, and you have saved a migration and learned which settings were wrong. A surprising share of scanner dissatisfaction is a configuration problem: nobody tuned it after the initial rollout, the gate was set at a threshold that made people stop reading, and the tool is blamed for a policy decision.

Sometimes it loses clearly, and now you have a defensible reason for the switch rather than a feeling, which is what you need when someone asks why the budget moved.

The concession

Every evaluation is partly political. There is a champion, a sunk cost, a relationship with a rep, and a renewal date that is not aligned with when you would ideally decide. Pretending the decision is purely technical produces a technically correct recommendation that loses to the org chart.

So do the measurement, and also be honest about the constraint. If the real question is whether to consolidate several tools into one platform, the precision comparison above matters less than integration and contract structure, and you should evaluate for that instead of running a detection comparison whose result will not change the outcome.

The implication

The properties that determine whether a scanner is still being used in a year are precision on your code, what it does after it finds something, and how long it makes people wait. Those three are measurable in about two weeks.

Almost everything else on a standard evaluation matrix is either the same across vendors or will not matter by March. Cut the matrix to the questions with variance, and the decision gets both faster and better.

Never miss an update

Weekly insights on software supply chain security, delivered to your inbox.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.