[MDR 10.2] Build Sheet: The Evaluation Scorecard and Proof-of-Value Plan

Most MDR evaluations are decided by a demonstration and a reference call, which between them test the vendor sales engineering and their choice of referee. A scorecard you can defend to a board, and a provider-neutral proof-of-value plan that measures what actually comes back.


Most managed detection evaluations are decided by a demonstration and a reference call, which between them test the vendor’s sales engineering and their choice of referee. This build sheet replaces that with a scorecard you can defend to a board and a proof-of-value plan that makes something happen in a trial tenant and measures what comes back. It is provider-neutral by design: the same plan runs against all of them, which is the only way the comparison means anything.

Do the readiness work first. Running a proof of value against a tenant with sparse telemetry measures your instrumentation rather than their analysts.


Step 1: Decide the architecture question before you invite anyone

Write down, before the first call, whether you are buying a second security stack, operators for the one you own, or a provider who leaves your sensors in place and replaces the data platform above them. That single answer eliminates most of the field and turns an unmanageable comparison into a shortlist of two or three. It also stops the evaluation drifting, which is the usual failure: three months in, you are comparing a platform against a service and wondering why the scores do not resolve.

Record the decision and the reasoning in one paragraph, get it agreed by whoever owns the security budget and whoever owns the Microsoft estate, and treat it as fixed unless something in the evaluation genuinely invalidates it.


Step 2: Build the scorecard

Seven axes, weighted to your environment rather than to a generic template. The weights below are a starting point for a Microsoft-centric organisation with hybrid identity; adjust them and record why you did.

AxisWhat you are scoringSuggested weight
Response authorityWhat they may do unilaterally, how it is scoped, whether you can withdraw it, where the audit trail lives25
Identity depthHybrid handling, whether they close the synchronisation gap, detect-and-respond versus inline prevention20
Tenant integrationWhich products they consume, permission model, what the credential can reach15
Telemetry residency and exitWhere data lands, retention, what you keep when the contract ends15
Commercial commitmentsWhat the published number measures, whether an obligation exists in the contract, incident response scope10
Overlap costWhat you pay for twice against your current licensing10
AI claimsGovernance of your users’ AI usage, and mechanism behind any operational AI claim5

Score each axis out of five against evidence rather than assertion, and require a source for every score above three. A claim in a demonstration is not evidence. Documentation, a contract clause, or something you watched happen in the trial is.

Authority carries the largest weight deliberately. It is the axis with the widest genuine spread across providers, it is the hardest to change after signature, and it is the one that determines what happens in your environment when nobody from your team is awake.


Step 3: Ask the seven questions, in writing

Send these to every provider before any demonstration and require written answers. The demonstration then becomes a chance to verify the answers rather than a chance to be impressed.

What actions can your analysts take in our tenant without contacting us first, how is that scope defined, and can we change it after signature without renegotiating. What does your published response figure measure, specifically which event starts the clock and which stops it, and is that figure a contractual obligation with a remedy. Where does our telemetry physically reside, how long do you retain it, and what do we receive on termination and in what format. For a hybrid account compromise, describe your containment sequence step by step, including how you handle directory synchronisation. What permissions does your integration require in our tenant, and can you provide the full list before we consent. Which of our existing Microsoft licences does your service duplicate, and what do you recommend we stop using. And on artificial intelligence, distinguish clearly between anything you offer for governing our users’ AI usage and any AI you use inside your own operations, describing the mechanism and where a human makes the decision.

How they answer is itself data. A provider who responds precisely, including where the honest answer is a limitation, is showing you how they will communicate during an incident. One who deflects every question to a call is showing you that too.


Step 4: Run the proof of value

Scope it to a defined set of machines and accounts, run it against every finalist over the same period, and generate the same events for each. Use a dedicated test account and a small group of pilot devices, with a change record and your own security team informed. Do not tell the provider which day you will run the scenarios.

Five scenarios, all benign, all detectable, all reflecting real intrusion behaviour rather than malware samples.

ScenarioWhat you generateWhat a competent provider should return
Impossible travelSign in as the test account from your normal location, then within the hour from a VPN endpoint in another continentAn alert naming both sign-ins, with device and location context, and a stated recommendation
Illicit consentConsent the test account to a benign multi-tenant application requesting mail read permissionsIdentification of the grant, the permissions requested, and whether the app is known
Mailbox persistenceCreate an inbox rule forwarding to an external address, and a second rule moving messages to a rarely-read folderBoth rules surfaced, with the forwarding one treated as higher severity
Endpoint discoveryOn a pilot device, run a sequence of built-in enumeration commands against the domainCorrelation of the commands into one behavioural detection rather than five unrelated alerts
Hybrid containmentAsk them to contain the test account, then observe what actually happens to it in both directoriesAccount disabled in Active Directory and Entra, sessions and refresh tokens revoked, and it stays disabled after a synchronisation cycle

The last scenario is the one that separates providers, and it is the one nobody runs. Ask for containment and then check the account thirty minutes later, after a sync cycle has passed. If it is enabled again, you have learned something no reference call would have told you.


Step 5: Measure the response, not the alert

For each scenario record four things, and record them from your own timestamps rather than from the provider’s reporting: the time you generated the event, the time you were first notified, what the notification actually contained, and what action was taken or recommended.

The interval that matters is the second minus the first, and it is the number you have been trying to get all along, measured in your environment rather than in their marketing. Comparing that figure across three providers running the same scenarios in the same week is worth more than every published metric in this market combined.

Judge the content as hard as the timing. A notification saying suspicious sign-in detected is worth very little. One naming the account, both locations, the device, whether multifactor was satisfied, what the account accessed afterwards and what they recommend you do is worth a great deal. You are buying analysis, so evaluate the analysis.

Note anything they missed entirely, and ask why. The answer is either a coverage gap on their side or a telemetry gap on yours, and both are useful findings. If it is yours, that is the readiness work returning value.


Step 6: Take the commercial terms apart before you score them

Model three years, not one. Include the licence cost of anything the provider requires you to add, the licence cost of anything you will continue paying for and stop using, the incident response arrangement and what happens on the second incident of a year, and the exit cost implied by wherever your telemetry lives.

Read the warranty document if a warranty forms part of the case, because the exclusions are not on the product page. Establish whether threat hunting, incident response and log retention are included or separate lines, since providers differ and the assumption is usually wrong in the direction that flatters the quote.


The decision record

Produce one page at the end: the architecture decision from step one and why, the scorecard with weights and evidence, the measured response times and content quality from the proof of value, the three-year cost including duplication, the exit terms as written into the agreement, and the specific limitations you accepted knowingly.

That last section is the one that pays off later. Every engagement has limitations, and an organisation that recorded them at signature is in a very different position at the post-incident review from one that discovers them during it. Write down what this provider will not see, what they will not do without asking, and what you decided to live with, then diary a review for the point at which any of those assumptions might have changed.

Re-run the two questions that go stale fastest at every renewal: whether the provider still has the same owner, and whether your own tenant now produces telemetry it did not when you signed. Both change more often than contracts do.


MDR
‹ Previous: [MDR 10.1] Build Sheet: MDR Readiness in Your Own Tenant