Agent Foskett Academy • Microsoft Sentinel • Module 3 • Lesson 32

Lesson 32 — Testing a Detection Before Enabling It

The KQL runs without an error.

That does not mean the detection works.

Before an analytics rule starts feeding the SOC queue, prove that it finds the activity you expect, ignores the activity you do not want, produces useful entities and alert context, and behaves properly across the rule's real execution window.

A query that runs is not automatically a detection that works.
Agent Foskett Microsoft Sentinel testing detections before enabling analytics rules lesson
What you will learn

Validate the complete detection before it reaches an analyst.

Historical query testing
True and false positive validation
Rule simulation and result review
Entity and alert-context checks

Learning objectives

  • Separate query validation from analytics-rule validation.
  • Test detection logic against historical evidence.
  • Validate expected true positives and false positives.
  • Check thresholds, entities and alert enrichment.
  • Use rule simulation before enabling scheduled detections.

The investigation

A new rule has been written to detect repeated failed sign-ins followed by a successful sign-in.

The query runs. The engineer enables it. Ten minutes later the SOC queue starts filling with incidents nobody expected.

There are two things to test

TestQuestion
Detection logicDoes the KQL identify the behaviour we intended to detect?
Analytics rule behaviourDoes the configured rule turn those results into useful alerts and incidents?
Agent Foskett rule:

Passing the first test does not guarantee you pass the second.

The pre-enablement test workflow

Detection hypothesis │ ▼ Write KQL │ ▼ Run against history │ ▼ Inspect matching evidence │ ┌────────┴────────┐ ▼ ▼ Expected hits Unexpected hits │ │ └────────┬────────┘ ▼ Tune and retest │ ▼ Configure analytics rule │ ▼ Simulate rule results │ ▼ Validate alert configuration │ ▼ Enable

Start with the hypothesis

Write down exactly what the detection is supposed to identify before testing it.

For example: “Detect an account with repeated failed sign-ins followed by a successful sign-in from the same IP address within a short period.”

Define the expected evidence

What fields should prove the hypothesis? User, IP address, timestamps, result codes and application context may all matter.

If the query cannot return the evidence needed to investigate the behaviour, it is not ready.

Run the detection as a hunting query first

Before creating the analytics rule, execute the KQL interactively across a useful historical period.

Review failed sign-in patterns
  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
SigninLogs
| where TimeGenerated > ago(7d)
| where ResultType != "0"
| summarize FailedSignIns = count(),
            FirstSeen = min(TimeGenerated),
            LastSeen = max(TimeGenerated)
    by UserPrincipalName, IPAddress
| where FailedSignIns >= 10

Do not stop at the row count. Open the results and ask whether the matched activity actually represents the behaviour you intended to detect.

Find known true positives

If you have a previous confirmed incident or controlled test event, use it as a validation case.

The detection should find the evidence you already know it is supposed to find.

Find known false positives

Run the same logic across known benign activity such as approved scanners, service accounts or expected administrative behaviour.

If those events match, decide whether the rule needs tuning before production.

Test the boundaries

Good testing includes cases around the detection threshold — not only obvious malicious and obvious benign examples.

Test caseExpected result
9 failed sign-ins when threshold is 10Should not match.
10 failed sign-insShould match if all other conditions are satisfied.
50 failures from an approved scannerDepends on the documented exception logic.
10 failures spread across several usersShould not behave like 10 failures against one account unless that is the detection design.
Matching activity outside the lookup windowShould not unexpectedly combine with the current execution window.

Test missing and unusual values

Real telemetry is rarely perfect. Check how the query behaves when an IP address, account field, device name or other expected value is empty.

One null field should not silently break an entire detection path.

Test time properly

The analytics rule will execute over a defined lookup period and frequency. Test using the same time assumptions.

A query that works across seven days of manually selected data may behave very differently when it executes every five minutes over a ten-minute window.

Test the rule configuration — not only the KQL

After the query behaves correctly, validate the rest of the analytics rule.

Query frequency and lookup period
Alert threshold
Entity mapping
Custom details and dynamic alert details
Alert and incident grouping

Use Results simulation

For scheduled analytics rules, Microsoft Sentinel provides a Results simulation area while configuring rule logic.

It can show how the rule would have behaved against recent data so you can review the expected alert volume before enabling it.

Look at the volume

If your “high-confidence” detection would have generated hundreds of alerts during the simulation period, investigate why.

The simulator is giving you a preview of tomorrow's SOC queue.

Validate entity mapping

A detection can find the correct activity and still produce a poor investigation experience if its entities are missing or mapped incorrectly.

Query result │ ┌───────────┼───────────┐ ▼ ▼ ▼ Account IP Host │ │ │ └───────────┼───────────┘ ▼ Sentinel entities │ ▼ Investigation / correlation

Confirm that the fields returned by the query actually populate the entity identifiers configured in the rule.

Validate custom details

If the alert is supposed to surface a source application, event count, device or risk value, make sure those fields exist in the final query output.

Renaming or removing a projected field can leave enrichment mappings useless.

Read the alert like an analyst

Ask whether someone seeing the alert for the first time can understand what happened and where to investigate next.

A technically correct detection with no useful context simply moves the work downstream.

Keep a simple test matrix

ScenarioExpectedActual
Known malicious sampleDetectedRecord result
Known benign scannerExcluded or handled as designedRecord result
Below thresholdNo detectionRecord result
At thresholdDetectionRecord result
Missing entity valueGraceful behaviourRecord result
Normal business activityNo excessive alertingRecord result

Change one thing at a time

If you alter the threshold, lookup period, exclusions and grouping simultaneously, you may not know which change fixed — or broke — the detection.

Controlled tuning produces evidence you can explain later.

Retest after every meaningful change

A small KQL change can affect row counts, entity fields, thresholds and incident behaviour.

Testing is not a one-time ceremony before the first deployment.

Common mistake — testing only syntax

“No errors” means the query parsed and executed.

It does not mean the logic is correct, the threshold is sensible or the alert will help an analyst.

Common mistake — using only clean test data

Production telemetry contains exceptions, missing values, service accounts, scanners and strange edge cases.

Test against the messy evidence too.

Common mistake — enabling to see what happens

The SOC queue should not be your test environment.

Use historical execution, simulation and controlled validation before production alerting.

Common mistake — forgetting the analyst

A rule may detect exactly what the engineer intended but still create an alert with no useful entities or investigation context.

Test the final experience, not just the query.

Agent Foskett investigation exercise

Your new rule returns 42 historical matches.

Thirty are a known scanner, eight are normal administrator activity and four look genuinely suspicious.

  1. Do not enable the rule yet.
  2. Validate why the scanner and administrator activity are benign.
  3. Tune only the conditions supported by that evidence.
  4. Re-run the query and confirm the four suspicious cases still match.
  5. Use rule simulation to inspect expected alert volume.
  6. Check entities, custom details and incident behaviour before enabling.

Best practices

  • Write the detection hypothesis before testing.
  • Use historical data to understand real behaviour.
  • Test known true positives and known false positives.
  • Test boundaries, null values and timing assumptions.
  • Use Results simulation for scheduled rules.
  • Validate entities and alert context as part of the test.
  • Retest after meaningful detection changes.

Agent Foskett takeaway

Do not ask, “Does the query run?”

Ask, “Does the detection find the evidence I expect, reject the evidence I do not want, and give the analyst enough context to investigate?”

The logs already knew. Your test proves the rule understood them.

Lesson summary
Testing a Microsoft Sentinel detection means validating both the KQL logic and the configured analytics rule. Use historical evidence, true and false positive cases, boundary tests, rule simulation, entity mapping and alert context before enabling a detection in production.
Sentinel Academy Home

Continue learning

Continue through Module 3 — Analytics Rules.

Testing a Microsoft Sentinel Detection Before Enabling It

Microsoft Sentinel scheduled analytics rules should be tested against historical data, known true positives, known false positives, threshold boundaries and realistic timing before they are enabled for production alerting.

Microsoft Sentinel Lesson 32

This Agent Foskett Microsoft Sentinel Academy lesson explains detection validation, historical KQL testing, Results simulation, thresholds, entity mapping, alert enrichment and practical pre-production analytics rule testing.