Lesson 32 — Testing a Detection Before Enabling It
The KQL runs without an error.
That does not mean the detection works.
Before an analytics rule starts feeding the SOC queue, prove that it finds the activity you expect, ignores the activity you do not want, produces useful entities and alert context, and behaves properly across the rule's real execution window.

What you will learn
Validate the complete detection before it reaches an analyst.
Learning objectives
- Separate query validation from analytics-rule validation.
- Test detection logic against historical evidence.
- Validate expected true positives and false positives.
- Check thresholds, entities and alert enrichment.
- Use rule simulation before enabling scheduled detections.
The investigation
A new rule has been written to detect repeated failed sign-ins followed by a successful sign-in.
The query runs. The engineer enables it. Ten minutes later the SOC queue starts filling with incidents nobody expected.
There are two things to test
| Test | Question |
|---|---|
| Detection logic | Does the KQL identify the behaviour we intended to detect? |
| Analytics rule behaviour | Does the configured rule turn those results into useful alerts and incidents? |
Passing the first test does not guarantee you pass the second.
The pre-enablement test workflow
Start with the hypothesis
Write down exactly what the detection is supposed to identify before testing it.
For example: “Detect an account with repeated failed sign-ins followed by a successful sign-in from the same IP address within a short period.”
Define the expected evidence
What fields should prove the hypothesis? User, IP address, timestamps, result codes and application context may all matter.
If the query cannot return the evidence needed to investigate the behaviour, it is not ready.
Run the detection as a hunting query first
Before creating the analytics rule, execute the KQL interactively across a useful historical period.
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
SigninLogs
| where TimeGenerated > ago(7d)
| where ResultType != "0"
| summarize FailedSignIns = count(),
FirstSeen = min(TimeGenerated),
LastSeen = max(TimeGenerated)
by UserPrincipalName, IPAddress
| where FailedSignIns >= 10
Do not stop at the row count. Open the results and ask whether the matched activity actually represents the behaviour you intended to detect.
Find known true positives
If you have a previous confirmed incident or controlled test event, use it as a validation case.
The detection should find the evidence you already know it is supposed to find.
Find known false positives
Run the same logic across known benign activity such as approved scanners, service accounts or expected administrative behaviour.
If those events match, decide whether the rule needs tuning before production.
Test the boundaries
Good testing includes cases around the detection threshold — not only obvious malicious and obvious benign examples.
| Test case | Expected result |
|---|---|
| 9 failed sign-ins when threshold is 10 | Should not match. |
| 10 failed sign-ins | Should match if all other conditions are satisfied. |
| 50 failures from an approved scanner | Depends on the documented exception logic. |
| 10 failures spread across several users | Should not behave like 10 failures against one account unless that is the detection design. |
| Matching activity outside the lookup window | Should not unexpectedly combine with the current execution window. |
Test missing and unusual values
Real telemetry is rarely perfect. Check how the query behaves when an IP address, account field, device name or other expected value is empty.
One null field should not silently break an entire detection path.
Test time properly
The analytics rule will execute over a defined lookup period and frequency. Test using the same time assumptions.
A query that works across seven days of manually selected data may behave very differently when it executes every five minutes over a ten-minute window.
Test the rule configuration — not only the KQL
After the query behaves correctly, validate the rest of the analytics rule.
Use Results simulation
For scheduled analytics rules, Microsoft Sentinel provides a Results simulation area while configuring rule logic.
It can show how the rule would have behaved against recent data so you can review the expected alert volume before enabling it.
Look at the volume
If your “high-confidence” detection would have generated hundreds of alerts during the simulation period, investigate why.
The simulator is giving you a preview of tomorrow's SOC queue.
Validate entity mapping
A detection can find the correct activity and still produce a poor investigation experience if its entities are missing or mapped incorrectly.
Confirm that the fields returned by the query actually populate the entity identifiers configured in the rule.
Validate custom details
If the alert is supposed to surface a source application, event count, device or risk value, make sure those fields exist in the final query output.
Renaming or removing a projected field can leave enrichment mappings useless.
Read the alert like an analyst
Ask whether someone seeing the alert for the first time can understand what happened and where to investigate next.
A technically correct detection with no useful context simply moves the work downstream.
Keep a simple test matrix
| Scenario | Expected | Actual |
|---|---|---|
| Known malicious sample | Detected | Record result |
| Known benign scanner | Excluded or handled as designed | Record result |
| Below threshold | No detection | Record result |
| At threshold | Detection | Record result |
| Missing entity value | Graceful behaviour | Record result |
| Normal business activity | No excessive alerting | Record result |
Change one thing at a time
If you alter the threshold, lookup period, exclusions and grouping simultaneously, you may not know which change fixed — or broke — the detection.
Controlled tuning produces evidence you can explain later.
Retest after every meaningful change
A small KQL change can affect row counts, entity fields, thresholds and incident behaviour.
Testing is not a one-time ceremony before the first deployment.
Common mistake — testing only syntax
“No errors” means the query parsed and executed.
It does not mean the logic is correct, the threshold is sensible or the alert will help an analyst.
Common mistake — using only clean test data
Production telemetry contains exceptions, missing values, service accounts, scanners and strange edge cases.
Test against the messy evidence too.
Common mistake — enabling to see what happens
The SOC queue should not be your test environment.
Use historical execution, simulation and controlled validation before production alerting.
Common mistake — forgetting the analyst
A rule may detect exactly what the engineer intended but still create an alert with no useful entities or investigation context.
Test the final experience, not just the query.
Agent Foskett investigation exercise
Thirty are a known scanner, eight are normal administrator activity and four look genuinely suspicious.
- Do not enable the rule yet.
- Validate why the scanner and administrator activity are benign.
- Tune only the conditions supported by that evidence.
- Re-run the query and confirm the four suspicious cases still match.
- Use rule simulation to inspect expected alert volume.
- Check entities, custom details and incident behaviour before enabling.
Best practices
- Write the detection hypothesis before testing.
- Use historical data to understand real behaviour.
- Test known true positives and known false positives.
- Test boundaries, null values and timing assumptions.
- Use Results simulation for scheduled rules.
- Validate entities and alert context as part of the test.
- Retest after meaningful detection changes.
Agent Foskett takeaway
Do not ask, “Does the query run?”
Ask, “Does the detection find the evidence I expect, reject the evidence I do not want, and give the analyst enough context to investigate?”
The logs already knew. Your test proves the rule understood them.
Related Agent Foskett learning
Continue learning
Testing a Microsoft Sentinel Detection Before Enabling It
Microsoft Sentinel scheduled analytics rules should be tested against historical data, known true positives, known false positives, threshold boundaries and realistic timing before they are enabled for production alerting.
Microsoft Sentinel Lesson 32
This Agent Foskett Microsoft Sentinel Academy lesson explains detection validation, historical KQL testing, Results simulation, thresholds, entity mapping, alert enrichment and practical pre-production analytics rule testing.
