Define the behavior being tested
I would test a detection as a chain: source telemetry, parsing, rule evaluation, alert construction, and action routing. This illustrative design assumes versioned rules, an isolated test destination, and a controlled way to replay permitted records. Each scenario specifies the behavior, required fields, timing, and expected output.
A syntactically valid query is only one check. Elastic's public rule tests validate syntax, schemas, and metadata; this design adds behavioral fixtures and replay checks around those checks. A passing result applies to the tested input, configuration, and engine version, not every environment where the rule might run.
Build fixtures with honest labels
Synthetic fixtures exercise clear boundaries: a matching sequence, a similar benign sequence, missing fields, duplicate delivery, and events arriving outside the expected window. Preserve relationships and relative timing when sanitizing replay data. Record every substitution because removing an identifier or changing a distribution can change the behavior under test.
Store the scenario author, expected result, label rationale, and unresolved ambiguity beside the data. Analyst disposition is useful evidence, but an uninvestigated alert is not a confirmed false positive. Keep uncertain examples separate from a scored evaluation set and review labels when new evidence becomes available.
Run the path that production depends on
Pin the parser, field mappings, rule revision, engine settings, and fixture version. Where possible, replay source records through parsing before evaluation; injecting already normalized events skips a common failure point. Test the final generated query when a portable rule format is converted to a backend query.
Assert which events produced which alerts, including entity grouping and event windows. Record expected nonmatches too. Sigma's log-source declaration helps document required telemetry, but it does not establish that collection is enabled. A missing alert may indicate absent input, parser failure, suppression, or rule behavior. Report those outcomes separately.
Compare safely before switching actions
After fixture checks pass, run the candidate in shadow mode against the same permitted live input as the active version. Write candidate alerts to a comparison destination without response credentials. Compare event coverage, alert differences, processing cost, and reviewed samples across a stated observation interval.
Promotion changes which version may request actions. Fence stale workers with an activation generation checked at the action gateway. Use a shared action ledger keyed by logical alert and action, independently of rule revision. It rejects duplicate requests from replay or overlapping schedules; destination effects still require idempotency or reconciliation. Rollback restores an earlier version while preserving that ledger.
Watch telemetry and detection drift
Track source arrival gaps, required-field completeness, parse failures, evaluation lag, rule errors, suppressions, and alert-volume changes. Segment these by relevant source and rule version. A falling alert count could reflect quieter activity or broken collection; the count alone cannot distinguish them.
ATT&CK mappings organize the behaviors a rule addresses. They are not a percentage of attacks prevented or a substitute for exercised scenarios. Maintain separate records of mapped techniques, telemetry availability, tested procedures, and reviewed outcomes. For behavior-based detection, changes to the baseline population or learning window also require evaluation.
Choose evidence over a single score
Fixtures are repeatable but narrow. Sanitized replay is more representative, yet costly to maintain and limited by what was collected and labeled. Shadow comparison reveals operational differences without establishing the truth of every alert.
A smaller program can start with reviewed synthetic scenarios, versioned query checks, and a manual release checklist. Add replay and shadow infrastructure when rule volume or action risk justifies it. Regardless of scale, no alert is not proof of no intrusion, and a green test suite does not remove the need for investigation.