Detection platforms

Separate rule authoring from controlled execution, and make missed coverage visible.

Illustrative reference architecture 3 min read

Reference diagram

Authoring and detection execution

Authoring and release

  1. Rule projectPython, tests, input contract, owner
    build an immutable candidate by digest
  2. Built candidateTest and shadow the same artifact digest
    approve and promote without rebuilding
  3. Release artifactThe approved, tested artifact digest

Evaluation and investigation

  1. Scoped inputsTenant, event window, source progress
    authorize data and enrichment access
  2. Bounded executionIsolated worker; brokered enrichment
    persist findings before recording progress
  3. Analyst handoffVersioned evidence, grouping, delivery status

Connections between paths

  • Release artifact→Bounded execution

    Verify the approved artifact digest before execution

Illustrative reference architecture. Release approval governs code changes; runtime authorization governs each evaluation. Failed or disabled rules remain visible as coverage gaps.

Code with an operating contract

This illustrative platform accepts Python detections over security events. Authors can express custom logic, but a reviewed rule still needs limits, reproducible inputs, and an accountable owner. Assume multiple tenants, delayed events, changing schemas, and occasional retries. This is a reference design, not a description of a workplace system.

Each detection declares its input schemas, required fields, evaluation window, permitted enrichment, resource budget, and response owner. Unsupported inputs produce a coverage error, not a clean result.

Review and release an artifact

Authors submit code, dependency locks, example events, expected matches, and negative cases together. Tests include duplicate deliveries, missing fields, tenant boundaries, and events on either side of a window boundary. Review checks the detection's hypothesis as well as its implementation.

Build an immutable candidate containing pinned code, dependencies, and configuration, identified by digest. Run artifact tests and shadow evaluation against that same candidate, comparing its behavior with the active version. Approval and deployment promote the tested digest without rebuilding it; any artifact change requires a new build and evaluation. Rollback restores the previous artifact and its compatible state. Historical evaluations record enrichment snapshots or their versions so changing external data does not masquerade as a code change.

Give execution narrow authority

A coordinator supplies authorized tenant inputs and a pinned artifact to an isolated worker. Bound wall time, memory, input rows, state growth, and enrichment calls. Use a read-only filesystem and deny network access except through approved services. A Python subprocess or ordinary container alone is not the isolation contract; Kubernetes documents the limits of shared-kernel isolation.

Rules receive scoped enrichment responses rather than long-lived credentials. A broker rechecks tenant, operation, and resource permissions on each request. Logs and exceptions follow the same tenant boundary and avoid copying sensitive event payloads unnecessarily.

Define which events count

Stateful detections evaluate explicit event-time intervals, such as a half-open window that includes its start and excludes its end. Record source progress and an allowed-lateness policy. A watermark estimates progress; it does not prove that every earlier event arrived, as the Beam model explains.

Remove repeated source events before counting them, using tenant-scoped source identifiers. Late arrivals can revise a provisional finding within the stated policy. Older arrivals enter a bounded replay with a separate run identity and an explicit rule for amending prior findings.

Separate findings from notifications

Event deduplication prevents repeated input from inflating a result. Finding grouping combines related matches. Notification suppression controls how often people are interrupted; it must retain the underlying evidence and suppression reason. Panther's rule documentation illustrates Python detections and alert grouping, but grouping is not proof that input was processed once.

Commit a finding identity and its output before advancing durable progress. Send analysts the rule version, window, relevant evidence references, uncertainty, and a response guide. A match opens an investigation; it does not automatically authorize containment.

Give analysts a workbench

The illustrative workbench shows rule version, deployed digest, environment, enabled status, and run history together. Analysts can launch ad hoc tests over a stated interval using the same bounded runtime, with response actions disabled. Queued, running, failed, partial, and completed runs remain distinguishable.

Tables, charts, and exports reference the same recorded run, filters, and coverage. Viewing results does not permit editing code, starting tests, promoting releases, or exporting evidence: authorize each operation separately and recheck access when downloading. A changed rule creates a new candidate; rerunning history does not quietly change the deployed version.

Make failures and costs visible

Track evaluation lag, source gaps, execution timeouts, exception rates, state size, and findings awaiting delivery. Repeatedly failing rules can be disabled to protect capacity, but the interface must show when coverage stopped, why, and who owns recovery. A quiet rule is not necessarily a healthy rule.

Python enables expressive detections at the cost of dependency maintenance, isolation, and testing. A constrained declarative rule or scheduled SQL query may be easier to review and operate when it expresses the same detection clearly.

References

Search the site

Search experience, studies, articles, projects, and contributions.

Try a topic