Start with coverage and evidence
This illustrative security information and event management platform combines recent-event search with longer-lived analytical storage. It accepts security events from several source types and supports both scheduled detections and analyst queries. Sources can pause, retry, change format, or provide incomplete history.
Assume enforceable tenant access, stable source identities where available, and finite retention budgets. Faster queries are useful only when the reader can tell which sources, time range, and representations were actually searched.
Normalize without erasing lineage
Capture the original record with its source, tenant, event time, ingestion time, and receipt identifier. A versioned parser creates a normalized event and retains a reference to that original. Shared schemas such as OCSF provide common field definitions; they do not prove that a source timestamp or identity mapping is correct.
Distinguish missing, invalid, and genuinely empty fields. Quarantine incompatible records with a reason and expose their counts. Deduplication follows the source's identifier contract: identical payloads can represent separate legitimate events. Reprocessing writes a new transformation version so a parser correction does not silently rewrite evidence already used in a case.
Set separate storage contracts
Publish selected fields and recent intervals to the search index. Store normalized history in partitioned analytical files, with raw records retained for a separately defined period. Catalogue coverage by source, tenant, interval, and parser version. An index is a serving projection; an archive is useful only if a supported path can retrieve and interpret it.
Retention includes expiry and any required restoration delay. Security Lake's lifecycle controls illustrate storage transitions and deletion as explicit settings. This design also tracks derived copies and case evidence so removing a raw record does not leave an unexplained retention exception elsewhere.
Admit detection and investigation work
The detection control path versions rule definitions, owners, schedules, and intended source coverage. Each run receives an authorized scope and records its rule version, input interval, and completion status. Streaming evaluations and historical queries need explicit overlap and finding-reconciliation rules.
Analyst queries pass through a separate admission policy with tenant checks, time bounds, scan estimates, concurrency limits, and cancellation. Reserve capacity for scheduled detection so an expensive investigation does not starve it. Route work to the index or lake according to requested coverage and supported semantics, rather than returning the cheapest available subset without explanation.
Return completeness with results
One storage write may succeed while another fails. Track publication progress independently, retry safely, and reconcile against captured receipts. Acknowledgement at ingestion is not a general exactly-once guarantee across the index, lake, and alert destination.
Queries return missing partitions, failed shards, timeouts, and excluded sources alongside matches. A completed request can still be partial; Elasticsearch's async search response exposes this distinction. Detection treats incomplete required coverage as an execution problem, not as evidence that nothing suspicious happened.
Measure the cost of coverage
Track per-source arrival gaps, parser rejection, index lag, lake publication lag, retention headroom, bytes scanned, and detection deadlines missed. Reconcile captured and published record identities over comparable intervals, accounting for known filters and late arrivals; equal totals alone do not establish equal event sets.
Traces help diagnose processing delays and failures, but trace sampling may omit requests. Keep completeness checks independent of sampled telemetry and reconcile expected records against destination receipts. A successful trace does not prove that every expected event was stored.
Two stores add indexing cost, duplicated transformations, access-policy work, and reconciliation. A single search platform may suffice for a smaller retention window; a lake-first design may suit slower investigations. The choice follows the required response time and evidence contract, not the label of a next-generation SIEM.