Data freshness

Separate operational freshness from a reproducible analytical result.

Illustrative reference architecture 2 min read

Figure 02 / Two publication contracts

Shared event historyEvent ID · Event time · Ingestion time · Schema version

01 / Operational view

  1. Continuous processingHandle late and repeated events
    incremental updates
  2. Serving storeFast reads; provisional results
    show source progress and gaps
  3. Operational decisionsAct on recent activity

02 / Analytical view

  1. Scheduled transformationBound the interval and reconcile
    validate before publishing
  2. Versioned datasetRecord inputs and transformation version
    publish a complete version
  3. Reporting & analysisReproducible results; explicit corrections
Illustrative reference architecture. Both paths keep the same definitions where possible. Their freshness and completeness contracts differ.

Two different needs

An operational view may need recent activity even while some records are missing. A report may need a stable interval that can be reconciled and reproduced. This reference design gives those views different publication contracts while preserving a common event history.

Assume each event has an ID, event time, ingestion time, and a schema version. Records can arrive late or out of order. Corrections are possible. Raw history remains available for a defined recovery period. No particular throughput or latency target is assumed.

Give each view a contract

The operational path consumes events continuously and updates a serving store. It exposes processing success and progress for each source. One active source must not hide another source that has stopped.

The analytical path runs on a schedule, validates and reconciles its input, and publishes a named dataset version. Where the storage system supports it, publication exposes a complete version atomically. A reader can identify the input interval and transformation version behind the result.

A recent job timestamp does not establish that all expected source data has arrived.

Event time and late arrivals

Event time says when an event belongs in the business timeline. Processing time says when the system handled it. A watermark estimates progress through event time; it does not prove that every earlier event has arrived.

A lateness policy determines how long a result can change through the normal path and how later corrections are handled. The interface should distinguish provisional results from a published, reconciled version.

Schema changes and backfills

Adding an optional field is different from changing the meaning of an existing one. Preserve the original representation and state which schema versions each transformation accepts. A table format can maintain field identity through structural changes, but it cannot validate a business definition.

A backfill selects a bounded input range, records its code and schema versions, and writes to a staging destination. Before publication, compare record counts, duplicates, missing intervals, and relevant aggregates. Define a cutoff or merge rule for concurrent arrivals so the backfill does not replace newer data by accident.

When a second path is worth it

Two paths mean more operational work and a period when their views may disagree. They need shared definitions, clear ownership, and labels that explain those differences.

A streaming path earns its place when fresher data changes an action. If a scheduled update meets the need, a single scheduled pipeline can be easier to operate and explain.

References

Search the site

Search experience, studies, articles, projects, and contributions.

Try a topic