Data catalogs

Keep extracted metadata, curated meaning, and source freshness separate.

Illustrative reference architecture 3 min read

Reference diagram

Technical observations and reviewed meaning meet in the catalog

Extraction and identity

  1. Scoped connectorsSource metadata, credentials, extraction status
    Record source identities and versions
  2. Versioned metadataSchema changes, aliases, retirement evidence
    Publish collected facts and their age
  3. Catalog recordSource evidence, metadata versions, coverage

Curation and consumption

  1. Owner reviewDefinitions, classifications, semantic changes
    Publish the reviewed definition
  2. Curated meaningAccountable owner and effective version
    Serve with metadata access checks
  3. Permitted discoverySearch and API results with freshness limits

Connections between paths

  • Versioned metadata→Owner review

    Present changes and unresolved identity matches

  • Catalog record→Permitted discovery

    Supply scoped technical facts alongside curated meaning

Extraction does not approve meaning, and a lineage graph does not establish complete dependency coverage. Metadata access and access to the underlying data remain separate decisions.

Separate discovery from curation

This illustrative catalog collects metadata from databases, pipelines, and reporting tools. Connectors discover technical facts; accountable owners supply definitions and review classifications. Assume independent source schedules, incomplete connector coverage, and authenticated consumers. The catalog helps people choose and understand data without replacing the source systems.

Start with a defined extraction scope and narrow service credentials. Collect schema and operational metadata by default. Sample rows and query text require separate justification and access controls because they can expose protected values.

Keep identity stable as names change

Give each asset a catalog identifier linked to its source system, environment, and native identifier where available. Record connector version, extraction time, and schema revision. A table name alone cannot distinguish a rename from a dropped table recreated under the same name.

When the source lacks stable identifiers, use explicit reconciliation rules and flag ambiguous matches for review. Keep aliases and retirement history rather than silently merging records. Version extracted schema separately from editorial changes. OpenMetadata's version history illustrates recording changes to schemas, descriptions, tags, and ownership.

Give meaning an accountable owner

Each useful entry needs a responsible team, a description of its grain, important units, and approved uses. Keep extracted descriptions, suggested classifications, and reviewed definitions distinguishable. An ingestion run should not overwrite an owner's correction without an explicit precedence rule.

A field changing from elapsed milliseconds to seconds may retain the same data type. Schema comparison cannot detect that meaning reliably. Record semantic changes, review affected consumers, and retain the earlier definition with its effective period. A classification suggestion becomes policy input only through the agreed review process.

Show how much lineage is known

Attach source evidence, collection method, and observation time to lineage edges. Distinguish a relationship extracted from executed jobs, one parsed from query text, and one entered by a person. They have different limits.

Public lineage ingestion documentation shows how coverage depends on connectors, available queries, and identifiable source entities. Missing lineage is unknown coverage, not proof that a dataset has no consumers. An impact review should list unresolved dependencies instead of treating the visible graph as complete.

Make failed discovery visible

Track successful extraction by source and scope, not just the scheduler's last run. Display stale metadata and partial scans. A failed connection or reduced permission scope must not be interpreted as mass deletion. Retire records only after an explicit deletion signal or a complete, comparable inventory confirms removal.

Metadata freshness and data freshness are separate: a recent schema scan says nothing about whether yesterday's records arrived. Monitor extraction gaps, unresolved owners, rejected updates, and stale lineage, with someone assigned to each failure.

Control access and verify meaning at use

Authorize search, asset details, lineage neighbors, exports, and API access before returning sensitive metadata. Test discovery routes as well as individual records; OpenMetadata documents separate search authorization settings. Permission to read a catalog entry does not grant access to its data.

Consumers still check the source contract and freshness required for their task. A catalog adds connector maintenance and stewardship work. For a small collection, an owned, versioned dataset register may be sufficient.

References

Search the site

Search experience, studies, articles, projects, and contributions.

Try a topic