Separate discovery from curation
This illustrative catalog collects metadata from databases, pipelines, and reporting tools. Connectors discover technical facts; accountable owners supply definitions and review classifications. Assume independent source schedules, incomplete connector coverage, and authenticated consumers. The catalog helps people choose and understand data without replacing the source systems.
Start with a defined extraction scope and narrow service credentials. Collect schema and operational metadata by default. Sample rows and query text require separate justification and access controls because they can expose protected values.
Keep identity stable as names change
Give each asset a catalog identifier linked to its source system, environment, and native identifier where available. Record connector version, extraction time, and schema revision. A table name alone cannot distinguish a rename from a dropped table recreated under the same name.
When the source lacks stable identifiers, use explicit reconciliation rules and flag ambiguous matches for review. Keep aliases and retirement history rather than silently merging records. Version extracted schema separately from editorial changes. OpenMetadata's version history illustrates recording changes to schemas, descriptions, tags, and ownership.
Give meaning an accountable owner
Each useful entry needs a responsible team, a description of its grain, important units, and approved uses. Keep extracted descriptions, suggested classifications, and reviewed definitions distinguishable. An ingestion run should not overwrite an owner's correction without an explicit precedence rule.
A field changing from elapsed milliseconds to seconds may retain the same data type. Schema comparison cannot detect that meaning reliably. Record semantic changes, review affected consumers, and retain the earlier definition with its effective period. A classification suggestion becomes policy input only through the agreed review process.
Show how much lineage is known
Attach source evidence, collection method, and observation time to lineage edges. Distinguish a relationship extracted from executed jobs, one parsed from query text, and one entered by a person. They have different limits.
Public lineage ingestion documentation shows how coverage depends on connectors, available queries, and identifiable source entities. Missing lineage is unknown coverage, not proof that a dataset has no consumers. An impact review should list unresolved dependencies instead of treating the visible graph as complete.
Make failed discovery visible
Track successful extraction by source and scope, not just the scheduler's last run. Display stale metadata and partial scans. A failed connection or reduced permission scope must not be interpreted as mass deletion. Retire records only after an explicit deletion signal or a complete, comparable inventory confirms removal.
Metadata freshness and data freshness are separate: a recent schema scan says nothing about whether yesterday's records arrived. Monitor extraction gaps, unresolved owners, rejected updates, and stale lineage, with someone assigned to each failure.
Control access and verify meaning at use
Authorize search, asset details, lineage neighbors, exports, and API access before returning sensitive metadata. Test discovery routes as well as individual records; OpenMetadata documents separate search authorization settings. Permission to read a catalog entry does not grant access to its data.
Consumers still check the source contract and freshness required for their task. A catalog adds connector maintenance and stewardship work. For a small collection, an owned, versioned dataset register may be sufficient.