Start with separate sources
This example lets an analyst join a relational dataset with an event table through a shared query engine. Each source keeps its own availability, permissions, and update schedule. The service supports bounded analytical reads; it does not promise one transaction spanning both systems.
Every request has a server-verified user and tenant. Dataset owners define permitted columns, tenant boundaries, and acceptable load. A common SQL endpoint simplifies access, but those responsibilities still exist behind it.
Authorize before planning the read
The service checks catalog, table, column, and row access before returning protected data. Connector credentials have separate scopes for each source. If a connector uses a shared service account, the source may see only that account; the query service must enforce the caller's restrictions rather than assume the source knows the tenant.
Tenant filtering is mandatory policy, not a condition the analyst must remember. Test joins, views, metadata access, and cached results under different identities. Restrict direct access to connector credentials and alternate query endpoints that could bypass these checks.
Check where the work will run
The planner estimates row counts and decides which filters, projections, or aggregates a connector can send to its source. Pushdown depends on the connector and expression. A short SQL statement can still transfer a large table or create a large intermediate join.
Inspect the plan for representative queries and compare estimates with actual scans. This design requires bounded time ranges, limits returned rows and bytes, and puts interactive and scheduled queries in separate resource groups. Queue limits control admission; execution also needs memory, runtime, and scan limits. Unknown cost is a reason to narrow a request, not assume it is cheap.
Name the consistency limit
One source may be read at a stable snapshot while another changes during execution. Independent snapshots do not create a shared point in time. A join can therefore combine an updated record with an older event count without either source being faulty.
Record query start and finish times, source versions where available, and the requested data interval. Label results whose sources cannot supply stable versions. If a report requires a reproducible cutoff, select compatible published datasets or materialize a reconciled version before running it.
Cancel work and keep the evidence
If a required source fails, return an incomplete-query error rather than present the remaining rows as a complete answer. A retry may observe newer data, so keep it as a separate execution. Propagate cancellation to workers and connectors, and verify that expensive remote work actually stops; a closed browser connection proves little.
Track queue time, source latency, scanned and returned bytes, spills, rejected queries, and cancellation delay. Associate them with a query ID and access decision. Protect query logs because predicates can contain sensitive values. The source owner and query-service owner need an agreed response when a query overloads a dependency.
Know when to copy the data
Federation is useful when selective queries can reach data where it already lives. It adds dependency on source health and makes performance less predictable as queries change.
Repeated large joins, unstable source APIs, or strict reporting cutoffs can justify a materialized dataset instead. That alternative adds storage and refresh work, but moves expensive reads away from operational sources and gives consumers a published version to reference.