Platform operations

Separate workload boundaries, release controls, and recovery responsibilities.

Illustrative reference architecture 4 min read

Reference diagram

Keep release decisions connected to workload behavior

Change control

  1. Versioned releaseOwner, image, configuration, state changes
    Review boundaries
  2. Policy checksRBAC, workload identity, quotas
    Set rollout limits
  3. Rollout decisionReadiness, capacity, stop conditions

Operation and recovery

  1. Workload runtimesSeparate permissions and resource budgets
    Measure service behavior
  2. Signals and ownersErrors, backlog, dependency health
    Exercise recovery
  3. Restore verificationIsolated restore and application checks

Connections between paths

  • Rollout decision→Workload runtimes

    Deploy and drain within the agreed limits

  • Signals and owners→Rollout decision

    Continue, pause, or invoke recovery

Control-plane permissions and runtime isolation solve different problems. Service owners use workload signals to stop unsafe changes; scheduled restore exercises test the recovery agreement.

Make the operating agreement explicit

This reference platform runs an ingestion service, a query service, and scheduled jobs on Kubernetes. Their owners share infrastructure but have different latency, capacity, and recovery needs. Some persistent state lives in managed databases or attached storage.

Before choosing components, name who owns cluster upgrades, application releases, credentials, backups, and source dependencies. Each workload needs a recovery target and a contact who can decide whether to pause it. A shared dashboard does not settle those decisions during an incident.

Separate control access from runtime access

Kubernetes RBAC controls who can act on API resources. Namespaces organize those resources; they do not alone stop network traffic or isolate the underlying kernel. Runtime boundaries need network policy, restricted pod privileges, service identities, and appropriate node or cluster separation.

This design starts with separate namespaces and service accounts, denied cross-workload traffic unless required, and a network plugin that enforces the policy. Untrusted code may require stronger isolation. Platform administrators retain broad power, so control-plane access and audit records need their own review.

Give each workload the access it needs

A job that reads one source should not inherit the query service's credentials or the deployment controller's permissions. Bind external tokens to the intended service and operations. Prefer short-lived credentials where supported, and keep them out of images, release manifests, and ordinary logs.

Secrets require restricted access, encrypted storage, and a rotation procedure that the application can tolerate. Permission to create a pod can also enable access to secrets available in its namespace. Test that boundary, including deployment automation, rather than reviewing only direct secret-read permissions.

Give developer workspaces their own boundaries

In this reference design, browser workspaces have per-user compute and storage quotas, explicit access to source repositories and data services, and controls on forwarded ports. Scope credentials to the user and workload; keep provisioning credentials outside the workspace. Isolate developer runtimes according to the code they can execute.

Separate persistent working files from replaceable runtime resources. Define what survives stopping or rebuilding a workspace, including uncommitted work. Set idle shutdown, state retention, and cleanup rules. Offboarding revokes access and credentials, then handles retained files under that policy. Test restart and revocation rather than assuming that a closed browser ends access.

Limit the damage from a release

Resource requests and limits describe each workload; namespace quotas cap its admitted allocation. Application queues and concurrency limits are still needed to protect databases and external APIs. An autoscaler cannot create unlimited source capacity.

For a release, first check available capacity and compatibility with existing state. Roll out a small portion, inspect service behavior, then continue. Readiness gates traffic, while graceful shutdown stops new work and drains or checkpoints in-flight work before termination. Define a maximum drain period and how unfinished work resumes.

A failed rollout should stop automatically at an agreed condition. External effects need their own reconciliation procedure.

Track database changes independently of images

Track versioned migrations by stable identity and checksum. All migration workers use the same history and target lock to serialize changes. Duplicate jobs check recorded progress; changed content under an applied identity requires review. Expand the schema first, backfill and verify data, move consumers, then retire the old representation when dependencies allow.

After interruption, inspect the schema and migration history before resuming or repairing; transactional DDL support varies by database. Confirm that no worker remains active before clearing a stale lock. An image rollback does not reverse schema changes, so retain a compatible application version or an explicit forward-repair or restore procedure.

Verify recovery with a real restore

Watch user-visible errors and latency alongside queue age, saturation, restarts, unavailable replicas, and dependency failures. Route alerts to the owner able to act. Keep release and configuration identifiers with the timeline so responders can connect a behavior change to an actual deployment.

Periodically restore protected data into an isolated environment and verify application reads, permissions, counts, and the recoverable time range. A successful backup job does not establish that recovery works. Record elapsed recovery time and missing data against the workload's target, then fix the gap or revise the promise.

Choose sharing deliberately

Sharing infrastructure can reduce duplicated work, but it creates shared failure and upgrade dependencies. Isolation policies also need maintenance as services change.

A smaller installation may be easier to operate with a managed service or separate runtimes. Choose a common platform when its owners can maintain the boundaries and recovery procedures, not simply because every workload can be packaged as a container.

References

Search the site

Search experience, studies, articles, projects, and contributions.

Try a topic