Kafka Platform Operations Cheat Sheet: Streams on OpenShift

Recall operator ownership, access boundaries, delivery guarantees and recovery evidence before practising Kafka platform scenarios.

Use this reference to frame an answer, then test your reasoning with the free practice set . The specialization uses Red Hat Streams 3.2 and OpenShift 4.20; support boundaries matter as much as upstream feature availability.

Identify the owner and the boundary

Evidence or changeFirst distinction
Kafka or KafkaNodePool desired stateCluster Operator reconciliation; node-pool resources belong to KafkaNodePool.
KafkaTopic configurationTopic Operator ownership and authoritative Git state; partition counts cannot be reduced in place.
KafkaUser credentials and Kafka ACLsUser Operator reconciliation; generated credentials must reach the intended client securely.
Permission to change a custom resourceOpenShift RBAC authorizes the Kubernetes API; Kafka ACLs do not grant it.
A client can bootstrap but cannot consumeCheck every advertised broker address, DNS, routing, TLS name and network policy.
Resource Ready or generation observedCheck the condition and workload outcome; neither a processed generation nor a green dashboard proves business acceptance.

State the guarantee precisely

acks=all concerns the source cluster’s in-sync replicas. It does not wait for asynchronous replication to another site. min.insync.replicas and healthy replication affect availability and durability, but do not make an external database write atomic with a consumer offset commit.

A consumer can repeat an external effect after a crash between that effect and its offset commit. Use an appropriate idempotency or reconciliation design and test the crash boundary. Kafka Streams processing guarantees cover defined Kafka/state operations; name anything outside that boundary.

KRaft controllers own metadata consensus. Brokers store and serve records. Red Hat Streams 3.2 supports static controller quorums: do not treat upstream dynamic controller membership as a supported production scaling procedure for this baseline.

Read the capacity evidence

CalculationAssumption you must name
Peak rate × a capacity incrementState whether headroom is extra capacity above demand or a fraction of installed capacity kept unused.
Retained bytes × replicationDistinguish logical input, compression, replicas, reserve and recovery traffic.
Backlog ÷ net drain rateNet drain is processing rate minus continuing arrival rate; it must be positive.
State size ÷ restore throughputInclude coordination time and competing restoration traffic; standby lag differs from a full rebuild.
Aggregate broker capacityCheck per-broker placement, skew, failure domains and the proposed rebalance.

High key cardinality alone does not rule out a hot key. A partition’s ordered work cannot be divided among several active members of the same consumer group simply by adding consumers.

Keep integration assumptions explicit

Before changing a registry or serializer, record the required client library, API endpoint, wire framing and schema-ID lookup behavior. Include retained records in the migration test. A successful schema registration alone does not establish that old records remain readable.

For a topology test, choose input that forces the intended transformation. An already-uppercase status cannot demonstrate that case conversion works. Assert the output key and value, and distinguish the tested topology from the artifact actually deployed.

For an operator change, identify the desired-state resource, its reconciling operator and its generated outputs. Check the resulting workload behavior after reconciliation rather than treating the requested YAML as proof of success.

Verify recovery and promotion

  1. Define the authoritative source, expected result and approved recovery point.
  2. Preserve credentials and secrets through their controlled management mechanism.
  3. Validate client traffic with representative identities, endpoints and network-policy context.
  4. Reconcile business records after replay; consumer lag alone does not prove correct external effects.
  5. Test alerts through notification and runbook execution by the intended operator.
  6. Before failback, fence competing writers and establish the reconciled authoritative history.

Read the technical references for supported procedures. Continue in the practice bank and explain which observation would change your decision.