Kafka Platform Operations Cheat Sheet: Streams on OpenShift
Recall operator ownership, access boundaries, delivery guarantees and recovery evidence before practising Kafka platform scenarios.
Use this reference to frame an answer, then test your reasoning with the free practice set . The specialization uses Red Hat Streams 3.2 and OpenShift 4.20; support boundaries matter as much as upstream feature availability.
Identify the owner and the boundary
| Evidence or change | First distinction |
|---|---|
| Kafka or KafkaNodePool desired state | Cluster Operator reconciliation; node-pool resources belong to KafkaNodePool. |
| KafkaTopic configuration | Topic Operator ownership and authoritative Git state; partition counts cannot be reduced in place. |
| KafkaUser credentials and Kafka ACLs | User Operator reconciliation; generated credentials must reach the intended client securely. |
| Permission to change a custom resource | OpenShift RBAC authorizes the Kubernetes API; Kafka ACLs do not grant it. |
| A client can bootstrap but cannot consume | Check every advertised broker address, DNS, routing, TLS name and network policy. |
| Resource Ready or generation observed | Check the condition and workload outcome; neither a processed generation nor a green dashboard proves business acceptance. |
State the guarantee precisely
acks=all concerns the source cluster’s in-sync replicas. It does not wait for asynchronous replication to another site. min.insync.replicas and healthy replication affect availability and durability, but do not make an external database write atomic with a consumer offset commit.
A consumer can repeat an external effect after a crash between that effect and its offset commit. Use an appropriate idempotency or reconciliation design and test the crash boundary. Kafka Streams processing guarantees cover defined Kafka/state operations; name anything outside that boundary.
KRaft controllers own metadata consensus. Brokers store and serve records. Red Hat Streams 3.2 supports static controller quorums: do not treat upstream dynamic controller membership as a supported production scaling procedure for this baseline.
Read the capacity evidence
| Calculation | Assumption you must name |
|---|---|
| Peak rate × a capacity increment | State whether headroom is extra capacity above demand or a fraction of installed capacity kept unused. |
| Retained bytes × replication | Distinguish logical input, compression, replicas, reserve and recovery traffic. |
| Backlog ÷ net drain rate | Net drain is processing rate minus continuing arrival rate; it must be positive. |
| State size ÷ restore throughput | Include coordination time and competing restoration traffic; standby lag differs from a full rebuild. |
| Aggregate broker capacity | Check per-broker placement, skew, failure domains and the proposed rebalance. |
High key cardinality alone does not rule out a hot key. A partition’s ordered work cannot be divided among several active members of the same consumer group simply by adding consumers.
Keep integration assumptions explicit
Before changing a registry or serializer, record the required client library, API endpoint, wire framing and schema-ID lookup behavior. Include retained records in the migration test. A successful schema registration alone does not establish that old records remain readable.
For a topology test, choose input that forces the intended transformation. An already-uppercase status cannot demonstrate that case conversion works. Assert the output key and value, and distinguish the tested topology from the artifact actually deployed.
For an operator change, identify the desired-state resource, its reconciling operator and its generated outputs. Check the resulting workload behavior after reconciliation rather than treating the requested YAML as proof of success.
Verify recovery and promotion
- Define the authoritative source, expected result and approved recovery point.
- Preserve credentials and secrets through their controlled management mechanism.
- Validate client traffic with representative identities, endpoints and network-policy context.
- Reconcile business records after replay; consumer lag alone does not prove correct external effects.
- Test alerts through notification and runbook execution by the intended operator.
- Before failback, fence competing writers and establish the reconciled authoritative history.
Read the technical references for supported procedures. Continue in the practice bank and explain which observation would change your decision.