Skip to main content
RisingWave Cloud provides detailed alert playbooks to help you quickly diagnose and address issues. Each playbook entry includes the alert’s name, description, common triggers, diagnostic steps, and immediate remediation actions. These alerts are organized by their function categories. This guide is regularly updated to address emerging scenarios.

Service Health

Out of memory error

A RisingWave container was terminated due to an Out of Memory (OOM) condition in the cluster. See the dedicated Out of memory error playbook. Triggers
  • Large backfill operations.
  • Sudden spikes in workload.
Resolution Scale the node resources. For more information, see Configure node resources.

Streaming

Barrier pending for too long

No barrier has been committed in this cluster for longer than the configured threshold. See the dedicated Barrier pending for too long playbook. Triggers
  • Streaming graph bottlenecks. Typical causes include join amplification, insufficient resources, and suboptimal streaming queries such as over-window queries and joins.
  • Compaction write stalls that result in longer barrier synchronization times.
Diagnosis
  • Check CPU and memory utilization for all nodes. If the nodes are fully utilized, the cluster might have insufficient resources.
  • Run SHOW JOBS and check for jobs that are being created. Backfilling a newly created job can increase pressure on the cluster.
Resolution If either resources are maxed out, or backfilling is happening, scale out the cluster to alleviate the pressure.

Sink lag too large

Data for a particular sink has been pending in RisingWave’s internal log store for longer than the configured threshold. See the dedicated Sink lag too large playbook. Triggers
  • Slow external sink processing.
  • Insufficient sink parallelism.
Diagnosis Check the downstream of the sink to see if there’s any abnormality.

PostgreSQL CDC WAL lag

A PostgreSQL CDC source is more than 20 GiB behind the upstream WAL for at least 15 minutes. See the dedicated PostgreSQL CDC WAL lag playbook.

Compaction

Compaction back pressure

Back pressure from compaction detected in your cluster. See the dedicated Compaction back pressure playbook. Triggers
  • Insufficient compaction resources.
Diagnosis
  • Check compaction CPU usage.
  • Check the CPU ratio of Compute Nodes (including Streaming and Serving Nodes) and Compactor Nodes.
Resolution Scale the compactor out. For more information, see Scale a cluster manually.