Skip to content

Observability

CobaltCore's observability stack provides real-time visibility into every layer - compute, storage, networking, and OpenStack services. It combines Prometheus for metrics, Perses for dashboards, and Prysm for storage-specific monitoring and audit.

Components

ComponentRole
PrysmObservability CLI for Ceph and RGW - real-time monitoring, SMART disk health, log audit
PrometheusScrapes and stores time-series metrics from all CobaltCore components
PersesDashboard platform for visualizing metrics from Prometheus

Key metrics coverage

DomainWhat is monitored
ComputeHypervisor uptime, VM counts, migration status, HA agent telemetry
StorageCeph health, OSD status, pool usage, RGW throughput, Arbiter monitor reachability
OpenStackNova API latency, Neutron agent status, Cinder volume operations
NetworkingOVN controller status, OVS flow counts

Alerting

Alerts are defined as Prometheus rules. Critical thresholds monitored across the stack:

AlertCondition
CephHealthWarning / CephHealthErrorCluster health degradation
CephOSDNearFullOSD usage exceeds 85%
CephMonQuorumLostMonitor quorum lost
RGWHighErrorRateElevated 5xx rate on the RGW
HypervisorUnreachableHA agent stops reporting

See also

EU and German government funding logos

Funded by the European Union – NextGenerationEU.

The views and opinions expressed are solely those of the author(s) and do not necessarily reflect the views of the European Union or the European Commission. Neither the European Union nor the European Commission can be held responsible for them.