Skip to content

Prometheus

Prometheus collects and stores time-series metrics from all CobaltCore components. It is the central metrics store that feeds alerting rules and Perses dashboards.

Exporters

ExporterSourceWhat it covers
ceph-exporterCeph daemonsOSD stats, pool usage, cluster health, latency histograms
rook-ceph-mgrRook managerOperator status, daemon lifecycle events
radosgw-exporterRGWRequest rates, error rates, per-user and per-bucket bandwidth
kvm-ha-agentHypervisor nodesHypervisor uptime, VM instance counts, libvirt events
OpenStack exportersNova, Neutron, CinderAPI latency, queue depths, service health

Retention and storage

Metrics are retained according to the cluster-wide retention policy. Long-term storage uses Prometheus remote-write to an external TSDB (configured separately per deployment).

Alert rules

Key alerting rules across the CobaltCore stack:

RuleCondition
CephHealthWarning / CephHealthErrorCluster health degradation
CephOSDNearFullOSD usage exceeds 85%
CephMonQuorumLostMonitor quorum lost
RGWHighErrorRateElevated 5xx rate on the RGW
HypervisorUnreachableHA agent stops reporting for > threshold

INFO

Full alert rule definitions and scrape configuration are being documented. See the CobaltCore observability charts for current configuration.

See also

EU and German government funding logos

Funded by the European Union – NextGenerationEU.

The views and opinions expressed are solely those of the author(s) and do not necessarily reflect the views of the European Union or the European Commission. Neither the European Union nor the European Commission can be held responsible for them.