Skip to content

Monitoring

The operator exposes Prometheus metrics for observability. This guide covers the available metrics, how to enable scraping, and example alert rules.

Metrics Endpoint

The operator serves metrics on port 8080 at /metrics by default. This includes both controller-runtime default metrics and custom operator metrics.

Disabling Metrics

To disable the metrics endpoint entirely, set metrics.enabled: false in the Helm values:

metrics:
  enabled: false

This passes --metrics-bind-address=0 to the operator and removes the metrics port from the Deployment.

Securing the endpoint

By default the metrics endpoint serves plaintext HTTP to anything that can reach the pod. It carries no key material, but it does list every certificate the operator manages and when each expires - a ready-made inventory. Turn on authentication with:

metrics:
  secure: true

Each scrape then has to present a bearer token, which the operator verifies against the API server with a TokenReview and a SubjectAccessReview. The chart creates a {release}-metrics-reader ClusterRole granting get on /metrics; bind your scraper's ServiceAccount to it. Nothing is bound by default:

kubectl create clusterrolebinding prometheus-openvox-metrics \
  --clusterrole=<release>-openvox-operator-metrics-reader \
  --serviceaccount=monitoring:prometheus

The serving certificate

metrics.certManager.enabled Result
true (default) cert-manager issues the certificate from the chart's CA issuer, and the bundled ServiceMonitor verifies against it
false the operator generates a self-signed certificate at startup and a new one on every restart, so scrapers must skip verification
metrics.tls.certSecret set your own Secret with tls.crt and tls.key is mounted instead

The protection comes from the authentication filter, not from the certificate - which is why the self-signed path is usable and why insecureSkipVerify against an in-cluster endpoint that exposes no secrets is not the problem it looks like.

Turning this on breaks existing scrape configurations

The endpoint switches to HTTPS and starts requiring a token. Prometheus reports the target as down without saying why. The bundled ServiceMonitor is updated automatically; hand-written scrape configs are not.

Custom Metrics

Metric Type Labels Description
openvox_server_replicas_desired Gauge name, namespace Desired number of replicas for a Server CR
openvox_server_replicas_ready Gauge name, namespace Number of ready replicas for a Server CR
openvox_certificate_expiry_timestamp_seconds Gauge name, namespace Unix timestamp when a certificate or CA expires
openvox_crl_last_refresh_timestamp_seconds Gauge name, namespace Unix timestamp of the last successful CRL refresh for a CertificateAuthority

All series are removed when the resource they describe is deleted, so alerts do not keep firing for objects that no longer exist.

Controller-Runtime Metrics

These are automatically available:

Metric Description
controller_runtime_reconcile_total Total reconciliations per controller
controller_runtime_reconcile_errors_total Reconciliation errors per controller
controller_runtime_reconcile_time_seconds Reconciliation duration histogram
workqueue_depth Current work queue depth

Prometheus Integration

Metrics Service

A ClusterIP Service is created by default to expose the metrics endpoint:

metrics:
  enabled: true
  port: 8080
  service:
    enabled: true

ServiceMonitor

For clusters running the Prometheus Operator, enable the ServiceMonitor:

metrics:
  serviceMonitor:
    enabled: true
    interval: 30s
    labels: {}  # additional labels for ServiceMonitor selection

The ServiceMonitor selects the metrics Service by the standard app.kubernetes.io/name and app.kubernetes.io/instance labels.

Pod Annotations

For Prometheus setups that use annotation-based discovery instead of ServiceMonitor:

podAnnotations:
  prometheus.io/scrape: "true"
  prometheus.io/port: "8080"
  prometheus.io/path: "/metrics"

Example Alert Rules

Server Degraded

Alert when a Server CR has fewer ready replicas than desired for more than 5 minutes:

groups:
  - name: openvox
    rules:
      - alert: OpenVoxServerDegraded
        expr: openvox_server_replicas_ready < openvox_server_replicas_desired
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "OpenVox Server {{ $labels.name }} is degraded"
          description: "{{ $labels.name }} in {{ $labels.namespace }} has {{ $value }} ready replicas (desired: {{ with printf `openvox_server_replicas_desired{name=\"%s\",namespace=\"%s\"}` $labels.name $labels.namespace | query }}{{ . | first | value }}{{ end }})"

Certificate Expiring

Alert when a certificate expires within 30 days:

      - alert: OpenVoxCertExpiringSoon
        expr: openvox_certificate_expiry_timestamp_seconds - time() < 30 * 24 * 3600
        for: 1h
        labels:
          severity: warning
        annotations:
          summary: "OpenVox certificate {{ $labels.name }} expires soon"
          description: "{{ $labels.name }} in {{ $labels.namespace }} expires in {{ $value | humanizeDuration }}"

Missing CRL series

A stale CRL is alertable only while the series exists. If the operator never refreshed the CRL - it crashed early, or the CA never became ready - there is no series at all and the staleness rule below stays silent. Alert on the absence separately:

      - alert: OpenVoxCRLMetricMissing
        expr: absent(openvox_crl_last_refresh_timestamp_seconds)
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "No CRL refresh has been recorded"
          description: "The operator has not refreshed any CRL since it started. Revocations are not reaching agents."

Stale CRL

A CRL that is no longer refreshed means revoked agents keep being accepted. Alert when the last successful refresh is more than a day old:

      - alert: OpenVoxCRLStale
        expr: time() - openvox_crl_last_refresh_timestamp_seconds > 24 * 3600
        for: 1h
        labels:
          severity: critical
        annotations:
          summary: "OpenVox CRL for {{ $labels.name }} is stale"
          description: "The CRL of {{ $labels.name }} in {{ $labels.namespace }} was last refreshed {{ $value | humanizeDuration }} ago, so revoked agents may still be accepted"

CA Expiring

CAs and certificates share one metric

openvox_certificate_expiry_timestamp_seconds is written by both the Certificate and the CertificateAuthority controller, with the same name/namespace labels and nothing that says which kind a series belongs to. Telling them apart in a query means matching on the name.

The rule below uses .*-ca, which is a convention rather than a guarantee: it also catches a Certificate that happens to end in -ca, and it misses a CertificateAuthority named otherwise. Replace the matcher with your actual CA names if you rely on the distinction.

Alert when a CA certificate expires within 90 days:

      - alert: OpenVoxCAExpiringSoon
        expr: openvox_certificate_expiry_timestamp_seconds{name=~".*-ca"} - time() < 90 * 24 * 3600
        for: 1h
        labels:
          severity: critical
        annotations:
          summary: "OpenVox CA {{ $labels.name }} expires soon"
          description: "CA {{ $labels.name }} in {{ $labels.namespace }} expires in {{ $value | humanizeDuration }}"