Monitoring¶
The operator exposes Prometheus metrics for observability. This guide covers the available metrics, how to enable scraping, and example alert rules.
Metrics Endpoint¶
The operator serves metrics on port 8080 at /metrics by default. This includes both controller-runtime default metrics and custom operator metrics.
Disabling Metrics¶
To disable the metrics endpoint entirely, set metrics.enabled: false in the Helm values:
This passes --metrics-bind-address=0 to the operator and removes the metrics port from the Deployment.
Securing the endpoint¶
By default the metrics endpoint serves plaintext HTTP to anything that can reach the pod. It carries no key material, but it does list every certificate the operator manages and when each expires - a ready-made inventory. Turn on authentication with:
Each scrape then has to present a bearer token, which the operator verifies
against the API server with a TokenReview and a SubjectAccessReview. The chart
creates a {release}-metrics-reader ClusterRole granting get on /metrics;
bind your scraper's ServiceAccount to it. Nothing is bound by default:
kubectl create clusterrolebinding prometheus-openvox-metrics \
--clusterrole=<release>-openvox-operator-metrics-reader \
--serviceaccount=monitoring:prometheus
The serving certificate¶
metrics.certManager.enabled |
Result |
|---|---|
true (default) |
cert-manager issues the certificate from the chart's CA issuer, and the bundled ServiceMonitor verifies against it |
false |
the operator generates a self-signed certificate at startup and a new one on every restart, so scrapers must skip verification |
metrics.tls.certSecret set |
your own Secret with tls.crt and tls.key is mounted instead |
The protection comes from the authentication filter, not from the
certificate - which is why the self-signed path is usable and why
insecureSkipVerify against an in-cluster endpoint that exposes no secrets is
not the problem it looks like.
Turning this on breaks existing scrape configurations
The endpoint switches to HTTPS and starts requiring a token. Prometheus reports the target as down without saying why. The bundled ServiceMonitor is updated automatically; hand-written scrape configs are not.
Custom Metrics¶
| Metric | Type | Labels | Description |
|---|---|---|---|
openvox_server_replicas_desired |
Gauge | name, namespace |
Desired number of replicas for a Server CR |
openvox_server_replicas_ready |
Gauge | name, namespace |
Number of ready replicas for a Server CR |
openvox_certificate_expiry_timestamp_seconds |
Gauge | name, namespace |
Unix timestamp when a certificate or CA expires |
openvox_crl_last_refresh_timestamp_seconds |
Gauge | name, namespace |
Unix timestamp of the last successful CRL refresh for a CertificateAuthority |
All series are removed when the resource they describe is deleted, so alerts do not keep firing for objects that no longer exist.
Controller-Runtime Metrics¶
These are automatically available:
| Metric | Description |
|---|---|
controller_runtime_reconcile_total |
Total reconciliations per controller |
controller_runtime_reconcile_errors_total |
Reconciliation errors per controller |
controller_runtime_reconcile_time_seconds |
Reconciliation duration histogram |
workqueue_depth |
Current work queue depth |
Prometheus Integration¶
Metrics Service¶
A ClusterIP Service is created by default to expose the metrics endpoint:
ServiceMonitor¶
For clusters running the Prometheus Operator, enable the ServiceMonitor:
metrics:
serviceMonitor:
enabled: true
interval: 30s
labels: {} # additional labels for ServiceMonitor selection
The ServiceMonitor selects the metrics Service by the standard app.kubernetes.io/name and app.kubernetes.io/instance labels.
Pod Annotations¶
For Prometheus setups that use annotation-based discovery instead of ServiceMonitor:
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8080"
prometheus.io/path: "/metrics"
Example Alert Rules¶
Server Degraded¶
Alert when a Server CR has fewer ready replicas than desired for more than 5 minutes:
groups:
- name: openvox
rules:
- alert: OpenVoxServerDegraded
expr: openvox_server_replicas_ready < openvox_server_replicas_desired
for: 5m
labels:
severity: warning
annotations:
summary: "OpenVox Server {{ $labels.name }} is degraded"
description: "{{ $labels.name }} in {{ $labels.namespace }} has {{ $value }} ready replicas (desired: {{ with printf `openvox_server_replicas_desired{name=\"%s\",namespace=\"%s\"}` $labels.name $labels.namespace | query }}{{ . | first | value }}{{ end }})"
Certificate Expiring¶
Alert when a certificate expires within 30 days:
- alert: OpenVoxCertExpiringSoon
expr: openvox_certificate_expiry_timestamp_seconds - time() < 30 * 24 * 3600
for: 1h
labels:
severity: warning
annotations:
summary: "OpenVox certificate {{ $labels.name }} expires soon"
description: "{{ $labels.name }} in {{ $labels.namespace }} expires in {{ $value | humanizeDuration }}"
Missing CRL series¶
A stale CRL is alertable only while the series exists. If the operator never refreshed the CRL - it crashed early, or the CA never became ready - there is no series at all and the staleness rule below stays silent. Alert on the absence separately:
- alert: OpenVoxCRLMetricMissing
expr: absent(openvox_crl_last_refresh_timestamp_seconds)
for: 15m
labels:
severity: warning
annotations:
summary: "No CRL refresh has been recorded"
description: "The operator has not refreshed any CRL since it started. Revocations are not reaching agents."
Stale CRL¶
A CRL that is no longer refreshed means revoked agents keep being accepted. Alert when the last successful refresh is more than a day old:
- alert: OpenVoxCRLStale
expr: time() - openvox_crl_last_refresh_timestamp_seconds > 24 * 3600
for: 1h
labels:
severity: critical
annotations:
summary: "OpenVox CRL for {{ $labels.name }} is stale"
description: "The CRL of {{ $labels.name }} in {{ $labels.namespace }} was last refreshed {{ $value | humanizeDuration }} ago, so revoked agents may still be accepted"
CA Expiring¶
CAs and certificates share one metric
openvox_certificate_expiry_timestamp_seconds is written by both the
Certificate and the CertificateAuthority controller, with the same
name/namespace labels and nothing that says which kind a series belongs
to. Telling them apart in a query means matching on the name.
The rule below uses .*-ca, which is a convention rather than a guarantee:
it also catches a Certificate that happens to end in -ca, and it misses a
CertificateAuthority named otherwise. Replace the matcher with your actual
CA names if you rely on the distinction.
Alert when a CA certificate expires within 90 days:
- alert: OpenVoxCAExpiringSoon
expr: openvox_certificate_expiry_timestamp_seconds{name=~".*-ca"} - time() < 90 * 24 * 3600
for: 1h
labels:
severity: critical
annotations:
summary: "OpenVox CA {{ $labels.name }} expires soon"
description: "CA {{ $labels.name }} in {{ $labels.namespace }} expires in {{ $value | humanizeDuration }}"