Enable alert rules¶
This guide walks through enabling Prometheus and Loki alert rules for your JAAS deployment.
The JIMM application ships a set of alert rules that are provisioned automatically to Prometheus through the same relation used for monitoring. The rules travel over the metrics-endpoint relation, so no integration is required beyond the ones described in Enable monitoring.
Note
This guide covers alert rules for JIMM only. For alert rules for the other components of your deployment — for example OpenFGA, PostgreSQL or Vault — see the documentation of the corresponding charms.
The current alert rules cover:
Alert |
Fires when |
Severity |
|---|---|---|
|
JIMM’s |
Critical |
|
Authentication failures exceed ~12 per minute for 5 minutes |
High |
|
JIMM experiences sustained errors calling the Juju API |
Warning |
|
The p95 database query latency is above 1 second for 10 minutes |
Warning |
|
JIMM experiences sustained errors calling Vault |
Warning |
|
JIMM experiences sustained errors calling OpenFGA |
Warning |
For the up-to-date list of rules and their exact expressions, see the alert rules in the JIMM charm source.
Prerequisites¶
A running JAAS deployment
A deployed COS stack, integrated with JAAS as described in Enable monitoring
Integrate with Prometheus¶
The alert rules are transferred over the monitoring relations:
JIMM endpoint |
Interface |
Alert rules |
|---|---|---|
|
|
Prometheus alert rules based on |
The metrics-endpoint relation is the same one used for monitoring — no additional integration is required beyond the ones described in Enable monitoring.
To be notified when an alert fires, also integrate the Prometheus application of your COS stack with Alertmanager.
Verify¶
To check which alert rules are loaded in Prometheus, open the Prometheus web UI of your COS deployment and check Status > Rules. Rules sourced from each related charm appear as groups named <model>_<model-uuid>_<application>_<rule-group>; the JIMM groups are named after the alerts listed above.
Firing alerts are visible in the Alertmanager UI of your COS deployment, and in the Alertmanager Operator Overview dashboard in Grafana.
Note
Alert rules based on counters (for example, authentication failures or Juju API errors) only evaluate once the corresponding time series exist. On an idle deployment most JIMM alerts would remain in an inactive state until JIMM starts handling traffic.