Provided metrics¶
LXD provides a number of instance metrics and internal metrics. See How to monitor metrics for instructions on how to work with these metrics.
Instance metrics¶
The following instance metrics are provided:
Metric |
Description |
|---|---|
|
Total number of effective CPUs |
|
Total number of CPU time used (in seconds) |
|
Total number of bytes read |
|
Total number of completed reads |
|
Total number of bytes written |
|
Total number of completed writes |
|
Available space (in bytes) |
|
Free space (in bytes) |
|
Size of the file system (in bytes) |
|
Amount of anonymous memory on active LRU list |
|
Amount of memory on active LRU list |
|
Amount of file-backed memory on active LRU list |
|
Amount of cached memory |
|
Amount of memory waiting to be written back to the disk |
|
Amount of free memory for |
|
Amount of used memory for |
|
Amount of anonymous memory on inactive LRU list |
|
Amount of memory on inactive LRU list |
|
Amount of file-backed memory on inactive LRU list |
|
Amount of mapped memory |
|
Amount of available memory |
|
Amount of free memory |
|
Amount of used memory |
|
The number of out-of-memory kills |
|
Amount of anonymous and swap cache memory |
|
Amount of cached file system data that is swap-backed |
|
Amount of reclaimable slab memory |
|
Amount of used swap memory |
|
Amount of unevictable memory |
|
Amount of memory queued for syncing to disk |
|
Amount of received bytes on a given interface |
|
Amount of received dropped bytes on a given interface |
|
Amount of received errors on a given interface |
|
Amount of received packets on a given interface |
|
Amount of transmitted bytes on a given interface |
|
Amount of transmitted dropped bytes on a given interface |
|
Amount of transmitted errors on a given interface |
|
Amount of transmitted packets on a given interface |
|
Number of running processes |
Internal metrics¶
The following internal metrics are provided:
Metric |
Description |
|---|---|
|
Total number of completed requests. See API rates metrics. |
|
Number of requests currently being handled. See API rates metrics. |
|
Total number of bytes allocated (even if freed) |
|
Number of bytes allocated and still in use |
|
Number of bytes used by the profiling bucket hash table |
|
Total number of frees |
|
Number of bytes used for garbage collection system metadata |
|
Number of goroutines that currently exist |
|
Number of heap bytes allocated and still in use |
|
Number of heap bytes waiting to be used |
|
Number of heap bytes that are in use |
|
Number of allocated objects |
|
Number of heap bytes released to OS |
|
Number of heap bytes obtained from system |
|
Total number of pointer lookups |
|
Total number of |
|
Number of bytes in use by |
|
Number of bytes used for |
|
Number of bytes in use by |
|
Number of bytes used for |
|
Number of heap bytes when next garbage collection will take place |
|
Number of bytes used for other system allocations |
|
Number of bytes in use by the stack allocator |
|
Number of bytes obtained from system for stack allocator |
|
Number of bytes obtained from system |
|
Number of running operations |
|
Number of configured replicators in the project. See Replicator metrics. |
|
Whether the last replicator run ended in the given status. See Replicator metrics. |
|
Time of the last successful replicator run (in seconds since the epoch). See Replicator metrics. |
|
Creation time of the oldest snapshot replicated by the last successful run (in seconds since the epoch). See Replicator metrics. |
|
Daemon uptime (in seconds) |
|
Number of active warnings |
API rates metrics¶
The API rates metrics include lxd_api_requests_completed_total and lxd_api_requests_ongoing. These metrics can be consumed by an observability tool deployed externally (for example, the Canonical Observability Stack or another third-party tool) to help identify failures or overload on a LXD server. You can set thresholds on the observability tools for these metrics’ values to trigger alarms and take programmatic actions.
These metrics consider all endpoints in the LXD REST API, with the exception of the / endpoint. Requests using an invalid URL are also counted. Requests against the metrics server are also counted. Both introduced metrics include a label entity_type based on the main entity type that the endpoint is operating on.
lxd_api_requests_ongoing contains the number of requests that are not yet completed by the time the metrics are queried. A request is considered completed when the response is returned to the client and any asynchronous operations spawned by that request are done. lxd_api_requests_completed_total contains the number of completed requests. This metric includes an additional label named result based on the outcome of the request. The label can have one of the following values:
error_server, for errors on the server side, this includes responses with HTTP status codes from 500 to 599. Any failed asynchronous operations also fall into this category.error_client, for responses with HTTP status codes from 400 to 499, indicating an error on the client side.succeeded, for endpoints that executed successfully.
Replicator metrics¶
The replicator metrics report the health of replicators, so that replication failures and the current recovery point objective (RPO) are visible in an observability tool rather than only through the API.
Replicator state is global to the cluster, so these metrics are reported only by the cluster member that is currently the database leader. Scraping every member therefore yields one sample per replicator rather than one per member.
lxd_replicators carries a project label and reports the number of replicators configured in that project. A sample is emitted for every project, using 0 for projects that have no replicators, so that unprotected projects can be identified.
lxd_replicator_last_run_status carries project, name and status labels. A sample is emitted for each of the four possible statuses (Pending, Running, Completed and Failed), with a value of 1 for the status of the last run and 0 for the others. Pending means the replicator has never run. Because every status is always present, a status change never leaves a stale series behind.
lxd_replicator_last_success_timestamp and lxd_replicator_last_success_oldest_snapshot_timestamp carry project and name labels and report times in seconds since the epoch. The first is the completion time of the last successful run, and the second is the creation time of the oldest snapshot that run replicated. A value of 0 means no such time has been recorded yet.
Some example queries:
# Number of replicators whose last run failed
sum(lxd_replicator_last_run_status{status="Failed"})
# Breakdown by status
sum by (status) (lxd_replicator_last_run_status)
# Projects with no replication configured
lxd_replicators == 0
# Time since the last successful run, in seconds
time() - lxd_replicator_last_success_timestamp
# Current RPO, in seconds
time() - lxd_replicator_last_success_oldest_snapshot_timestamp
# Alert: last successful run is older than an hourly schedule allows
time() - lxd_replicator_last_success_timestamp > 3600
# Replicators that have never succeeded
lxd_replicator_last_success_timestamp == 0