Monitoring Unify with Prometheus metrics
Unify exposes its metrics in Prometheus format on the metrics endpoint (default http://<host>:9090/metrics). Every metric is a per-broker instance except where noted elsewhere here. In a cluster deployment, each node exposes its own values and dashboards normally sum them across nodes. mqtt_retained_messages is the one exception as described below.
Event-driven counters update in real time. Gauges that require scanning connected clients or shared states (mqtt_client_inflight_utilization, mqtt_clients_send_quota_saturated, mqtt_wildcard_subscriptions, mqtt_topics_tracked, mqtt_retained_messages, and the broker-level connection/subscription gauges) refresh on an internal sampler every 10 seconds by default. Set METRICS_SAMPLE_INTERVAL to change the interval. The sampler refreshes connector metrics every 15 seconds.
The following sections example queries in PromQL, the query language used by Prometheus and Grafana. Note the following conventions:
- Counters (*_total) only ever increase. Query them with rate(...[5m]) to get a per-second rate or increase(...[1h]) to get a total over a window. The raw counter value is rarely useful on its own.
- Gauges are point-in-time values. Query them directly and wrap them in sum(...) to aggregate across cluster nodes.
Message throughput
These are the primary load indicators. The rate of mqtt_messages_received_total is the inbound publish load and mqtt_messages_sent_total is delivery load (received rate × average fan-out). The byte counters distinguish "many small messages" from "few large messages", which is relevant for capacity planning. A growing gap between received rate and sent rate can indicate messages with no subscribers, ACL denials, or drops (see Message drops and rejections).
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_messages_received_total | Total PUBLISH messages received from clients, broken down by QoS level. | Counter | qos (0/1/2) |
mqtt_messages_sent_total | Total PUBLISH messages delivered to subscribers, by QoS. One inbound message fanned out to N subscribers counts N times here. | Counter | qos |
mqtt_message_bytes_received_total | Total payload bytes of PUBLISH messages received from clients, by QoS. | Counter | qos |
mqtt_message_bytes_sent_total | Total wire bytes of PUBLISH packets sent to subscribers, by QoS. | Counter | qos |
mqtt_published_total | Legacy counter of all published messages, any QoS. This counter remains for backward compatibility; prefer mqtt_messages_received_total. | Counter | — |
mqtt_packets_sent_total | Legacy counter of all MQTT packets sent (all packet types, not just PUBLISH). | Counter | — |
mqtt_packets_read_total | Legacy counter of all MQTT packets read (all packet types). | Counter | — |
Example queries:
# Inbound publish rate (messages/sec) across the cluster
sum(rate(mqtt_messages_received_total[5m]))
# Delivery rate broken down by QoS level
sum by (qos) (rate(mqtt_messages_sent_total[5m]))
# Average fan-out: how many subscribers each published message reaches
sum(rate(mqtt_messages_sent_total[5m])) / sum(rate(mqtt_messages_received_total[5m]))
# Outbound bandwidth (bytes/sec) — capacity planning
sum(rate(mqtt_message_bytes_sent_total[5m]))
# Average inbound payload size (bytes per message)
sum(rate(mqtt_message_bytes_received_total[5m])) / sum(rate(mqtt_messages_received_total[5m]))Delivery latency
The broker's internal delivery time indicates how long a message resides in the broker before it goes out to a subscriber. It does not include the network time between broker and client. Watch the p95/p99: a rising tail under constant load indicates broker or downstream pressure (such as slow subscribers or NATS backpressure). For clusters, when the publisher and subscriber connect to different nodes, the measurement spans the inter-node NATS hop and includes any clock skew between nodes.
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_publish_e2e_latency_seconds | Broker-side latency from receiving a PUBLISH to handing it to a subscriber's connection, by QoS. Buckets range from 0.5 ms to 5 s. | Histogram | qos |
Example queries:
# p99 delivery latency (99% of messages arrive faster than this)
histogram_quantile(0.99, sum by (le) (rate(mqtt_publish_e2e_latency_seconds_bucket[5m])))
# p95 latency per QoS level — expect QoS 1/2 to be slower than QoS 0
histogram_quantile(0.95, sum by (le, qos) (rate(mqtt_publish_e2e_latency_seconds_bucket[5m])))
# Mean latency (useful as a trend line, hides the tail)
sum(rate(mqtt_publish_e2e_latency_seconds_sum[5m])) / sum(rate(mqtt_publish_e2e_latency_seconds_count[5m]))
# Fraction of messages delivered within 10 ms (SLO-style check)
sum(rate(mqtt_publish_e2e_latency_seconds_bucket{le="0.01"}[5m]))
/ sum(rate(mqtt_publish_e2e_latency_seconds_count[5m]))Message drops and rejections
In a healthy system, reasons stay flat at zero. Any non-zero rate is message loss or rejection and identifies the cause directly. The reason label points at either a misbehaving client (packet_too_large, acl_denied, or quota_exceeded) or a consumption problem (queue_full, inflight_expired, or packet_id_exhausted).
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_messages_dropped_total | Total messages the broker dropped or rejected, by reason. The broker pre-initializes all reason series to 0, so a flat zero line means "nothing dropped" (not "no data"). | Counter | reason |
The following table shows what each drop reason indicates.
reason | Meaning | Indicates |
|---|---|---|
queue_full | A subscriber's outbound queue was full. | That subscriber is consuming slower than messages arrive for it (slow consumer). |
inflight_expired | An in-flight QoS 1/2 message expired or the broker abandoned it before acknowledgment. | Subscribers are not acknowledging in time possibly due to disconnected, stuck, or overloaded clients. |
packet_id_exhausted | The broker aborted a QoS delivery because the client's send window had no free packet IDs. | The client's entire QoS window sits in unacknowledged messages. |
acl_denied | An access-control rule rejected a publish. | A client attempted to publish to a topic for which it lacks permission (also see mqtt_acl_denied_total). |
packet_too_large | A packet exceeded the configured maximum packet size. | A client is sending payloads above the configured limit. |
quota_exceeded | Receive-maximum or quota enforcement rejected a packet. | A client is publishing faster than its negotiated QoS receive window allows. |
Example queries:
# Drop rate by reason — the primary "is anything being lost?" panel
sum by (reason) (rate(mqtt_messages_dropped_total[5m]))
# Total messages dropped in the last hour, by reason
sum by (reason) (increase(mqtt_messages_dropped_total[1h]))
# Drop percentage relative to delivery volume
100 * sum(rate(mqtt_messages_dropped_total[5m])) / sum(rate(mqtt_messages_sent_total[5m]))
# Alert condition: any drops at all in the last 15 minutes
sum(increase(mqtt_messages_dropped_total[15m])) > 0QoS reliability and backpressure
These metrics reveal QoS delivery health before messages start dropping. Resends mean subscribers failed to acknowledge within the retry interval, which is the earliest sign of a struggling consumer or lossy network. A steadily climbing mqtt_inflight_messages means acknowledgments are not keeping up with deliveries. The utilization histogram and saturated-clients gauge identify how close clients are to their send-window limit. Clients that sit near 1.0 stop receiving new QoS messages entirely until they acknowledge, which surfaces as packet_id_exhausted drops if it persists.
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_qos_resends_total | Total QoS 1/2 PUBLISH retransmissions to subscribers. | Counter | qos (1/2) |
mqtt_inflight_messages | QoS 1/2 messages currently in flight. The broker has sent these and is waiting for acknowledgment. | Gauge | — |
mqtt_client_inflight_utilization | Distribution of per-client outbound QoS-window utilization (consumed ÷ negotiated) across all connected clients. 1.0 means a client's send window is fully exhausted. | Histogram | — |
mqtt_clients_send_quota_saturated | Number of connected clients whose outbound QoS window is at least 25% consumed. | Gauge | — |
Example queries:
# Resend ratio: retransmissions as a fraction of QoS 1/2 deliveries
sum(rate(mqtt_qos_resends_total[5m]))
/ sum(rate(mqtt_messages_sent_total{qos=~"1|2"}[5m]))
# In-flight backlog trend — sustained growth means acks are falling behind
sum(mqtt_inflight_messages)
# Number of clients close to their send-window limit right now
sum(mqtt_clients_send_quota_saturated)
# Fraction of client samples with more than 90% of the QoS window consumed
1 - (
sum(rate(mqtt_client_inflight_utilization_bucket{le="0.9"}[5m]))
/ sum(rate(mqtt_client_inflight_utilization_count[5m]))
)Connections and sessions
mqtt_clients_connected is the basic fleet-size or availability signal. The disconnect reason breakdown separates normal churn (client_initiated) from problems:
- A spike in keepalive_timeout or network_error points at network instability or dying devices
- Recurring takeover means two devices share a client ID (a common fleet misconfiguration)
- A protocol_error indicates a broken client implementation.
The connects_total labels show reconnect behavior. A high rate with session_present="true" is a reconnect storm of existing clients rather than fleet growth.
Metric | Descripiton | Type | Labels |
|---|---|---|---|
mqtt_clients_connected | Currently connected MQTT clients. | Gauge | — |
mqtt_clients_disconnected | Disconnected persistent sessions the broker is retaining (clients that connected with a session expiry and may return). | Gauge | — |
mqtt_connects_total | Client connections established, by clean-session flag and whether the client resumed an existing session. | Counter | clean, session_present |
mqtt_disconnects_total | Client disconnections, by reason (see below). | Counter | reason |
mqtt_sessions_expired_total | Persistent sessions the broker removed after they reached their expiry. | Counter | — |
Disconnect reasons:
reason | Meaning |
|---|---|
client_initiated | Normal, client-requested disconnect. |
keepalive_timeout | The client stopped sending within its keepalive interval (silent death, network partition). |
takeover | Another client connected with the same client ID and took over the session. |
protocol_error | Malformed packet or protocol violation. |
server_initiated | The broker closed the connection (e.g. shutdown). |
network_error | Transport-level error (connection reset, TLS failure, etc.). |
other | Any cause not covered above. |
Example queries:
# Connected fleet size across the cluster (and per node with: by (instance))
sum(mqtt_clients_connected)
# Abnormal disconnect rate by reason — excludes normal client-requested disconnects
sum by (reason) (rate(mqtt_disconnects_total{reason!="client_initiated"}[5m]))
# Connect rate — a sudden spike is a reconnect storm or a fleet rollout
sum(rate(mqtt_connects_total[1m]))
# Share of connections that resumed an existing session
sum(rate(mqtt_connects_total{session_present="true"}[5m]))
/ sum(rate(mqtt_connects_total[5m]))
# Session takeovers in the last hour — duplicate client IDs in the fleet
sum(increase(mqtt_disconnects_total{reason="takeover"}[1h]))Authentication and access control
A rising outcome="failure" rate indicates clients with bad credentials including misconfigured devices, expired tokens, or a brute-force attempts. The cached label shows how much authentication load hits the database versus the cache. A low cache ratio under high connect rates puts pressure on PostgreSQL. mqtt_acl_denied_total shows clients trying to publish or subscribe outside their permitted topics. This is usually a misconfigured device or an incorrectly scoped ACL rule. Investigate either case.
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_auth_total | Authentication attempts, by outcome and whether the broker validated the credential from the in-memory auth cache rather than the database. | Counter | outcome (success/failure), cached |
mqtt_acl_denied_total | Denied ACL permission checks by action. | Counter | action (publish/subscribe) |
Example queries:
# Authentication failure rate (attempts/sec)
sum(rate(mqtt_auth_total{outcome="failure"}[5m]))
# Failure percentage of all auth attempts — spikes suggest bad credentials or an attack
100 * sum(rate(mqtt_auth_total{outcome="failure"}[5m])) / sum(rate(mqtt_auth_total[5m]))
# Auth cache hit ratio — a low value under load means the database is doing the work
sum(rate(mqtt_auth_total{outcome="success",cached="true"}[5m]))
/ sum(rate(mqtt_auth_total{outcome="success"}[5m]))
# ACL denials by action (publish vs subscribe) over the last hour
sum by (action) (increase(mqtt_acl_denied_total[1h]))Subscriptions and topics
Subscription counts track consumer-side scale. Wildcard subscriptions get their own metric because each one matches many topics and multiplies fan-out cost. A rising wildcard count with rising delivery latency is a common correlation. mqtt_topics_tracked reflects the size of the topic namespace. Unbounded growth usually means clients are embedding unique values (timestamps or UUIDs) in topic names.
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_subscriptions | Currently active subscriptions on the broker. | Gauge | — |
mqtt_subscribe_total | Subscriptions the broker granted, split by whether the topic filter contains a wildcard (+/#). | Counter | wildcard (true/false) |
mqtt_unsubscribe_total | Unsubscriptions, with the same split. | Counter | wildcard |
mqtt_wildcard_subscriptions | Active wildcard subscriptions across connected clients. | Gauge | — |
mqtt_topics_tracked | Topics the broker currently tracks in its local topic tree. | Gauge | — |
Example queries:
# Active subscriptions across the cluster
sum(mqtt_subscriptions)
# Wildcard share of active subscriptions — high values multiply fan-out cost
sum(mqtt_wildcard_subscriptions) / sum(mqtt_subscriptions)
# Topic namespace growth over the last 24 hours — sustained growth suggests
# unique values (timestamps, UUIDs) leaking into topic names
sum(delta(mqtt_topics_tracked[24h]))
# Subscribe churn rate, split by wildcard vs exact filters
sum by (wildcard) (rate(mqtt_subscribe_total[5m]))Retained messages
Retained messages persist indefinitely, the broker delivers them to every new matching subscriber, and the gauge tracks storage growth in the shared store. Comparing the add rate against a flat gauge distinguishes "clients refreshing retained values on the same topics" (healthy) from "retained messages accumulating on ever-new topics" (growth to watch).
Metric | Description | Type | Labels |
|---|---|---|---|
mqtt_retained_messages | Active retained messages in the cluster-shared store. This value is cluster-global: every node reports the same number, so aggregate it with max, not sum, in cluster queries. | Gauge | — |
mqtt_retained_added_total | Retained messages that clients set on this broker instance. | Counter | — |
Example queries:
# Retained messages in the cluster — max, NOT sum (every node reports the same
# cluster-global value; summing would multiply it by the node count)
max(mqtt_retained_messages)
# Rate at which clients set retained messages (per node, summed)
sum(rate(mqtt_retained_added_total[5m]))
# Net growth of the retained store over the last hour — near zero while the add
# rate is non-zero means existing topics are being refreshed, not accumulating
max(mqtt_retained_messages) - max(mqtt_retained_messages offset 1h)Connector (bridge) metrics
connector_connected is the connector health signal. 0 for an enabled connector means the remote endpoint is unreachable. A repeatedly resetting uptime_seconds or a climbing reconnect_attempts_total means the link is flapping. queue_depth is the backpressure indicator. A persistently growing queue means the remote endpoint cannot absorb the local publish rate and queue_full drops follow if it saturates. The forwarded counters, split by direction, verify data is actually flowing each way across the bridge.
These cover the MQTT-to-MQTT and Kafka connectors (bridges). Every metric carries a connector_id label enabling you to track each configured connector independently.
Metric | Description | Type | Labels |
|---|---|---|---|
connector_connected | 1 if the connector is currently connected to its remote endpoint, 0 otherwise. | Gauge | connector_id |
connector_uptime_seconds | Seconds since the connector last (re)connected. | Gauge | connector_id |
connector_messages_forwarded_total | Messages forwarded across the bridge. direction="in" is remote → local broker (ingress); direction="out" is local broker → remote (egress). | Counter | connector_id, direction |
connector_messages_dropped_total | Messages the connector dropped: loop (message would echo back to its origin), oversized (exceeds the remote's size limit), queue_full (connector's outbound buffer full). | Counter | connector_id, reason |
connector_routing_errors_total | Transient errors while routing a message (decode, marshal, or publish failures). | Counter | connector_id |
connector_reconnect_attempts_total | Reconnect attempts to the remote endpoint. | Counter | connector_id |
connector_queue_depth | Current depth of the connector's outbound queue. | Gauge | connector_id |
connector_last_error_timestamp_seconds | Unix timestamp of the connector's most recent error; 0 if no error has occurred. | Gauge | connector_id |
Example queries:
# Which connectors are currently down
connector_connected == 0
# Flapping detection: reconnect attempts in the last 15 minutes, per connector
sum by (connector_id) (increase(connector_reconnect_attempts_total[15m]))
# Throughput per connector and direction (in = remote→local, out = local→remote)
sum by (connector_id, direction) (rate(connector_messages_forwarded_total[5m]))
# Backpressure: connectors whose outbound queue is growing over the last 10 minutes
deriv(connector_queue_depth[10m]) > 0
# Connectors that logged an error in the last 5 minutes
(time() - connector_last_error_timestamp_seconds < 300)
and (connector_last_error_timestamp_seconds > 0)Process and runtime metrics
These metrics separate application-level symptoms from resource-level causes. Rising delivery latency alongside rising GC pause time or CPU saturation indicates the broker instance needs more resources. Latency growth with flat resource usage points at a downstream cause (slow consumers, NATS, or database).
The metrics endpoint also serves the standard Prometheus collectors:
- go_* Go runtime health
- memory in use (go_memstats_*)
- goroutine count (go_goroutines)
- garbage-collection pause durations (go_gc_duration_seconds)
- process_* OS-level process stats
- CPU time (process_cpu_seconds_total)
- resident memory (process_resident_memory_bytes)
- open file descriptors (process_open_fds each MQTT connection consumes one. Monitor against process_max_fds at high connection counts).
- go_build_info The Go module and version of the running binary.
Example queries:
# CPU usage per broker instance (1.0 = one full core)
rate(process_cpu_seconds_total[5m])
# Resident memory per instance
process_resident_memory_bytes
# File-descriptor headroom — each MQTT connection consumes one descriptor
process_open_fds / process_max_fds
# Goroutine count — a steady climb with flat connection count suggests a leak
go_goroutines