Monitoring and Troubleshooting
The node writes structured logs to stdout and serves a health endpoint and Prometheus metrics on port 9090. Together they answer the three questions that matter in production: is it connected, is it publishing, and is anything being dropped.
Logs
Logs are JSON, one object per line, on stdout; the container runtime owns retention. Set the level with LE_LOGGING_LEVEL. At info the node logs startup and session state, connectivity, birth-set summaries, host application state changes, and shutdown. Per-device births, Litmus Edge events, subscription activity, and buffering progress are debug.
Every line carries a component field, which is the fastest way to filter:
component | Covers |
|---|---|
lifecycle | Startup and shutdown |
config | Configuration loading |
nats | The Litmus Edge NATS broker connection |
mqtt | Broker connection and reconnects |
sparkplug | Birth sets, publishing, commands |
forward | Store and forward |
observability | The metrics server |
Other fields appear where they apply: node, group, device, tag, topic, seq, bdSeq, server, and count.
{"component":"sparkplug","dbirth_published":4,"dbirth_skipped":0,"devices":4,"duration_ms":120,"group":"Plant1","level":"info","msg":"birth set published","node":"Line1","status":"success","time":"2026-07-24T10:30:00.123456789Z"}Health Endpoint
curl -i http://<host>:9090/healthz200 with body ok means the node has a connected MQTT session — and, when host application gating is enabled, that the host application is online. Anything else returns 503 with body unavailable.
This is readiness rather than liveness: a healthy process waiting for its host application reports 503 by design. If your orchestrator distinguishes the two, wire this endpoint to readiness.
Metrics
Prometheus metrics are served at /metrics on the same port, all prefixed le_sparkplug_. Both endpoints are unauthenticated, so bind them to an internal interface or front them with a proxy. Turn them off entirely with OBS_ENABLED=false, or move them with OBS_ADDR.
Publishing
Metric | Type | Meaning |
|---|---|---|
nbirth_total | Counter | NBIRTH messages published. Climbing steadily means a rebirth loop |
ndata_total, ddata_total | Counter | Data messages published |
ddata_metrics_total | Counter | Individual metrics inside those data messages |
ddata_batch_size | Histogram | Metrics per data message |
mqtt_publish_duration_seconds | Histogram | Publish duration, labelled by message_type |
Connection and Session
Metric | Type | Meaning |
|---|---|---|
mqtt_reconnect_total | Counter | Reconnect attempts |
pha_state | Gauge | Host application state: -1 unknown, 0 offline, 1 online |
nats_decode_errors_total | Counter | Messages from Litmus Edge the node could not decode |
api_request_duration_seconds | Histogram | Litmus Edge API calls, labelled by endpoint and status |
operation_duration_seconds | Histogram | Internal operations, labelled by operation |
Commands
Metric | Type | Meaning |
|---|---|---|
ncmd_total | Counter | Node commands, labelled type: rebirth, next_server, unknown |
dcmd_total | Counter | Device commands received |
rebirth_suppressed_total | Counter | Rebirth requests dropped because one was already running, labelled reason |
Store and Forward
Metric | Type | Meaning |
|---|---|---|
sf_backlog_size | Gauge | Messages waiting on disk |
sf_dropped_total | Counter | Buffered messages discarded instead of replayed, labelled reason |
Sources
Metric | Type | Meaning |
|---|---|---|
devicehub_entity_count | Gauge | Latest counts, labelled entity: devices, devices_published, registers, assigned_tags |
devicehub_devices_filtered_total | Counter | Devices dropped by a publish filter, labelled reason: source_disabled, exclude, not_in_include |
digital_twins_instances_filtered_total | Counter | Twin instances dropped by a publish filter, labelled reason |
digital_twin_quarantine_total | Counter | Twin instances skipped for want of a usable contract, labelled reason |
digital_twin_attribute_fallback_total | Counter | Attributes registered as String for want of a type signal |
digital_twin_attribute_inference_total | Counter | Attributes typed from a non-DeviceHub source |
digital_twin_attribute_quarantine_total | Counter | Attributes dropped while planning the contract |
digital_twin_attribute_value_mismatch_total | Counter | Runtime values that did not match the birth datatype |
digital_twin_attribute_value_skipped_total | Counter | Values skipped after such a mismatch |
Scrape Configuration
scrape_configs:
- job_name: le-sparkplug
static_configs:
- targets: ['le-sparkplug:9090']
scrape_interval: 15sWorth Alerting On
Expression | What it means |
|---|---|
rate(le_sparkplug_mqtt_reconnect_total[5m]) > 0.1 | Broker instability, or host application state flapping |
le_sparkplug_pha_state == 0 for 5 minutes | The host application is down, and data is being buffered or dropped |
le_sparkplug_sf_backlog_size > 10000 | The drain is not keeping up with the outage |
increase(le_sparkplug_sf_dropped_total[15m]) > 0 | Buffered data was discarded, which happens for a removed or filtered device |
rate(le_sparkplug_nats_decode_errors_total[5m]) > 0 | Unreadable messages from Litmus Edge |
increase(le_sparkplug_nbirth_total[15m]) > 5 | Repeated rebirths, so host applications keep rebuilding their tag lists |
increase(le_sparkplug_digital_twin_quarantine_total[15m]) > 0 | A twin stopped publishing because its contract no longer compiles |
increase(le_sparkplug_digital_twin_attribute_value_mismatch_total[15m]) > 0 | Runtime values disagree with a declared datatype |
Troubleshooting
Start with the node's logs and, where a symptom involves the broker, a capture of the Sparkplug namespace. Setting LE_LOGGING_LEVEL=debug and restarting adds per-device and per-event detail.
mosquitto_sub -h <broker host> -t 'spBv1.0/#' -F '%t'No Node Appears in the Host Application
Read the logs in order. Each component logs connected when its half of the connection succeeds.
Log state | Cause | Fix |
|---|---|---|
No connected from nats | The node cannot reach the Litmus Edge NATS broker | Inside Litmus Edge, check the Docker gateway, or set EDGE_DOCKER_GATEWAY_IP. Outside, check the API Account key and that the NATS proxy is enabled. See Deployment |
No connected from mqtt | The broker is unreachable, or credentials or certificates are rejected | Check MQTT_URL, credentials, and the TLS settings in Deployment |
waiting for primary host application and nothing after | The host application has not reported online | Confirm PRIMARY_HOST_ID matches the host exactly, and clear any retained STATE message left from testing. See Configuration |
birth set published with devices as 0 | The node connected but Litmus Edge returned no devices | Confirm devices exist and are enabled in DeviceHub, and that filters are not excluding them |
A Device or Tag Is Missing
- A whole device. Check the publish filters. le_sparkplug_devicehub_devices_filtered_total names the reason; compare devicehub_entity_count{entity="devices"} with {entity="devices_published"}.
- One tag on a device. Either its value type is not one the node maps, or its name has no alphanumeric characters. Both are logged; the second logs invalid tag name with the tag. See Sparkplug Data.
- A removed tag still visible in the host application. The node publishes DDEATH before a shrunk tag set for exactly this reason. If the stale tag persists, request a rebirth, or confirm the host processed the DDEATH.
A Digital Twin Publishes No Data
- Is there a DBIRTH for the twin at all? If not, the instance is disabled, has no topic, or its contract compiled to nothing — check le_sparkplug_digital_twin_quarantine_total.
- Does the DBIRTH contain Attributes/… metrics, or only Device Properties/Digital Twin/… and Device Control/Rebirth? Only the latter means no dynamic attribute produced a metric.
- Are values arriving but not appearing? Check le_sparkplug_digital_twin_attribute_value_mismatch_total, whose labels name the metric, the expected datatype, and what arrived.
Every Digital Twin Attribute Is a String
No type signal reached the node, so it registered the honest fallback rather than guessing. Confirm it with the typeSource property on the birth metrics, which reads untyped, and count the affected attributes with le_sparkplug_digital_twin_attribute_fallback_total. Sparkplug Data explains the fix in the model editor.
Data Stops After a Reconnect
Check whether the node is publishing at all: le_sparkplug_ddata_total should still be climbing.
- Climbing, but the host application shows nothing — the host is rejecting the data, and its own logs will say why.
- Flat, with le_sparkplug_pha_state at 0 — the host application went offline and the node stopped publishing by design.
- Flat, with reconnects climbing — the session is cycling. Check the broker and, with several brokers configured, whether the node is walking the list.
The Host Application Reports a Sequence Gap
Almost always two nodes sharing one Sparkplug identity, whose sequence counters interleave. Confirm only one container runs with your Group ID and Node ID — including any leftover from an upgrade — then restart the surviving one so the host sees a clean session.
The Node Reads as Stale After Every Restart
The data directory is not persistent, so the session counter restarts and the host reads each new session as out of order. Mount /le-sparkplug/data on a volume and keep it across upgrades. See Deployment.
Duplicate Data After a Restart
Expected after an unclean stop with store and forward enabled. A buffered message is deleted only after the broker acknowledges it, so a message published just before the process died is replayed once. Replayed metrics are marked historical, which is how a host application distinguishes them.
Buffered Data Never Arrives
- le_sparkplug_sf_backlog_size stays flat: the node still cannot publish. Check the broker connection and, if gating is enabled, host application state.
- le_sparkplug_sf_dropped_total{reason="unbirthed_device"} is climbing: the data belongs to a device with no birth certificate in this session, because it was removed in Litmus Edge or excluded by a filter. That data is discarded by design.
- The backlog fell without data arriving: messages may have passed their LE_TTL_DURATION.
Certificate Errors on the Broker Connection
The node logs certificate problems and keeps running. Check the mqtt component against the message table in Deployment, which names each decode and key-pair failure and its cause.
The Metrics Endpoint Is Unreachable
OBS_ENABLED may be false, the port may not be published from the container, or OBS_ADDR may have moved it. Verify from inside the container first, then check the port mapping.
What to Include in a Support Request
- The node version, and whether it runs inside or outside Litmus Edge
- Litmus Edge version
- Logs at debug covering a restart
- A broker capture of spBv1.0/# covering the same window
- The host application's own logs, when it is involved
- Your Group ID and Node ID, and the number of devices and twins expected