Configuration
Every setting can be supplied as an environment variable or as a field in a JSON file. Environment variables override file values key by key.
The Marketplace application exposes the settings most deployments need as form parameters, and applies the defaults below to the rest. Settings it does not expose are set in the configuration file, or as an environment variable on a container you manage yourself: multiple brokers, publish filters, client ID and timeout overrides, and buffer-file rotation limits.
The node reads config/config.json under its configuration path, which is /le-sparkplug/data unless CONFIG_PATH says otherwise. A missing file is not an error — the node then runs on environment values and defaults alone. Only EDGE_API_TOKEN has no default.
Configuration is read once at startup. There is no reload — restart the container to apply a change.
Litmus Edge Connection
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
EDGE_API_TOKEN | edge_api_token | none, required | Authenticates to the Litmus Edge REST and GraphQL APIs |
EDGE_EXTERNAL | edge_external | false | true when the node runs outside Litmus Edge |
EDGE_HOSTNAME | edge_hostname | the detected Docker gateway | Address of the Litmus Edge instance. Required when EDGE_EXTERNAL=true |
EDGE_DOCKER_GATEWAY_IP | edge_docker_gateway_ip | read from the container's default route | Docker gateway used to reach Litmus Edge from inside the device |
EDGE_ACCESS_ACCOUNT_API_KEY | edge_access_account_api_key | none | API Account key. In external mode the node builds tls://admin:<key>@<EDGE_HOSTNAME>:4222 from it |
NATS_URL | nats_url | generated | Explicit NATS broker URL. Overrides the generated one |
CONFIG_PATH | config_path | /le-sparkplug/data | Base path for the configuration file and persistent state |
Sparkplug Identity
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
GROUP_ID | group_id | le-sparkplug | Sparkplug Group ID |
NODE_ID | node_id | the container hostname | Sparkplug Node ID, unique within the group |
/, +, #, and whitespace in either value are replaced with _, because those characters are reserved in MQTT topics.
MQTT Broker
These fields describe a single broker and are ignored when MQTT_SERVERS is set. For the certificate workflow, see Deployment.
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
MQTT_URL | mqtt.servers[].url | tcp://mqtt:1883 | Broker URL. Schemes: tcp, ssl, mqtt, mqtts |
MQTT_USERNAME or MQTT_USER | mqtt.servers[].user | empty | Broker username |
MQTT_PASSWORD | mqtt.servers[].password | empty | Broker password |
MQTT_CA_CERT | mqtt.servers[].ca_cert | empty | Base64-encoded CA certificate |
MQTT_CLIENT_CERT | mqtt.servers[].client_cert | empty | Base64-encoded client certificate |
MQTT_CLIENT_KEY | mqtt.servers[].client_key | empty | Base64-encoded client private key |
MQTT_SKIP_CERT_VERIFY | mqtt.servers[].skip_cert_verify | false | Accepts any broker certificate |
MQTT Client and Failover
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
MQTT_SERVERS | mqtt.servers | unset | JSON array of broker objects. Replaces the single-broker fields, and enables failover across brokers |
MQTT_CLIENT_ID | mqtt.client_id | <GROUP_ID>-<NODE_ID> | MQTT client identifier |
MQTT_CLIENT_ID_SUFFIX | mqtt.client_id_suffix | empty | Appended to the client ID after _ |
MQTT_CONNECT_TIMEOUT | mqtt.connect_timeout | 30s | Time allowed for a connection attempt |
MQTT_WRITE_TIMEOUT | mqtt.write_timeout | 3s | Time allowed for a publish to complete |
MQTT_KEEPALIVE | mqtt.keepalive | 30 | Seconds of inactivity before a keepalive packet |
Diagnostics
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
OBS_ENABLED | obs_enabled | true | Serves /metrics and /healthz |
OBS_ADDR | obs_addr | :9090 | Bind address for those endpoints |
LE_LOGGING_LEVEL | le_logging_level | info | One of debug, info, warn, error |
Invalid Values
Most invalid values are logged and replaced with the default, so the node still starts. Timeouts, keepalive, log level, and the store-and-forward numbers all behave that way. These, by contrast, stop startup:
Condition | Message |
|---|---|
EDGE_API_TOKEN missing or blank | [EDGE_API_TOKEN] is UNDEFINED |
EDGE_EXTERNAL not a boolean | [EDGE_EXTERNAL] is Invalid |
PRIMARY_HOST_SUPPORT not a boolean | [PRIMARY_HOST_SUPPORT] is Invalid |
LE_STORE_AND_FORWARD not a boolean | [LE_STORE_AND_FORWARD] is Invalid |
OBS_ENABLED not a boolean | [OBS_ENABLED] is Invalid |
DH_ENABLED or DT_ENABLED not a boolean | [DH_ENABLED] is Invalid, [DT_ENABLED] is Invalid |
A filter pattern is not a valid glob | [DH_DEVICES_INCLUDE] …, and the equivalent for the other three |
MQTT_SERVERS is not valid JSON | [Invalid JSON Array [MQTT_SERVERS]] |
MQTT_SERVERS contains an entry without a URL | [MQTT_SERVERS] contains no valid server URLs |
EDGE_HOSTNAME missing in external mode | [EDGE_HOSTNAME] is UNDEFINED for external mode |
EDGE_ACCESS_ACCOUNT_API_KEY missing in external mode with no NATS_URL | [EDGE_ACCESS_ACCOUNT_API_KEY] is UNDEFINED for external mode when [NATS_URL] is unset |
Configuration File
{
"edge_api_token": "<Litmus Edge API token>",
"group_id": "Plant1",
"node_id": "Line1",
"mqtt": {
"servers": [
{ "url": "ssl://broker-a.example.com:8883", "user": "sparkplug", "password": "<password>", "ca_cert": "<Base64 CA certificate>" },
{ "url": "tcp://broker-b.example.com:1883" }
],
"client_id_suffix": "line1",
"connect_timeout": "30s",
"write_timeout": "3s",
"keepalive": 30
},
"primary_host_support": true,
"primary_host_id": "scada-host",
"le_store_and_forward": true,
"le_ttl_duration": 24,
"devicehub": {
"enabled": true,
"exclude": ["test_*"]
},
"digital_twins": {
"enabled": true
},
"obs_enabled": true,
"le_logging_level": "info"
}Primary Host Application
A Sparkplug host application publishes its own state to the broker so edge nodes know whether anyone is listening. With primary host support enabled, the node publishes nothing until that host reports online, and stops the moment it reports offline. Use it when a single host application owns the data; leave it off when several applications consume the node independently.
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
PRIMARY_HOST_SUPPORT | primary_host_support | false | Waits for the host application before publishing |
PRIMARY_HOST_ID | primary_host_id | le-spb when support is enabled | Host application ID in the STATE topic. Must match exactly what the host publishes |
The node subscribes to spBv1.0/STATE/<PRIMARY_HOST_ID> and reads a JSON payload:
{"online":true,"timestamp":1778750000000}flowchart TD
SUB[Subscribe to STATE topic] --> WAIT[Wait]
WAIT --> ON{online?}
ON -- true --> BIRTH[Publish NBIRTH and DBIRTH, start data, drain buffer]
ON -- false --> END[Publish NDEATH, disconnect, move to next broker]On online, the node subscribes to Litmus Edge events, publishes the full birth set, starts streaming data, and drains any store-and-forward backlog.
On offline, it ends its MQTT session and moves to the next broker in its list, as the Sparkplug specification requires — this is how a Sparkplug node follows a host application that has failed over. With one broker configured it reconnects to the same one and waits again. Meanwhile no NBIRTH, DBIRTH, or DDATA is published, /healthz returns 503 unavailable, and device data is buffered when store and forward is enabled and dropped otherwise.
State messages are accepted only if their timestamp is newer than the last one seen; an older one is ignored and logged as PHA STATE timestamp conflict; ignoring. This protects against retained state arriving out of order after a reconnect. A fresh online message with a newer timestamp from an already-online host counts as a new host session, and the node republishes the whole birth set.
To confirm the handshake, look for primary host application online in the logs, le_sparkplug_pha_state at 1, and /healthz returning 200.
Store and Forward
When the node cannot publish, device data is dropped by default. Store and forward writes it to disk instead and republishes it once publishing resumes, marked so host applications do not mistake it for live values.
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
LE_STORE_AND_FORWARD | le_store_and_forward | false | Buffers device data on disk while publishing is blocked |
LE_TTL_DURATION | le_ttl_duration | 12 | Hours a buffered message stays valid. Expired messages are removed by the database, so a long outage discards the oldest data first |
LE_MAX_SIZE_MB | le_max_size_mb | 1000 | Size in MB of each on-disk buffer file before the database rotates to a new one |
LE_MAX_NUM_ENTRIES | le_max_num_entries | 1000000 | Entries in each buffer file before the database rotates to a new one |
Buffering starts whenever the node cannot publish — no connected MQTT session, or host application gating enabled with the host not online.
flowchart LR
TAG[Tag update] --> CAN{Can publish?}
CAN -- yes --> PUB[Publish DDATA]
CAN -- no --> SF{Store and forward on?}
SF -- yes --> DISK[Write to disk]
SF -- no --> DROP[Drop]
DISK --> DRAIN[Replay when publishing resumes]Only data messages are buffered. Birth and death certificates are session state, rebuilt on reconnect rather than replayed. The buffer lives at <group>-<node>_db/ under /le-sparkplug/data, which must be on the persistent volume described in Deployment.
Replay
A drain runs when the host application comes online, and every 5 minutes while a backlog exists. Messages leave disk in batches of 256 so a large backlog does not block live data. Each one is rewritten before it goes out:
- Its metrics are marked historical, so host applications record history without moving their live values
- Its datatype fields are stripped, matching the Sparkplug rule for data messages
- It receives the next sequence number of the current session, not the one it had when captured
A drain stops early if the connection drops again or the host application goes offline; whatever is left stays on disk for the next attempt.
A message is deleted from disk only after the broker acknowledges it, so a node that stops between the publish and the delete republishes that message on the next drain. Host applications must tolerate a duplicate, which the historical flag makes straightforward.
The node also drops stored messages it can no longer publish honestly: if a device has no birth certificate in the current session — removed in Litmus Edge, or excluded by a publish filter — its buffered messages are deleted rather than replayed, because data before birth breaks the Sparkplug contract. Those drops are counted by le_sparkplug_sf_dropped_total{reason="unbirthed_device"}.
To watch a cycle end to end, set LE_LOGGING_LEVEL=debug and stop the broker: buffered writes log stored from the forward component and le_sparkplug_sf_backlog_size climbs. Start the broker again and the node logs draining stored DDATA with a count field as the gauge falls to 0.
Publish Filters
By default the node publishes every DeviceHub device and every enabled Digital Twin instance it finds. Filters narrow that set, keeping test equipment out of production namespaces. They apply to whole devices and instances, not to individual tags or attributes.
Environment variable | config.json field | Default | Effect |
|---|---|---|---|
DH_ENABLED | devicehub.enabled | true | Publishes DeviceHub devices. false stops DeviceHub discovery entirely |
DH_DEVICES_INCLUDE | devicehub.include | empty | Glob patterns of device names to publish |
DH_DEVICES_EXCLUDE | devicehub.exclude | empty | Glob patterns of device names to skip |
DT_ENABLED | digital_twins.enabled | true | Publishes Digital Twin instances. false stops Digital Twin discovery entirely |
DT_INSTANCES_INCLUDE | digital_twins.include | empty | Glob patterns of instance names to publish |
DT_INSTANCES_EXCLUDE | digital_twins.exclude | empty | Glob patterns of instance names to skip |
Patterns match the DeviceHub device name or the Digital Twin instance name. As environment variables they are comma-separated; in a configuration file they are JSON lists. They support * for any sequence of characters, ? for a single character, and [abc] for a character class. Regular expressions are not supported, and a pattern that is not a valid glob stops startup.
The node applies them in this order:
- An exclude match drops the device. Exclude always wins over include.
- An empty include list admits everything that survived the excludes.
- A non-empty include list admits only names matching one of its patterns.
Goal | Setting |
|---|---|
Publish only production lines | DH_DEVICES_INCLUDE=line1_*,line2_* |
Publish everything except test rigs | DH_DEVICES_EXCLUDE=test_*,*_sandbox |
Publish one named device | DH_DEVICES_INCLUDE=press_002 |
Publish DeviceHub devices but no twins | DT_ENABLED=false |
Publish twins for one model's instances | DT_INSTANCES_INCLUDE=hydraulic_press_* |
A filtered device is absent, not marked offline: the node never publishes a DBIRTH for it, so host applications never learn it exists. Buffered data for a device you have just filtered out is discarded rather than replayed, for the reason given under Store and Forward above.
Filters are evaluated on every metadata refresh, so a device that starts matching an exclude pattern disappears from the next birth set — but because configuration is read only at startup, changing a filter takes a restart. le_sparkplug_devicehub_devices_filtered_total and le_sparkplug_digital_twins_instances_filtered_total count each drop with its reason, and the logs name each filtered device at debug level.
A mismatched `PRIMARY_HOST_ID` looks exactly like a broken deployment: the node connects to the broker, logs `waiting for primary host application`, and never publishes a birth certificate. Check the ID first, then clear any retained `STATE` message left from earlier testing by publishing an empty retained message to the same topic.`LE_MAX_SIZE_MB` and `LE_MAX_NUM_ENTRIES` size the buffer's individual files; they are not a total cap on disk use. The practical bound on the backlog is `LE_TTL_DURATION` together with the free space on the volume. Size that volume for your worst expected outage, and shorten the TTL if the data loses value quickly.Replay order is not capture order. Messages are keyed by a random identifier, so the backlog drains in arbitrary order. Each message still carries its original timestamps, which is what a historian records.