---
title: Configuration
slug: solutions/configuration
docTags: 
createdAt: 2026-08-31T22:46:13.619Z
---

Every setting can be supplied as an environment variable or as a field in a JSON file. Environment variables override file values key by key.

The Marketplace application exposes the settings most deployments need as form parameters, and applies the defaults below to the rest. Settings it does not expose are set in the configuration file, or as an environment variable on a container you manage yourself: multiple brokers, publish filters, client ID and timeout overrides, and buffer-file rotation limits.

The node reads `config/config.json` under its configuration path, which is `/le-sparkplug/data` unless `CONFIG_PATH` says otherwise. A missing file is not an error — the node then runs on environment values and defaults alone. Only `EDGE_API_TOKEN` has no default.

Configuration is read once at startup. There is no reload — restart the container to apply a change.

## Litmus Edge Connection

| Environment variable          | `config.json` field           | Default                                 | Effect                                                                                             |
| ----------------------------- | ----------------------------- | --------------------------------------- | -------------------------------------------------------------------------------------------------- |
| `EDGE_API_TOKEN`              | `edge_api_token`              | none, required                          | Authenticates to the Litmus Edge REST and GraphQL APIs                                             |
| `EDGE_EXTERNAL`               | `edge_external`               | `false`                                 | `true` when the node runs outside Litmus Edge                                                      |
| `EDGE_HOSTNAME`               | `edge_hostname`               | the detected Docker gateway             | Address of the Litmus Edge instance. Required when `EDGE_EXTERNAL=true`                            |
| `EDGE_DOCKER_GATEWAY_IP`      | `edge_docker_gateway_ip`      | read from the container's default route | Docker gateway used to reach Litmus Edge from inside the device                                    |
| `EDGE_ACCESS_ACCOUNT_API_KEY` | `edge_access_account_api_key` | none                                    | API Account key. In external mode the node builds `tls://admin:<key>@<EDGE_HOSTNAME>:4222` from it |
| `NATS_URL`                    | `nats_url`                    | generated                               | Explicit NATS broker URL. Overrides the generated one                                              |
| `CONFIG_PATH`                 | `config_path`                 | `/le-sparkplug/data`                    | Base path for the configuration file and persistent state                                          |

## Sparkplug Identity

| Environment variable | `config.json` field | Default                | Effect                                     |
| -------------------- | ------------------- | ---------------------- | ------------------------------------------ |
| `GROUP_ID`           | `group_id`          | `le-sparkplug`         | Sparkplug Group ID                         |
| `NODE_ID`            | `node_id`           | the container hostname | Sparkplug Node ID, unique within the group |

`/`, `+`, `#`, and whitespace in either value are replaced with `_`, because those characters are reserved in MQTT topics.

## MQTT Broker

These fields describe a single broker and are ignored when `MQTT_SERVERS` is set. For the certificate workflow, see [Deployment](#).

| Environment variable           | `config.json` field               | Default           | Effect                                             |
| ------------------------------ | --------------------------------- | ----------------- | -------------------------------------------------- |
| `MQTT_URL`                     | `mqtt.servers[].url`              | `tcp://mqtt:1883` | Broker URL. Schemes: `tcp`, `ssl`, `mqtt`, `mqtts` |
| `MQTT_USERNAME` or `MQTT_USER` | `mqtt.servers[].user`             | empty             | Broker username                                    |
| `MQTT_PASSWORD`                | `mqtt.servers[].password`         | empty             | Broker password                                    |
| `MQTT_CA_CERT`                 | `mqtt.servers[].ca_cert`          | empty             | Base64-encoded CA certificate                      |
| `MQTT_CLIENT_CERT`             | `mqtt.servers[].client_cert`      | empty             | Base64-encoded client certificate                  |
| `MQTT_CLIENT_KEY`              | `mqtt.servers[].client_key`       | empty             | Base64-encoded client private key                  |
| `MQTT_SKIP_CERT_VERIFY`        | `mqtt.servers[].skip_cert_verify` | `false`           | Accepts any broker certificate                     |

## MQTT Client and Failover

| Environment variable    | `config.json` field     | Default                | Effect                                                                                               |
| ----------------------- | ----------------------- | ---------------------- | ---------------------------------------------------------------------------------------------------- |
| `MQTT_SERVERS`          | `mqtt.servers`          | unset                  | JSON array of broker objects. Replaces the single-broker fields, and enables failover across brokers |
| `MQTT_CLIENT_ID`        | `mqtt.client_id`        | `<GROUP_ID>-<NODE_ID>` | MQTT client identifier                                                                               |
| `MQTT_CLIENT_ID_SUFFIX` | `mqtt.client_id_suffix` | empty                  | Appended to the client ID after `_`                                                                  |
| `MQTT_CONNECT_TIMEOUT`  | `mqtt.connect_timeout`  | `30s`                  | Time allowed for a connection attempt                                                                |
| `MQTT_WRITE_TIMEOUT`    | `mqtt.write_timeout`    | `3s`                   | Time allowed for a publish to complete                                                               |
| `MQTT_KEEPALIVE`        | `mqtt.keepalive`        | `30`                   | Seconds of inactivity before a keepalive packet                                                      |

## Diagnostics

| Environment variable | `config.json` field | Default | Effect                                  |
| -------------------- | ------------------- | ------- | --------------------------------------- |
| `OBS_ENABLED`        | `obs_enabled`       | `true`  | Serves `/metrics` and `/healthz`        |
| `OBS_ADDR`           | `obs_addr`          | `:9090` | Bind address for those endpoints        |
| `LE_LOGGING_LEVEL`   | `le_logging_level`  | `info`  | One of `debug`, `info`, `warn`, `error` |

## Invalid Values

Most invalid values are logged and replaced with the default, so the node still starts. Timeouts, keepalive, log level, and the store-and-forward numbers all behave that way. These, by contrast, stop startup:

| Condition                                                                 | Message                                                                                 |
| ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| `EDGE_API_TOKEN` missing or blank                                         | `[EDGE_API_TOKEN] is UNDEFINED`                                                         |
| `EDGE_EXTERNAL` not a boolean                                             | `[EDGE_EXTERNAL] is Invalid`                                                            |
| `PRIMARY_HOST_SUPPORT` not a boolean                                      | `[PRIMARY_HOST_SUPPORT] is Invalid`                                                     |
| `LE_STORE_AND_FORWARD` not a boolean                                      | `[LE_STORE_AND_FORWARD] is Invalid`                                                     |
| `OBS_ENABLED` not a boolean                                               | `[OBS_ENABLED] is Invalid`                                                              |
| `DH_ENABLED` or `DT_ENABLED` not a boolean                                | `[DH_ENABLED] is Invalid`, `[DT_ENABLED] is Invalid`                                    |
| A filter pattern is not a valid glob                                      | `[DH_DEVICES_INCLUDE] …`, and the equivalent for the other three                        |
| `MQTT_SERVERS` is not valid JSON                                          | `[Invalid JSON Array [MQTT_SERVERS]]`                                                   |
| `MQTT_SERVERS` contains an entry without a URL                            | `[MQTT_SERVERS] contains no valid server URLs`                                          |
| `EDGE_HOSTNAME` missing in external mode                                  | `[EDGE_HOSTNAME] is UNDEFINED for external mode`                                        |
| `EDGE_ACCESS_ACCOUNT_API_KEY` missing in external mode with no `NATS_URL` | `[EDGE_ACCESS_ACCOUNT_API_KEY] is UNDEFINED for external mode when [NATS_URL] is unset` |

## Configuration File

```json
{
  "edge_api_token": "<Litmus Edge API token>",
  "group_id": "Plant1",
  "node_id": "Line1",
  "mqtt": {
    "servers": [
      { "url": "ssl://broker-a.example.com:8883", "user": "sparkplug", "password": "<password>", "ca_cert": "<Base64 CA certificate>" },
      { "url": "tcp://broker-b.example.com:1883" }
    ],
    "client_id_suffix": "line1",
    "connect_timeout": "30s",
    "write_timeout": "3s",
    "keepalive": 30
  },
  "primary_host_support": true,
  "primary_host_id": "scada-host",
  "le_store_and_forward": true,
  "le_ttl_duration": 24,
  "devicehub": {
    "enabled": true,
    "exclude": ["test_*"]
  },
  "digital_twins": {
    "enabled": true
  },
  "obs_enabled": true,
  "le_logging_level": "info"
}
```

## Primary Host Application

A Sparkplug host application publishes its own state to the broker so edge nodes know whether anyone is listening. With primary host support enabled, the node publishes nothing until that host reports online, and stops the moment it reports offline. Use it when a single host application owns the data; leave it off when several applications consume the node independently.

| Environment variable   | `config.json` field    | Default                          | Effect                                                                               |
| ---------------------- | ---------------------- | -------------------------------- | ------------------------------------------------------------------------------------ |
| `PRIMARY_HOST_SUPPORT` | `primary_host_support` | `false`                          | Waits for the host application before publishing                                     |
| `PRIMARY_HOST_ID`      | `primary_host_id`      | `le-spb` when support is enabled | Host application ID in the `STATE` topic. Must match exactly what the host publishes |

The node subscribes to `spBv1.0/STATE/<PRIMARY_HOST_ID>` and reads a JSON payload:

```json
{"online":true,"timestamp":1778750000000}
```

```mermaid
flowchart TD
    SUB[Subscribe to STATE topic] --> WAIT[Wait]
    WAIT --> ON{online?}
    ON -- true --> BIRTH[Publish NBIRTH and DBIRTH, start data, drain buffer]
    ON -- false --> END[Publish NDEATH, disconnect, move to next broker]
```

On **online**, the node subscribes to Litmus Edge events, publishes the full birth set, starts streaming data, and drains any store-and-forward backlog.

On **offline**, it ends its MQTT session and moves to the next broker in its list, as the Sparkplug specification requires — this is how a Sparkplug node follows a host application that has failed over. With one broker configured it reconnects to the same one and waits again. Meanwhile no `NBIRTH`, `DBIRTH`, or `DDATA` is published, `/healthz` returns `503 unavailable`, and device data is buffered when store and forward is enabled and dropped otherwise.

State messages are accepted only if their timestamp is newer than the last one seen; an older one is ignored and logged as `PHA STATE timestamp conflict; ignoring`. This protects against retained state arriving out of order after a reconnect. A fresh `online` message with a newer timestamp from an already-online host counts as a new host session, and the node republishes the whole birth set.

To confirm the handshake, look for `primary host application online` in the logs, `le_sparkplug_pha_state` at `1`, and `/healthz` returning `200`.

## Store and Forward

When the node cannot publish, device data is dropped by default. Store and forward writes it to disk instead and republishes it once publishing resumes, marked so host applications do not mistake it for live values.

| Environment variable   | `config.json` field    | Default   | Effect                                                                                                                              |
| ---------------------- | ---------------------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `LE_STORE_AND_FORWARD` | `le_store_and_forward` | `false`   | Buffers device data on disk while publishing is blocked                                                                             |
| `LE_TTL_DURATION`      | `le_ttl_duration`      | `12`      | Hours a buffered message stays valid. Expired messages are removed by the database, so a long outage discards the oldest data first |
| `LE_MAX_SIZE_MB`       | `le_max_size_mb`       | `1000`    | Size in MB of each on-disk buffer file before the database rotates to a new one                                                     |
| `LE_MAX_NUM_ENTRIES`   | `le_max_num_entries`   | `1000000` | Entries in each buffer file before the database rotates to a new one                                                                |

Buffering starts whenever the node cannot publish — no connected MQTT session, or host application gating enabled with the host not online.

```mermaid
flowchart LR
    TAG[Tag update] --> CAN{Can publish?}
    CAN -- yes --> PUB[Publish DDATA]
    CAN -- no --> SF{Store and forward on?}
    SF -- yes --> DISK[Write to disk]
    SF -- no --> DROP[Drop]
    DISK --> DRAIN[Replay when publishing resumes]
```

Only data messages are buffered. Birth and death certificates are session state, rebuilt on reconnect rather than replayed. The buffer lives at `<group>-<node>_db/` under `/le-sparkplug/data`, which must be on the persistent volume described in [Deployment](#).

### Replay

A drain runs when the host application comes online, and every 5 minutes while a backlog exists. Messages leave disk in batches of 256 so a large backlog does not block live data. Each one is rewritten before it goes out:

- Its metrics are marked historical, so host applications record history without moving their live values
- Its datatype fields are stripped, matching the Sparkplug rule for data messages
- It receives the next sequence number of the current session, not the one it had when captured

A drain stops early if the connection drops again or the host application goes offline; whatever is left stays on disk for the next attempt.

A message is deleted from disk only after the broker acknowledges it, so a node that stops between the publish and the delete republishes that message on the next drain. Host applications must tolerate a duplicate, which the historical flag makes straightforward.

The node also drops stored messages it can no longer publish honestly: if a device has no birth certificate in the current session — removed in Litmus Edge, or excluded by a publish filter — its buffered messages are deleted rather than replayed, because data before birth breaks the Sparkplug contract. Those drops are counted by `le_sparkplug_sf_dropped_total{reason="unbirthed_device"}`.

To watch a cycle end to end, set `LE_LOGGING_LEVEL=debug` and stop the broker: buffered writes log `stored` from the `forward` component and `le_sparkplug_sf_backlog_size` climbs. Start the broker again and the node logs `draining stored DDATA` with a `count` field as the gauge falls to 0.

## Publish Filters

By default the node publishes every DeviceHub device and every enabled Digital Twin instance it finds. Filters narrow that set, keeping test equipment out of production namespaces. They apply to whole devices and instances, not to individual tags or attributes.

| Environment variable   | `config.json` field     | Default | Effect                                                                          |
| ---------------------- | ----------------------- | ------- | ------------------------------------------------------------------------------- |
| `DH_ENABLED`           | `devicehub.enabled`     | `true`  | Publishes DeviceHub devices. `false` stops DeviceHub discovery entirely         |
| `DH_DEVICES_INCLUDE`   | `devicehub.include`     | empty   | Glob patterns of device names to publish                                        |
| `DH_DEVICES_EXCLUDE`   | `devicehub.exclude`     | empty   | Glob patterns of device names to skip                                           |
| `DT_ENABLED`           | `digital_twins.enabled` | `true`  | Publishes Digital Twin instances. `false` stops Digital Twin discovery entirely |
| `DT_INSTANCES_INCLUDE` | `digital_twins.include` | empty   | Glob patterns of instance names to publish                                      |
| `DT_INSTANCES_EXCLUDE` | `digital_twins.exclude` | empty   | Glob patterns of instance names to skip                                         |

Patterns match the DeviceHub device name or the Digital Twin instance name. As environment variables they are comma-separated; in a configuration file they are JSON lists. They support `*` for any sequence of characters, `?` for a single character, and `[abc]` for a character class. Regular expressions are not supported, and a pattern that is not a valid glob stops startup.

The node applies them in this order:

1. An exclude match drops the device. Exclude always wins over include.
2. An empty include list admits everything that survived the excludes.
3. A non-empty include list admits only names matching one of its patterns.

| Goal                                    | Setting                                  |
| --------------------------------------- | ---------------------------------------- |
| Publish only production lines           | `DH_DEVICES_INCLUDE=line1_*,line2_*`     |
| Publish everything except test rigs     | `DH_DEVICES_EXCLUDE=test_*,*_sandbox`    |
| Publish one named device                | `DH_DEVICES_INCLUDE=press_002`           |
| Publish DeviceHub devices but no twins  | `DT_ENABLED=false`                       |
| Publish twins for one model's instances | `DT_INSTANCES_INCLUDE=hydraulic_press_*` |

A filtered device is absent, not marked offline: the node never publishes a `DBIRTH` for it, so host applications never learn it exists. Buffered data for a device you have just filtered out is discarded rather than replayed, for the reason given under Store and Forward above.

Filters are evaluated on every metadata refresh, so a device that starts matching an exclude pattern disappears from the next birth set — but because configuration is read only at startup, changing a filter takes a restart. `le_sparkplug_devicehub_devices_filtered_total` and `le_sparkplug_digital_twins_instances_filtered_total` count each drop with its reason, and the logs name each filtered device at debug level.

A mismatched \`PRIMARY\_HOST\_ID\` looks exactly like a broken deployment: the node connects to the broker, logs \`waiting for primary host application\`, and never publishes a birth certificate. Check the ID first, then clear any retained \`STATE\` message left from earlier testing by publishing an empty retained message to the same topic.\`LE\_MAX\_SIZE\_MB\` and \`LE\_MAX\_NUM\_ENTRIES\` size the buffer's individual files; they are not a total cap on disk use. The practical bound on the backlog is \`LE\_TTL\_DURATION\` together with the free space on the volume. Size that volume for your worst expected outage, and shorten the TTL if the data loses value quickly.Replay order is not capture order. Messages are keyed by a random identifier, so the backlog drains in arbitrary order. Each message still carries its original timestamps, which is what a historian records.
