SignalK (espos_sk) — discovery, access token, delta stream, inbound¶
M3: find the server, get and keep a token. M4: stream published values as
deltas over a WebSocket, buffer them while offline, reconcile metadata,
publish device health. M7: subscribe to paths and families, receive values
and meta, send PUT requests and raw frames — what a display or controller
needs on top of a sensor. Plus one HTTP client for everything else an
application asks the server over REST (espos_sk_http.h).
Discovery¶
espos_sk browses _signalk-http._tcp via mDNS every sk.discover_s
(default 60 s, immediately when WiFi comes up) and keeps up to
ESPOS_SK_MAX_SERVERS (12) servers with their TXT records (self, roles,
swname, swvers). Entries that
drop out of one query survive two intervals (mDNS is lossy).
GET /api/v1/sk/servers lists them; sk_servers SSE events fire after
every pass.
The responder the queries go through — and the device's own
<hostname>.local, _http._tcp and _espos._tcp records — is
espos_wifi's (wifi.md, "mDNS"); espos_sk only browses. A pass
waits for espos_mdns_is_ready() (the station reports connected a few
milliseconds before the responder's ESPOS_EVENT_MDNS_READY reaches it,
and an empty first pass would only be retried a whole interval later) and
returns nothing without a link. Built with CONFIG_ESPOS_NET_MDNS=n there
is no responder to browse with: discovery is off and sk.server_host must
be set.
Which server is used:
sk config |
choice |
|---|---|
server_host set |
that host:port (manual, for networks without mDNS) |
server_self set |
the discovered server with that self URN |
| neither | sticky: the server our token / pending request belongs to; else the discovered master with the lowest self URN; else any |
Discovered servers that turn out unreachable (wrong subnet, gone) are skipped for five minutes so one dead entry cannot block the machine.
Token state machine¶
sk_token_sm.c is pure C over an injected port (HTTP calls, storage,
timer, clock) and is unit-tested for every transition
(test/host/espos_sk_test). Verified against a real signalk-server 2.31
(the flow, status codes and body shapes below are what it actually
returns).
NO_SERVER ──server known──▶ evaluate:
pending href for this server? ──▶ REQUESTED (resume polling)
stored token for this self? ──▶ VERIFYING ── 200 ▶ APPROVED
else ──▶ IDLE: POST /signalk/v1/access/requests
{clientId, description, permissions}
202 {state:PENDING, href} ▶ REQUESTED, href persisted
404 (security disabled) ▶ OPEN (no token needed; re-POSTed every 60 s —
GET /self is blind to security when allow_readonly is on)
403 (device requests off) ▶ DENIED
400 "already requested" ▶ ERROR, retry in 60 s
unreachable / 5xx ▶ ERROR, backoff 10 s → 5 min
REQUESTED: GET href every 5 s, ×1.5 up to 60 s
state PENDING ▶ keep polling
COMPLETED + APPROVED + token▶ token persisted (keyed by self) ▶ VERIFYING
COMPLETED + DENIED ▶ DENIED (no auto retry; UI offers "request again")
404 / 500 "not found" ▶ server lost it ▶ IDLE (request again)
VERIFYING / APPROVED: GET /signalk/v1/api/self with Bearer
200 ▶ APPROVED (self URN learned/updated), re-check every check_s
401 / 403 ▶ token dropped ▶ IDLE (request again)
APPROVED + any other SK call reporting 401/403 (espos_sk_report_unauthorized) ▶ IDLE
Design points from the plan, all implemented:
clientIdis a v4 UUID generated once and stored in theskstateNVS namespace, never in the exported configuration; it survives config import/export and factory-reset only wipes it because the whole partition goes.- Tokens are keyed by the server's
selfURN. A server that changes address keeps its token (discovery re-resolves the host by self); a reinstalled server (new self) gets a fresh request; a token learned for a manual host without mDNS has its self filled in from the first successful verify. - A pending
hrefis persisted with the server it belongs to; a reboot mid-approval resumes polling instead of creating a duplicate request. - Manual token paste:
POST /api/v1/sk/token {"token": "…"}→ verified immediately. - Secrets: the token never appears in any API response or SSE event; the store lives in the same (optionally encrypted) NVS partition as the config.
Delta stream (M4)¶
espos_sk_publish_number/string/bool/json(path, value) is the whole app
API: thread-safe, never blocks, works before WiFi is up (but only after espos_sk_start() — earlier calls return ESP_ERR_INVALID_STATE). Values are for
vessels.self; the source label is espos.<hostname>.
Pipeline (sk_delta.c, pure C, unit-tested; sk_ws.c = the transport
task):
- Batching window (
sk.batch_ms, default 100 ms): everything published inside one window becomes one delta message with one update; a path published twice in a window keeps the last value. Numbers use the shortest round-trip representation. - Ring buffer while the stream is down (
sk.buffer_msgs/sk.buffer_kb, default 128 messages / 32 KiB): oldest messages are dropped first and counted (ws.dropped). Windows keep closing while offline, so a path's history survives, not just its latest value. - Drain after (re)connect at
sk.drain_per_s(default 20/s) so the server is not swamped by a backlog; new values queue behind the backlog so ordering per path is preserved.
The WebSocket task (espos_skws) runs when sk.ws_enabled, WiFi is up, a
server is selected and the token state allows streaming (approved, or the
server has security off). It connects to
ws://<host>:<port>/signalk/v1/stream?subscribe=none with
Authorization: Bearer <token>, consumes the hello, then sends deltas as
text frames. A 401 on connect calls espos_sk_report_unauthorized() (the
token machine re-verifies / re-requests); any other failure backs off with
the shared WiFi backoff curve (ws.next_retry_s). Config changes to the
stream keys are picked up live; ws_enabled=false closes the socket and
keeps buffering.
Meta reconciliation. espos_sk_declare_meta(path, meta_json,
period_ms) records metadata for a NON-standard path (spec paths belong to
the server). On every (re)connect the task GETs
/signalk/v1/api/vessels/self/<path>/meta; if the server has none it
PUTs ours, otherwise the server's copy — possibly edited by the user —
wins. period_ms > 0 adds timeout (2.5× the period, in seconds), the one
field the device really owns. ws.meta.declared/reconciled show progress.
Device health. Every sk.health_s (default 10 s, 0 = off) the task
publishes espos.<hostname>.{uptime,freeHeap,minFreeHeap,internalFree,largestBlock,rssi,
wifiReconnects,skReconnects,resetReason} with declared meta, so a
dashboard sees the device without any app code.
Wire facts that cost time (signalk-server 2.31): client text frames must
be sent with the FIN bit (WS_TRANSPORT_OPCODES_TEXT |
WS_TRANSPORT_OPCODES_FIN) or the server closes the socket after the
first frame; subscribe=none still delivers the hello; meta GET is
404 when unset and PUT takes {"value": {…}}.
Inbound (M7)¶
int h = espos_sk_subscribe("navigation.*", 1000, on_update, NULL); /* family */
espos_sk_subscribe("environment.mode", 0, on_update, NULL); /* exact */
espos_sk_unsubscribe(h);
espos_sk_put("navigation.anchor.maxRadius", "30", on_put_done, NULL);
espos_sk_send_raw("{\"context\":\"vessels.self\",\"updates\":[…]}");
- Subscriptions are exact paths or families (
prefix.*,prefix*,*). The stream is opened withsubscribe=none&sendMeta=all; after the hello (and after every reconnect) one{"context":"vessels.self", "subscribe":[{"path","period","format":"delta","policy":"instant", "minPeriod"}]}frame carries every subscription; new ones while connected go out incrementally,espos_sk_unsubscribesendsunsubscribewhen nothing else covers the pattern. Up toESPOS_SK_MAX_SUBS(48). - Delivery:
sk_parse.c(pure C, cJSON) turns each frame into items —path,value_json(verbatim JSON text: numbers, strings, objects,null),timestamp,$source/source.label,context— plus meta items (meta_jsonset, value NULL) when the server sendsmeta. The callback runs on the stream task; copy what you need and return (a display marshals to its UI thread — never block, never call anespos_sk_*function that could wait on the stream). Frames are reassembled up toCONFIG_ESPOS_SK_RX_FRAME_MAX(16 KiB); larger ones are dropped with a log line. - PUT (outbound):
{"context":"vessels.self","requestId":<uuid4>,"put": {"path","value"}}; the response (stateCOMPLETED/FAILED,statusCode,message) is matched by requestId and handed to the callback; no answer in 10 s →"TIMEOUT". Up to 8 in flight;ESP_ERR_INVALID_STATEwhen the stream is down (nothing is queued across reconnects — a control action must not fire minutes later). A real server without a handler answersCOMPLETEDwithstatusCode 405 "PUT not supported for …". - Raw frames (
espos_sk_send_raw) go out ahead of buffered deltas — e.g. an inboundnotifications.*delta withstate:"normal"to acknowledge an alarm. - Status:
ws.in {subs, frames, received},ws.put {pending, ok, failed}; RESTPOST /api/v1/sk/put {"path","value"}(202; last answer underGET /api/v1/sk/put) for scripts. - The example app subscribes to
app.watch_pathand logs each update. - PUT (inbound) — the server operating this device — is a separate mechanism with its own handler table; see Inbound PUT (control).
Verified 2026-08-18 against signalk-server 2.31 on the ESP32-P4: 586
updates in ~40 s of navigation.* from N2K sources, satellitesInView
objects of several KiB reassembled, PUT round trip (405 from a server
without handlers).
Inbound PUT (control)¶
Everything above is this device asking the server for something. This section is the other direction: the server asking this device to change something — a switch operated from a phone, a setpoint moved from a plotter. Without it a device can only ever report.
static esp_err_t set_bilge(const char *path, const char *value_json, void *arg)
{
if (strcmp(value_json, "true") == 0) { gpio_set_level(RELAY, 1); return ESP_OK; }
if (strcmp(value_json, "false") == 0) { gpio_set_level(RELAY, 0); return ESP_OK; }
return ESP_ERR_INVALID_ARG; /* answered COMPLETED 400 */
}
espos_sk_put_handler_register("electrical.switches.bilge.state", set_bilge, NULL);
espos_sk_publish_bool("electrical.switches.bilge.state", false); /* REQUIRED, see below */
The handler's return value becomes the answer:
| return | answer |
|---|---|
ESP_OK |
COMPLETED 200 |
ESP_ERR_INVALID_ARG |
COMPLETED 400 |
| anything else | COMPLETED 502 |
ESPOS_SK_PUT_PENDING |
PENDING 202, and you call espos_sk_put_respond() later |
| no handler for the path | COMPLETED 405 |
Three things that are easy to get wrong¶
The device must publish the path first. signalk-server routes a PUT to a
device by the (path, $source) pairs it has seen that connection publish
(processUpdates in src/interfaces/ws.ts). A path this device has never
published does not exist as a PUT target, and the request is answered 405 by
the server without ever reaching the device. Publish the current state once at
boot, and again on every change.
put arrives as an ARRAY. The server writes
{"requestId","context","put":[{"path","value"}]} — an array, even for one
path. That is not the shape a client sends outbound (espos_sk_put writes a
single object), and a parser that only understands the object form silently
sees no requests at all. espOS accepts both.
The reply state must be COMPLETED or PENDING. signalk-server's
isWsRequestReply() accepts a state of exactly COMPLETED, PENDING or
null and silently ignores anything else. A reply with
"state":"FAILED" — the obvious spelling for a failure — is dropped without
a word, and the client then waits out the server's full 60-second timeout. A
failure is COMPLETED with a 4xx/5xx statusCode.
espOS answers every request, including one whose path has no handler and one whose items are unusable. Silence is the worst answer: it costs the client 60 seconds and tells it nothing.
Flushing before sleep¶
Deltas are batched (sk.batch_ms), so a value published a millisecond before
esp_deep_sleep_start() is still sitting in the buffer when the radio goes
down. espos_sk_flush() closes the batch and waits — ESP_OK when nothing is
left, ESP_ERR_TIMEOUT if the deadline passed, ESP_ERR_INVALID_STATE when
the stream is down and nothing can drain. Never call it from the stream task
or a subscription callback.
The message the stream task is writing at that moment counts as pending until
the write returns, so ESP_OK means everything was handed to the socket. It
does not mean the server has acknowledged it: on a live link the TCP stack
sends within milliseconds, and code about to cut the radio should allow it
that moment.
The graph nodes (espos_sk_flow)¶
espos_sk_flow is the Signal K end of the data-flow graph — the
same calls as above, as nodes.
| Node | Direction |
|---|---|
sk::Output<T>(path[, Meta]) |
publish |
sk::Listener<T>(path) |
receive a value from the server |
sk::PutHandler<T>(path) |
let the server change something here |
sk::PutRequest<T>(path) |
ask the server to change something |
sk::Notify(key, message) |
raise/clear a device condition |
sk::NetRssi, sk::IpAddress |
what the network says about itself |
Metadata is impossible to get wrong¶
Never send metadata for a path in the Signal K specification. The server
already knows that navigation.speedOverGround is metres per second; a device
that declares it anyway can only get it wrong, and then every dashboard on the
boat is wrong.
So Output has no units argument. A units string cannot be passed without
constructing a Meta, and constructing a Meta is the statement "this path
is mine, nobody else knows what it means" — which is exactly when metadata is
correct:
sk::Output<float> sog("navigation.speedOverGround"); // spec: no meta, ever
sk::Output<float> pv("sensors.solar.0.voltage", sk::Meta{"V"}); // ours: meta declared
The rule is not documented and hoped for; it is unspeakable.
Output<std::optional<T>> publishes JSON null when disengaged — "this
sensor has nothing right now", which is different from zero and different
from stale.
Receiving¶
Listener and PutHandler receive on the stream task and neither emits
there: both post into a Mailbox, so the emit happens on the flow task like
every other node. That is why you wire from .out():
auto& depth = g.make<sk::Listener<float>>("environment.depth.belowTransducer");
depth.out() >> shallow_alarm;
A switch a phone can operate¶
auto& req = g.make<sk::PutHandler<bool>>("electrical.switches.bilge.state");
auto& relay = g.make<espos::sensors::GpioOutput>("relay", 22);
auto& state = g.make<sk::Output<bool>>("electrical.switches.bilge.state");
req.out() >> relay >> state; // PUT -> pin -> publish what the pin did
Publishing at the end of the chain is what registers this device as the path's source, without which the server has nowhere to route the PUT. It also publishes what the pin actually did rather than what was asked for — they differ when the pin failed to open, which is exactly when a dashboard must not lie.
PutHandler answers 200 as soon as the value is accepted into the mailbox,
not once the chain has run: the alternative is to block the stream task until
the flow task finishes, and one slow consumer would then stall every other
frame on the connection.
HTTP requests to the server¶
espos_sk_http.h is the one way an application talks HTTP to the selected
server. Four hand-rolled copies of "GET a SignalK REST node" in one firmware
had two of them rebooting the device; this is the version that does not.
#include "espos_sk_http.h"
espos_sk_http_resp_t r;
if (espos_sk_http_get("/signalk/v1/applicationData/global/my-app/1/layout.json", NULL, &r) == ESP_OK
&& r.status == 200 && !r.truncated) {
apply_layout(r.body, r.len); /* NUL-terminated, malloc'ed */
}
espos_sk_http_resp_free(&r);
char *value = NULL; /* GET …/vessels/self/navigation/position → "value" member */
if (espos_sk_get_value("navigation.position", &value) == ESP_OK) { /* {"latitude":…,"longitude":…} */ }
free(value);
char *meta = NULL; /* GET …/navigation/speedOverGround/meta */
if (espos_sk_get_meta("navigation.speedOverGround", &meta) == ESP_OK) { /* {"units":"m/s",…} */ }
free(meta);
espos_sk_http_opts_t o = { .timeout_ms = 3000, .max_body = 512 };
espos_sk_http_post("/plugins/my-plugin/api/thing", "{\"on\":true}", &o, &r); /* PUT, DELETE likewise */
espos_sk_http_resp_free(&r);
char url[ESPOS_SK_URL_MAX];
espos_sk_url("/signalk/v1/api", url, sizeof(url)); /* http(s)://host:port/signalk/v1/api */
espos_sk_ws_url("/signalk/v1/stream", url, sizeof(url)); /* ws(s)://… */
- Reply:
status(0 when nothing arrived),body(malloc'ed, NUL-terminated,""for an empty reply),len,truncated. The call returnsESP_OKwhenever a reply arrived — a 404 or 500 is a successful call; checkstatus.ESP_ERR_INVALID_STATEmeans no server is selected,ESP_ERR_TIMEOUTthat no connection slot came free, anything else is the transport erroresp_http_clientreported. Alwaysespos_sk_http_resp_free(). - Options (
NULLor zeroed = defaults): 6 s timeout, 16 KiB body cap,Authorization: Bearerfrom the current token, 401/403 reported,Accept: application/json.no_authdrops the header,no_report_unauthorizedkeeps a 401 from touching the token machine,acceptoverrides the header (""= none). - Body cap: the body is collected in
HTTP_EVENT_ON_DATAand stops atmax_body; beyond ittruncatedis set and the rest is drained and discarded. Never parse a truncated body — treat it as "the reply was too big" (espos_sk_get_value/metareturnESP_ERR_INVALID_SIZE). The buffer grows with the reply, so the cap costs nothing for small documents. - Token: snapshotted per call from
espos_sk_get_token(), sent as a header — never in a query string, where it would end up in every proxy and server log. No token, no header, which is what a server running without security expects. A 401/403 while a token was sent callsespos_sk_report_unauthorized(): the token machine re-verifies and, if the server really has dropped the device, requests access again. - Scheme: from the selected server's
tlsflag,http/wsorhttps/wss— decided per server bysk.schemeand, underauto, by what the server advertised or a probe found (below). Certificates go through the same trust store as the delta stream.espos_sk_url()/espos_sk_ws_url()build the URL for code that opens its own connection (the BLE gateway's control socket does). - Concurrency:
CONFIG_ESPOS_SK_HTTP_MAX_CONCURRENT(default 2, range 1–8) bounds requests in flight — the token machine's and the meta reconciliation's own calls included. Each open request is a socket plus, over TLS, ~20 KB of RAM; a display fetching one value per widget on a layout change would otherwise open dozens at once. A caller over the limit waits up to its owntimeout_msfor a slot. - Threading: blocking, on the caller's task, for up to
timeout_mswaiting for a slot plustimeout_mson the wire, with ~2 KiB of its stack. Call from an application task. Never from the SK stream task (theespos_sk_subscribe/espos_sk_putcallbacks), anESPOS_EVENThandler, a Bluetooth stack callback or an HTTP URI handler.
The two crash patterns it avoids¶
Both were reproduced on the ESP32-P4 against signalk-server, both are inside
esp_http_client, and both are why the helper insists on one particular
shape — a fresh client per call and esp_http_client_perform() only:
esp_http_client_open()→fetch_headers()→read()leaves the client'scache_data_in_fetch_hdrflag set. When the body arrives in the same TCP segment as the headers — which is every small SignalK reply — the next step hitsassert(orig_raw_data == raw_data)inhttp_on_bodyand the device reboots. Onlyperform()clears the flag.- Reusing one handle across
perform()calls (set_url()per path, one connection for a batch) desyncs the same two pointers and trips the same assert mid-batch (seen onnavigation.anchor.*paths).
perform() with a new handle per request enters neither path; the body is
delivered through the event handler, which is also where the size cap lives.
sk_http.c has used this shape for the token legs since M3; the meta
reconciliation and the public API now share that single implementation.
TLS (https / wss)¶
On by default (CONFIG_ESPOS_SK_TLS=y) and, with sk.scheme at its
default auto, used whenever the server says it speaks it. What changed: the
device now has somewhere to put a certificate no public root signed, which is
what every boat server has.
Which scheme (sk.scheme)¶
auto | http | https, default auto. http and https force one. What
auto does depends on how the server was found:
- discovered — signalk-server advertises
_signalk-https._tcpinstead of_signalk-http._tcpwhen itssslsetting is on (src/interfaces/rest.js), so the server has already told us. The device browses both types; no probe, no round trip. - manual host — nothing has been advertised, so one probe:
GET http://host:port/signalkwith redirects switched off. A30xto anhttps://Location means TLS, and the port is taken from the Location if it names one (an SSL-enabled server often listens elsewhere). Nothing answering on the plain port is tried once over https before concluding the host is down. The probe carries no token — it is aimed at an address that has not been established as our server yet, and handing the credential to whatever answers is exactly what must not happen.
No answer on either scheme is not an answer. The host may be down, still booting, or not reachable yet, and on a boat the server usually comes up after its devices. The device uses plain http as a guess, does not remember it, and asks again when the network comes up and after every attempt that runs on the guess. For the same reason a manual host is not selected, or probed, before there is a network at all.
An answer is a field on the chosen server (espos_sk_server_t::tls), cached
per (host, port) and re-decided only when the selection changes. It is not
restart_required any more: sk_ws.c destroys and rebuilds its transport pair
when the scheme changes, so a server that gains or loses TLS is followed live.
GET /sk/status reports what is actually in use as server.scheme, which
under auto is the only place to read it.
Which certificates are trusted (sk.tls_trust)¶
tofu | ca | bundle, default tofu. There is no accept-anything mode.
tofu— trust on first use, the way ssh does it. The first connection that works pins what the server presented; every later one must match.
Two shapes of anchor, because certificates get renewed:
- CA anchor — the highest
CA:TRUEcertificate in the chain, plus the normalised set of the leaf's dNSName/IP SANs. A renewal signed by the same CA for the same names is accepted with nobody pressing anything, which is what makes a 90-day certificate survivable on a device in a locker. Binding the SAN set as well as the CA matters: a private CA that signs one host would otherwise vouch for every other name it ever signs. -
Leaf anchor — the SHA-256 of the leaf itself, used when the chain has no CA or the leaf carries no SAN. signalk-server's own generated self-signed certificate lands here, and a renewal then needs one deliberate "trust the new certificate".
-
ca— the same machinery with the anchor supplied instead of captured: put the issuing CA insk.ca_pem(a blob; base64 over the REST API) orPUT /api/v1/sk/tls/ca {"pem": "-----BEGIN CERTIFICATE-----…"}. Setting it anchors the device immediately, so a fleet that pre-seeds the key connects on the first try with nothing to capture. Invalid PEM is a400, not a device that quietly stops connecting. -
bundle— the public Mozilla roots and nothing pinned: the pre-S1 behaviour, for a server with a real certificate from a real CA.
The common-name check is switched off outside bundle mode. It would add
nothing — the certificate has already been matched by fingerprint and SAN set —
and it fails on a boat server reached by IP, which is most of them.
Stash, then commit¶
A verify callback runs during a handshake that has not finished. Pinning
there would let a machine-in-the-middle answering a first connect plant its own
certificate as the anchor and be trusted from then on. So the callback only
fills a capture slot in RAM, and espos_sk_tls_commit() writes it to NVS after
the connection has proved itself — an HTTP status from the far end, or a
WebSocket 101. Every handshake starts by discarding whatever the last one
left.
The anchor lives in the skstate NVS namespace next to the token
(tls_kind, tls_fp, tls_ca, tls_san, tls_cn, tls_self, tls_at):
device state, so it does not travel with a configuration export. sk.ca_pem
is configuration and does.
When it does not match: cert_error¶
The token machine gains a state. token.state == "cert_error" means the
transport was refused, not the credential:
- the token is kept. Dropping it would mean a fresh approval in the server's admin UI after every renewal.
- retry is a flat 60 s, not the exponential ladder an unreachable server gets. The fix arrives from outside — the server renews, or somebody presses the button — and should be noticed within a minute rather than an hour.
- the stream stays down (
espos_sk_stream_allowed()is false). Falling back to plaintext would send the token to whoever answered. espos_healthcarries it asskCertificate(ALARM, non-fatal — a reboot would not fetch a new certificate), which also publishes it as a SignalK notification.
GET /api/v1/sk/tls shows the pinned identity next to the presented one, which
is what makes it actionable: an operator who can compare the two fingerprints
knows whether this is their own renewal. DELETE /api/v1/sk/tls forgets the
anchor and retries at once — the SignalK page's "Trust the new certificate"
button.
The plaintext 401 rule¶
Over TLS nothing on the path can inject an answer, so a 401 on a request that
carried our token is the server's and clears it, as before.
Over plaintext it is not. A captive portal, a proxy, a router's own login
page — all of them answer 401, and throwing the token away over one costs a
trip to the server's admin UI to approve the device again. So the first
unauthorised answer on a plaintext connection buys a second opinion instead:
keep the token, ask again in 5 s, and only clear if that is refused too. Any
200 resets the count. (SensESP's should_clear_token_on_status, same
reasoning.)
Heap¶
A TLS handshake wants around 20 KB of contiguous internal RAM, and that is the pool an ESP32 with a WiFi stack is short of — total free heap is the wrong number on a PSRAM board, where tens of megabytes hide the few kilobytes that matter.
- One handshake at a time, device-wide. Two at once is where the pool runs
out; it also gives the esp-tls attach hook, which takes no user pointer, a
well-defined capture slot. A caller waits
CONFIG_ESPOS_SK_TLS_HANDSHAKE_TIMEOUT_MS(10 s) for it, then retries on its own backoff. - Pre-flight. Below
CONFIG_ESPOS_SK_TLS_MIN_FREE_BLOCK_KB(24 KB) of largest free internal block the handshake is deferred andtlsMemoryis raised as a WARN. What that avoids is not a clean out-of-memory error: mbedTLS failing mid-handshake leaves a half-built session and a socket behind, and on a reconnect loop it starves the WiFi stack of the same pool. - The verify leg is skipped on TLS servers. The WebSocket upgrade carries the same token and rejects it just as plainly, so the leg would only buy a second handshake per reconnect.
sdkconfig.d/espos.defaultsturns on mbedTLS's dynamic buffers (MBEDTLS_DYNAMIC_BUFFER,DYNAMIC_FREE_PEER_CERT,DYNAMIC_FREE_CONFIG_DATA) and drops the server side (MBEDTLS_TLS_CLIENT_ONLY);SSL_IN_CONTENT_LENstays at 16384, which a SignalK subscription burst needs.
Cost¶
Flash, measured on esp32c6 as CONFIG_ESPOS_SK_TLS=y minus =n on the same
tree:
| build shape | y | n | cost |
|---|---|---|---|
with espos_ota (the reference app) |
1 485 308 | 1 475 280 | ~10 KB |
without it (the tls_server example) |
1 437 772 | 1 358 004 | ~78 KB |
A firmware that already has espos_ota links mbedTLS and the certificate
bundle for the https image source, so all it pays here is the trust store, the
transports and the verify path. Without it, most of the 78 KB is the Mozilla
bundle — which the trust store still needs, as the fallback for bundle mode.
RAM: ~20 KB while a connection is open. Note that the mbedTLS dynamic-buffer
settings above cost about 17 KB of flash in both columns, because they are in
sdkconfig.d/espos.defaults and apply to the OTA client too; they are what
gives the RAM back between handshakes.
Testing it against a real server¶
Turn SSL on in signalk-server (Server → Settings → SSL, or "ssl": true in
settings.json) and restart it. It generates a self-signed certificate and
advertises _signalk-https._tcp.
- The device's SignalK page should show the server with scheme
httpsand, after the first connect, a pinned certificate (leaf anchor — a generated self-signed certificate has no CA and often no SAN). - Delete signalk-server's certificate and restart it so it generates a new
one. The device goes to
cert_errorwithin a minute, keeps its token, and the page shows the pinned fingerprint next to the presented one. - Press Trust the new certificate. It reconnects within a few seconds.
- For the CA path, issue the server's certificate from a CA of your own and
put that CA in
sk.ca_pem; re-issuing the leaf for the same names should not interrupt anything, and re-issuing it for a different name should givecert_errorwith "names different hosts".
Task stacks¶
Two tasks, two Kconfig options. Both are at their historical sizes; what changed is that they are now options with measured numbers behind them rather than literals in the source.
CONFIG_ESPOS_SK_TASK_STACK(default 12288, range 6144–16384) — the client task: discovery, the token state machine and its HTTP legs (access request, poll, verify), plus the https probe.CONFIG_ESPOS_SK_WS_TASK_STACK(default 8192, range 6144–16384) — the stream task: frame reassembly, delta parsing, inbound PUT dispatch, the notification sink — and metadata reconciliation, which issues an HTTP GET and PUT per path (reconcile_meta()is called fromws_task(), not from the client task).
Measured on an ESP32-C5 running the BLE gateway against a live server —
discovery, an approved token, meta for the published paths — with
CONFIG_FREERTOS_USE_TRACE_FACILITY=y and uxTaskGetStackHighWaterMark():
| task | allocated | peak use | never touched |
|---|---|---|---|
espos_sk |
12288 | 1488 | 10800 (88 %) |
espos_skws |
8192 | 5272 | 2920 (36 %) |
Neither default was lowered on the strength of that, and the reason is worth stating. The server in that run was plain HTTP on port 80, so no TLS handshake completed on either task — and mbedTLS is the deepest call path both of them have. A 12 % high-water reading taken with the deepest path never taken does not show the stack is oversized; it shows the measurement was incomplete.
So the numbers above are a floor, not a budget, and what to measure before lowering either default depends on the build:
CONFIG_ESPOS_SK_TLS=y(the default) — take the reading against awssserver, where the client task's token legs and the stream task'sespos_sk_http_get_meta()/put_meta()calls actually go through mbedTLS.CONFIG_ESPOS_SK_TLS=n—sk_tls.cis compiled out and the device can only speak http/ws, so no wss reading is possible. Measure the busiest workload that build does support instead: a reconnect that replays a full subscription burst and reconciles meta for every published path.
The floors are 6144 on both. 4096 would sit below the 5272 B the stream task was measured using, and an option whose range lets a build fault on the first large frame is a trap rather than a choice. A stack-protection fault here is a reboot loop, not a degraded mode.
API¶
GET /api/v1/sk/status— token/server/discovery status plus thewsstream object (see rest-api.md).POST /api/v1/sk/put {"path","value"}/GET /api/v1/sk/put— PUT over the stream and the last answer.POST /api/v1/sk/publish {"path","value"[,"meta","period_ms"]}— publish over HTTP.- SSE
sk_ws— thewsobject on every stream change. GET /api/v1/sk/servers— discovered servers,selectedflag.POST /api/v1/sk/discover— run a discovery pass now.POST /api/v1/sk/request— request again (from denied/error/open).POST /api/v1/sk/token {"token"}— manual token.POST /api/v1/sk/forget— drop the token and start over. A request that is still pending is kept and polled on (the server holds it anyway and refuses duplicates).GET /api/v1/sk/tls— the pinned certificate identity next to the one the server last presented, and why they did not match if they did not.DELETE /api/v1/sk/tls— forget the pinned certificate and retry now.PUT /api/v1/sk/tls/ca {"pem"}— supply the issuing CA (switchessk.tls_trusttoca).- SSE
sk_tls— the/sk/tlsdocument, on connect and whenever the anchor changes. - SSE
sk,sk_servers.
Testing¶
test/host/espos_sk_test: 32 Unity cases — token state machine, store, delta batching / ring / drain, frame parser and path patterns.test/host/espos_httpd_testSkTests: the real HTTP client and WebSocket against a Python mock of the signalk-server security API and stream endpoint (approve, deny, revoke, forget, security off, manual token, manual host, deltas + meta reconciliation, offline buffering with ordered drain; the HTTP helper through a harness probe: 200 with body and Bearer header, 404 as a reply, oversize body →truncated, value/meta/URL lookups, PUT through the same core, 401 → token machine re-requests);SkInboundTests: subscribe frames, value/meta delivery, exact vs family dispatch, PUT round trip / failure / timeout, raw frames, unsubscribe + resubscribe after reconnect, 9 KiB frame reassembly.- Against a real signalk-server on the host:
node bin/signalk-server -c <fresh config dir>from a checkout,POST /skServer/enableSecurity {"userId","password","type":"admin"}, restart, then approve withPUT /skServer/security/access/requests/<clientId>/approvedusing the admin cookie fromPOST /signalk/v1/auth/login; revoke withDELETE /skServer/security/devices/<clientId>.
Notifications¶
espos_sk_notify(key, state, message) raises or clears a SignalK notification
under notifications.espos.<label>.<key>:
espos_sk_notify("wakeService", ESPOS_SK_ALERT_WARN, "wake service unreachable");
espos_sk_notify("wakeService", ESPOS_SK_ALERT_NORMAL, ""); /* cleared */
For conditions the device knows about and an operator would want to see: memory pressure, an overheating chip, a service the firmware depends on having gone away. Without them these surface as a device that has quietly stopped doing its job, which looks identical to a hardware fault and is the expensive kind of problem to diagnose.
- Level-triggered and idempotent. Re-raising the same state and message
sends nothing, so a caller may poll and re-raise on every tick. The first
raise after boot always goes out, even if it is
NORMAL, because the server may still hold an alert from before a restart. keyis a short stable identifier (lowMemory,wakeService) -- it becomes part of the path, and the path is what a rule or dashboard keys on. Themessageis the human half and may change freely.methodis["visual"]for warn/alarm and[]on clear. What to do about it is the server's decision, not the device's.- Up to
CONFIG_ESPOS_HEALTH_MAX_CONDITIONSdistinct keys (default 8, range 1-32): the notification is a sink onespos_health, so its condition table is the cap and one key too many getsESP_ERR_NO_MEM. Deltas are buffered like any other while offline. Oversized keys or messages are rejected withESP_ERR_INVALID_SIZErather than truncated -- a clipped key would never match on the next call and would leak a slot.
espOS raises lowMemory itself, from espos_health's watchdog policy rather
than the SignalK tick, so it exists without SignalK; thresholds and the
restart rule are in health.md. Internal RAM is checked separately
because it is the scarce pool on targets with PSRAM -- tens of megabytes free
overall can hide an internal-RAM exhaustion that will take the radio down.