Skip to content

File espos_health.h

FileList > espos_health > include > espos_health.h

Go to the source code of this file

  • #include <stdbool.h>
  • #include <stddef.h>
  • #include <stdint.h>
  • #include "esp_err.h"

Classes

Type Name
struct espos_health_condition_t
struct espos_health_reset_record_t

Public Types

Type Name
typedef void(* espos_health_sink_t
enum espos_health_state_t

Public Functions

Type Name
esp_err_t espos_health_add_sink (espos_health_sink_t sink, void * arg)
Register a sink.
bool espos_health_fatal_alarm (espos_health_condition_t * out)
The condition, if any, the policy would restart for: an ALARM raised with ESPOS_HEALTH_F_REBOOT_ON_ALARM.
void espos_health_kick (void)
The calling task is alive: feed the task watchdog and stamp the policy's registry.
bool espos_health_last_reset (espos_health_reset_record_t * out)
The record the previous boot left when the policy restarted the device.
esp_err_t espos_health_policy_start (void)
Arm the device watchdog: a periodic tick (CONFIG_ESPOS_HEALTH_POLICY_TICK_S, 10 s) that reports lowMemory and taskStalled, counts consecutive ticks on which a fatal ALARM is held and, at CONFIG_ESPOS_HEALTH_POLICY_STRIKES of them, writes the reset record and restarts.
esp_err_t espos_health_remove_sink (espos_health_sink_t sink, void * arg)
Remove a sink.
esp_err_t espos_health_report (const char * key, espos_health_state_t state, const char * message)
Raise (or with ESPOS_HEALTH_NORMAL, clear) the condition key .
esp_err_t espos_health_report_ex (const char * key, espos_health_state_t state, const char * message, uint32_t flags)
espos_health_report() with flags (ESPOS_HEALTH_F_*).
esp_err_t espos_health_report_test (const char * key, espos_health_state_t state, const char * message, uint32_t ttl_ms)
Raise or clear a synthetic condition, for exercising the sinks a real fault would reach (buzzer, LED, SignalK notifications) without causing the fault.
void espos_health_reset (void)
Forget every condition and sink (tests).
size_t espos_health_snapshot (espos_health_condition_t * out, size_t max)
Copy the current conditions into out (at mostmax ).
const char * espos_health_state_str (espos_health_state_t state)
Human-readable state, for logs and sinks: "normal", "warn", "alarm".
bool espos_health_test_expire (void)
Clear the synthetic condition if its ttl has elapsed; otherwise do nothing.
esp_err_t espos_health_unwatch_task (void)
Stop watching the calling task (both the task watchdog and the policy).
esp_err_t espos_health_watch_task (const char * name, uint32_t timeout_ms)
Watch the calling task: subscribe it to the IDF task watchdog (a task that stops calling espos_health_kick() for CONFIG_ESP_TASK_WDT_TIMEOUT_S panics, which core-dumps, restarts and — before an OTA image is confirmed — rolls back) and register it with the policy, which raisestaskStalled as a fatal ALARM once the task has been silent fortimeout_ms .
espos_health_state_t espos_health_worst (void)
Worst state currently recorded — what a single status LED wants to know.

Macros

Type Name
define ESPOS_HEALTH_F_REBOOT_ON_ALARM (1u &lt;&lt; 0)
define ESPOS_HEALTH_KEY_MAX 24
define ESPOS_HEALTH_MSG_MAX 96
define ESPOS_HEALTH_TEST_PREFIX "test."
Keys a synthetic condition may use must start with this.
define ESPOS_HEALTH_TEST_TTL_MAX_MS (300u \* 1000u)
Longest ttl_ms accepted: a drill nobody clears must end on its own.

Public Types Documentation

typedef espos_health_sink_t

typedef void(* espos_health_sink_t) (const char *key, espos_health_state_t state, const char *message, void *arg);

enum espos_health_state_t

enum espos_health_state_t {
    ESPOS_HEALTH_NORMAL = 0,
    ESPOS_HEALTH_WARN = 1,
    ESPOS_HEALTH_ALARM = 2
};

Public Functions Documentation

function espos_health_add_sink

Register a sink.

esp_err_t espos_health_add_sink (
    espos_health_sink_t sink,
    void * arg
) 

Every condition recorded so far is replayed into it before this returns, so a sink that comes up late (espos_sk connects long after the first report) still learns the current state instead of waiting for the next change.

Returns:

ESP_ERR_NO_MEM when CONFIG_ESPOS_HEALTH_MAX_SINKS is exhausted, ESP_ERR_INVALID_STATE if (sink, arg) is already registered.


function espos_health_fatal_alarm

The condition, if any, the policy would restart for: an ALARM raised with ESPOS_HEALTH_F_REBOOT_ON_ALARM.

bool espos_health_fatal_alarm (
    espos_health_condition_t * out
) 

What a display wants to show while the strikes are still counting. out may be NULL.


function espos_health_kick

The calling task is alive: feed the task watchdog and stamp the policy's registry.

void espos_health_kick (
    void
) 

Cheap and lock-free; call it from the loop you want watched. A task that is not watched may call it harmlessly.


function espos_health_last_reset

The record the previous boot left when the policy restarted the device.

bool espos_health_last_reset (
    espos_health_reset_record_t * out
) 

Valid for the whole of this boot (every caller gets it; /system/info shows it as last_reset) and gone after the next restart, whatever its cause: the first call of a boot clears the stored copy. False when the last reset was not the policy's — power-on, a panic, an OTA reboot, a user's reboot.


function espos_health_policy_start

Arm the device watchdog: a periodic tick (CONFIG_ESPOS_HEALTH_POLICY_TICK_S, 10 s) that reports lowMemory and taskStalled, counts consecutive ticks on which a fatal ALARM is held and, at CONFIG_ESPOS_HEALTH_POLICY_STRIKES of them, writes the reset record and restarts.

esp_err_t espos_health_policy_start (
    void
) 

Idempotent. espos_start() calls this when espos_start_opts_t.health_watchdog is set (the default); call it yourself only when bringing espOS up by hand.

Runs on the esp_timer task: ticks are short, and the sinks a report reaches from it are held to the same rule as everywhere else (do not block).


function espos_health_remove_sink

Remove a sink.

esp_err_t espos_health_remove_sink (
    espos_health_sink_t sink,
    void * arg
) 

A call already in flight on another task may still complete after this returns.


function espos_health_report

Raise (or with ESPOS_HEALTH_NORMAL, clear) the condition key .

esp_err_t espos_health_report (
    const char * key,
    espos_health_state_t state,
    const char * message
) 

Sinks are called only when the state or the message actually changed. message may be NULL or "".

Returns:

ESP_OK also when nothing changed; ESP_ERR_INVALID_ARG for an empty key; ESP_ERR_INVALID_SIZE when key/message exceed the maxima above (rejected rather than truncated — a clipped key would never match on the next call, so every report would consume another slot); ESP_ERR_NO_MEM when CONFIG_ESPOS_HEALTH_MAX_CONDITIONS is exhausted.


function espos_health_report_ex

espos_health_report() with flags (ESPOS_HEALTH_F_*).

esp_err_t espos_health_report_ex (
    const char * key,
    espos_health_state_t state,
    const char * message,
    uint32_t flags
) 

The flags of a condition are those of its latest report; a report that changes only the flags is recorded but not fanned out — sinks see states and messages. Unknown flag bits are ESP_ERR_INVALID_ARG.


function espos_health_report_test

Raise or clear a synthetic condition, for exercising the sinks a real fault would reach (buzzer, LED, SignalK notifications) without causing the fault.

esp_err_t espos_health_report_test (
    const char * key,
    espos_health_state_t state,
    const char * message,
    uint32_t ttl_ms
) 

Reported through the same path as anything real, so a sink cannot tell the difference that is the point with two deliberate restrictions:

  • flags are always 0, so this can never arm the reboot path no matter what a real condition of the same name would carry. Structural, not a promise.
  • key must start with ESPOS_HEALTH_TEST_PREFIX, else ESP_ERR_INVALID_ARG.

At most ONE synthetic condition is active at a time: raising a second clears the first. A drill on a live boat should have a bounded blast radius, and a test script cannot leak fake alarms into the table by looping.

Parameters:

  • state ESPOS_HEALTH_NORMAL clears it (and disarms the ttl); WARN or ALARM raises it.
  • ttl_ms backstop for a test session that goes away, NOT the normal way to clear call again with NORMAL for that. Required when raising: 1..ESPOS_HEALTH_TEST_TTL_MAX_MS. Ignored when clearing.

Returns:

ESP_ERR_INVALID_ARG for a bad key, state or ttl; ESP_ERR_NO_MEM when the condition table is full (CONFIG_ESPOS_HEALTH_MAX_CONDITIONS note a key keeps its slot for the life of the boot, so reuse one key).


function espos_health_reset

Forget every condition and sink (tests).

void espos_health_reset (
    void
) 


function espos_health_snapshot

Copy the current conditions into out (at mostmax ).

size_t espos_health_snapshot (
    espos_health_condition_t * out,
    size_t max
) 

Parameters:

  • out receives the conditions, oldest first
  • max capacity of out in entries
  • out may be NULL to query the count only.

Returns:

how many conditions exist, which may exceed max.


function espos_health_state_str

Human-readable state, for logs and sinks: "normal", "warn", "alarm".

const char * espos_health_state_str (
    espos_health_state_t state
) 


function espos_health_test_expire

Clear the synthetic condition if its ttl has elapsed; otherwise do nothing.

bool espos_health_test_expire (
    void
) 

Cheap and idempotent.

The policy tick calls this, so on a device with espos_health_policy_start() running the backstop fires within a tick. Nothing else drives it, so a caller that reads the conditions should call this first rather than assume a tick has happened GET /api/v1/health does.

Returns:

true if a synthetic condition was cleared by this call.


function espos_health_unwatch_task

Stop watching the calling task (both the task watchdog and the policy).

esp_err_t espos_health_unwatch_task (
    void
) 


function espos_health_watch_task

Watch the calling task: subscribe it to the IDF task watchdog (a task that stops calling espos_health_kick() for CONFIG_ESP_TASK_WDT_TIMEOUT_S panics, which core-dumps, restarts and — before an OTA image is confirmed — rolls back) and register it with the policy, which raisestaskStalled as a fatal ALARM once the task has been silent fortimeout_ms .

esp_err_t espos_health_watch_task (
    const char * name,
    uint32_t timeout_ms
) 

Idempotent per task; a second call updates name and timeout. name is clipped to 15 characters. A watched task must call espos_health_unwatch_task() before it exits. ESP_ERR_NO_MEM when ESPOS_HEALTH_WATCHED_MAX tasks are watched.


function espos_health_worst

Worst state currently recorded — what a single status LED wants to know.

espos_health_state_t espos_health_worst (
    void
) 


Macro Definition Documentation

define ESPOS_HEALTH_F_REBOOT_ON_ALARM

#define ESPOS_HEALTH_F_REBOOT_ON_ALARM `(1u << 0)`

define ESPOS_HEALTH_KEY_MAX

#define ESPOS_HEALTH_KEY_MAX `24`

define ESPOS_HEALTH_MSG_MAX

#define ESPOS_HEALTH_MSG_MAX `96`

define ESPOS_HEALTH_TEST_PREFIX

Keys a synthetic condition may use must start with this.

#define ESPOS_HEALTH_TEST_PREFIX `"test."`

A real condition cannot be impersonated, and a reader can tell at a glance that an ALARM is a drill which matters when the fan-out reaches a chartplotter.


define ESPOS_HEALTH_TEST_TTL_MAX_MS

Longest ttl_ms accepted: a drill nobody clears must end on its own.

#define ESPOS_HEALTH_TEST_TTL_MAX_MS `(300u * 1000u)`



The documentation for this class was generated from the following file espos_health/include/espos_health.h