Skip to main content
Version: current

alert


cmr/alert​

Package: cmr
Type: Directory

An alert rule applies conditions to a set of devices, selected with labels, and runs its actions on the CMR controller when a condition becomes true. A rule fires when a device starts matching its conditions and fires again only after the device stops matching and then matches again. See CMR - Alert rules for how alert rules work.

FlagNameDescription
XdisabledAlert rule is disabled.
ArgumentTypeDescription
name ( unset )stringName of the alert rule.
labels ( unset )object

Select the devices, by labels, that the alert rule applies to. Supports + and - signs as AND and AND NOT operators, respectively; if no sign is provided, the OR operator is used.

When labels is not set, the rule applies to all connected devices (labels=all).

connected ( unset )bool

Alert when the device connects or disconnects.

Default: no.

cpu-above ( unset )num

Alert when CPU load is above the given percent value (0..100); a value equal to the threshold also fires the rule.

Default: unset.

cpu-below ( unset )num

Alert when CPU load is below the given percent value (0..100); it fires only for values strictly below the threshold.

Default: unset.

hdd-above ( unset )num

Alert when disk (HDD) load is above the given percent value (0..100); a value equal to the threshold also fires the rule.

Default: unset.

hdd-below ( unset )num

Alert when disk (HDD) load is below the given percent value (0..100); it fires only for values strictly below the threshold.

Default: unset.

mem-above ( unset )num

Alert when RAM load is above the given percent value (0..100); a value equal to the threshold also fires the rule.

Default: unset.

mem-below ( unset )num

Alert when RAM load is below the given percent value (0..100); it fires only for values strictly below the threshold.

Default: unset.

health-above ( unset )num

Alert when the health sensor named by health-value goes above the given value; a value equal to the threshold also fires the rule. Devices that do not report the sensor are ignored.

Default: unset.

health-below ( unset )num

Alert when the health sensor named by health-value goes below the given value; it fires only for values strictly below the threshold. Devices that do not report the sensor are ignored.

Default: unset.

health-value ( unset )stringName of the health sensor to compare, matching a name from /system/health/print, for example cpu-temperature. Uses the [cpu-temperature] style placeholder in the action text.
interface-change ( unset )alt { change: enum (any) , change: ubit (running, not-running, added, removed) }

Alert on an interface state change:

  • any (default) - Any state change.
  • running - An interface's link comes up.
  • not-running - An interface's link goes down.
  • added - An interface is first seen.
  • removed - An interface is no longer seen.

running and not-running follow the link state of the interface, not its admin state: enabling or disabling an interface with no connection does not fire, while an actual link loss or recovery does.

interface-type ( unset )alt { type: enum (any) , type: ubit (ethernet, wifi, bridge) }

Restrict interface-change alerts to one interface type:

  • any (default) - Interfaces of any type.
  • ethernet - Ethernet ports.
  • wifi - Wireless interfaces.
  • bridge - Bridge interfaces.
upgrade-available ( unset )bool

Alert when a newer RouterOS version is available for the device.

Default: no.

disconnected-more-than ( unset )time

Alert when a device has been disconnected for longer than the given time.

Default: unset.

upgrade-job-done ( unset )alt { cond: enum (yes) , cond: ubit (success, fail) }

Alert when an upgrade job finishes; the rule fires once for the whole job, when the last device of the job is done:

  • yes (default) - Any outcome.
  • success - The job finished successfully.
  • fail - The job failed.

A job that ends without installing an upgrade is treated as a failed job, so the yes and fail variants both fire.

upgrade-done ( unset )alt { cond: enum (yes) , cond: ubit (success, fail) }

Alert when the upgrade of a device finishes; the rule fires once for every device that finished its upgrade:

  • yes (default) - Any outcome.
  • success - The upgrade finished successfully.
  • fail - The upgrade failed.

When no newer version is available for a device, its upgrade still finishes as a failure (the [upgrade-error] is no upgrade available), so the yes and fail variants both fire.

unpaired-device-connected ( unset )bool

Alert when a device that is not paired to the controller connects.

Default: no.

rebooted ( unset )bool

Alert when the device reboots.

Default: no.

log-topics ( unset )multi { array-id, array-id, topic: super { ! , topic: enum } }Log topics to follow; prefix a topic with ! to exclude it.
log-regex ( unset )stringAlert on log lines the device sends to the controller whose message matches the given regular expression.
action.log ( unset )stringWrite a message to the log when the alert fires. The message may contain placeholders such as [device] or [cpu-usage], replaced with the values of the device that fired; an inapplicable placeholder is rendered as unknown.
action.log-topics ( unset )multi { array-id, topic: enum }

Log topics to write the message to.

Default: cmr,info.

action.log-prefix ( unset )stringPrefix for the log message. When set, the prefix is prepended to the message without a separator; when unset the message is written as it is.
action.script ( unset )enumRun a named system script when the alert fires. The script must allow itself to run with reduced rights; otherwise the action fails with executing script ... from cmr failed ... (not enough permissions). Set the script's dont-require-permissions to yes.
action.script-vars ( unset )multi { array-id, script-var: string }List of variable names to make available to the action.script: each variable is filled with the value for the device that fired. For example device, identity, address, board, version become $device, $identity, and so on in the script. A name with a hyphen (for example cpu-usage) cannot be read as an ordinary script variable.
action.http-url ( unset )stringHTTP webhook URL to call when the alert fires. On a failed call the alert writes alert "<name>" HTTP action failed: <reason> to the cmr topic with warning severity and counts the failure in action-failures; the other actions still run.
action.http-method ( unset )enum (get | post | put | delete | head | patch)

HTTP method used for the webhook:

  • get
  • post
  • put
  • delete
  • head
  • patch
action.http-body ( unset )stringHTTP request body sent to the webhook. Supports the same placeholders as the log message, substituted per device (for example [device] and [version]). No Content-Type header is added automatically; add one with action.http-headers when the receiver expects it, for example Content-Type: application/json for a JSON body.
action.http-headers ( unset )multi { header: string }Additional HTTP headers sent with the webhook, each as Header: value. The request also carries the controller's RouterOS <version> user-agent.
reset-on-disconnect ( unset )enum (yes | no) { yes:0, no:-1 }Clear (reset) the alert state when the device disconnects, so a reconnecting device fires the rule again: yes or no. Default: no (the fired state is kept across a disconnect).
severity ( unset )enum (critical | high | medium | low)

Severity of the alert:

  • critical
  • high
  • medium (default)
  • low
category ( unset )multi { array-id, category: enum }Categories assigned to the alert. The category is a grouping label: the server assigns one according to the rule's conditions (for example performance, availability, version, network, or logging), and it can be changed to any value.
Read-only ArgumentTypeDescription
devicesnumNumber of devices the alert rule matches (the devices selected by labels). When the rule covers all devices, this count includes the controller itself.
devices-onnumNumber of devices the alert is currently fired on.
firednumTotal number of times the alert has fired.
action-failuresnumNumber of times an alert action has failed, for example a webhook that could not be reached.