Tuning COS Alert Rules at Runtime with the Alert Knob

COS charms like Prometheus, Loki, and Mimir automatically collect alert rules from related applications, inject Juju topology matchers, and evaluate them. This works great on Day 1 - deploy, relate, and your alerts light up. But what about Day 2? What if you need to silence a noisy alert from a specific application, bump a threshold, or rename a rule without asking the charm author to cut a new release?

Until now, the answer was: you couldn’t do easily. Not without forking charm configs or resorting to workarounds.

Enter the Alert Knob: a new AlertRulesCustomization helper in cosl v1.11.0 that gives admins declarative, YAML-based control over relation-derived alert rules. It’s already rolled out across prometheus-k8s, loki-k8s, mimir-coordinator, and loki-coordinator.


What it does

The Alert Knob exposes a charm config option named alert_rule_customizations where you supply a YAML document describing two kinds of operations:

Operation What it does
remove Drop matching alert rules (or entire groups)
patch Modify matching rules: rename them, change the expression, adjust the for duration, or merge labels and annotations

Operations are applied in order: remove first, then patch. The original rules are never mutated; the library works on a deep copy.

Matching rules

Each operation has a where block with one or more selectors. All fields in a single where are ANDed together. Multiple entries in the remove or patch list give you OR semantics.

Selector Match type
alert Exact match on the alert name
group Exact match on the rule group name
labels Every key-value pair must exist in the rule
annotations Every key-value pair must exist in the rule

When group is the only selector in a where, the entire group is dropped, including any recording rules it contains. Otherwise, recording rules are never removed or patched.


Demo

Let’s say you’ve deployed Mimir with a cos-configuration-k8s charm that provides these alert rules:

# Rules as seen by the coordinator (relation-derived)
app-1:
  groups:
    - name: first_group_rules
      rules:
        - alert: AlwaysFiringDueToNumericValue
          expr: vector(1) > 0
        - alert: AlwaysFiringDueToAbsentMetric
          expr: absent(nonexistent_metric)
    - name: second_group_rules
      rules:
        - alert: HostDown
          expr: up < 1
          labels:
            juju_application: aval
            environment: production
            zone: can-west
        - alert: LowThroughput
          expr: throughput < 10
          labels:
            juju_application: aval
            environment: staging
    - name: third_group_rules
      rules:
        - alert: HighLatency
          expr: latency > 100
          labels:
            environment: production

Now you decide three things:

  1. The second_group_rules group is irrelevant — drop it entirely.
  2. The LowThroughput alert in staging is too noisy — drop it.
  3. The HostDown alert should have a longer for duration and a higher severity label.

Write your customizations:

# customizations.yaml
remove:
  - where:
      group: second_group_rules
  - where:
      labels:
        environment: staging

patch:
  - where:
      alert: HostDown
      labels:
        juju_application: aval
    set:
      for: 25m
      labels:
        severity: critical

Apply it:

juju config mimir alert_rule_customizations=@customizations.yaml

The resulting rules — after remove, then patch — look like this:

app-1:
  groups:
    - name: first_group_rules
      rules:
        - alert: AlwaysFiringDueToNumericValue
          expr: vector(1) > 0
        - alert: AlwaysFiringDueToAbsentMetric
          expr: absent(nonexistent_metric)
    - name: third_group_rules
      rules:
        - alert: HostDown
          expr: up < 1
          for: 25m
          labels:
            juju_application: aval
            environment: production
            zone: can-west
            severity: critical
        - alert: HighLatency
          expr: latency > 100
          labels:
            environment: production
  • second_group_rules is gone — the entire group, dropped by the group-only selector.
  • LowThroughput is gone — removed by its environment: staging label.
  • HostDown now has for: 25m and severity: critical from the patch.

Patching expressions

You can also replace a rule’s PromQL or LogQL expression. For example, to tighten a threshold:

patch:
  - where:
      alert: HighLatency
    set:
      expr: latency > 500

Renaming alerts

Rename an alert across all groups with a single patch entry:

patch:
  - where:
      alert: OldAlertName
    set:
      alert: NewAlertName

Safety first

The Alert Knob validates your customizations at two levels:

  1. Schema validation: the YAML is parsed and validated with Pydantic. Invalid keys, missing where blocks, or empty selectors immediately block the charm with a clear status message.

  2. Expression validation: after applying customizations, every resulting rule file is validated with cos-tool against the target query language (PromQL or LogQL). If any identifier produces invalid rules, the entire customization is discarded as a no-op and the charm is blocked. The original, unmodified rules continue to serve.

This means you can experiment safely. If your patched expression has a syntax error, nothing breaks. The charm’s debug logs will tell you what the issue is.


Available today

The Alert Knob is available in cosl >= 1.11.0 and has been integrated into:

For full API documentation and more examples, check out the rules_customization module in cos-lib. Also check out the docs. Happy alert tuning!