4 Commits
Author SHA1 Message Date
pipistrello a4aa0e90fe grafana: expose phone LDAP DNS root cause
Verbose Yealink traces show LDAP binds fail before authentication because the provisioned short hostname SERVERPDC cannot be resolved. The failures started before the separate AD unauthenticated-bind hardening, so credentials and that security change are not the cause.\n\nInclude the explicit domain-resolution signature in current-state, trend, top-phone, and operational-log queries. Update panel guidance to direct investigation toward a resolvable LDAP FQDN or stable IP.\n\nThe added LogQL alternative was executed successfully against production Loki 3.7.2.
2026-07-17 14:25:30 +03:00
pipistrello 1a42f89c11 grafana: separate phone health signals from firmware noise
The Yealink dashboard treated nearly every level-3 line as actionable, making routine boot and firmware chatter look like fleet-wide failures.\n\nReplace generic error counts with recent LDAP, provisioning, and internal-database signals. Add dedicated operational trend and log panels, and relabel raw severity as diagnostic-only while retaining web-password, registration, and DHCP context.\n\nValidated all 17 LogQL targets directly against Loki 3.7.2. In the inspected 24-hour sample, provisioning had settled, LDAP failures remained active, and one handset reported a malformed internal database.
2026-07-17 14:19:32 +03:00
pipistrelloandClaude Opus 4.8 f438f3845c Key Yealink phones by MAC, and stop pretending fleet-wide sums are thresholds
Two changes driven by the fleet syslog rollout (~4 -> 106 senders) and the
operator enabling Yealink's prepend-MAC on the autoprovisioned handsets.

1. Alloy: lift the MAC into a real `mac` label.

   UDP syslog offers only three possible identities: sender IP, RFC3164
   HOSTNAME, message body.
   - source_ip is DHCP (observed lease 1200s) so it is NOT stable: a re-IP
     looks like a new device, and a recycled IP silently inherits another
     phone's history.
   - HOSTNAME is useless: Yealink puts its subsystem there (sua/GUI/cfg/sys),
     and after prepend-MAC the bracketed MAC lands in that slot and is rejected
     as a hostname, so the label is simply absent on new lines.
   So the body is the only place identity exists. A loki.process regex stage
   extracts it. Non-Yealink senders (the D-Link switches) don't match and pass
   through with no mac label. Cardinality is safe: mac is 1:1 with a device and
   does not multiply against source_ip.

   Verified on a disposable alloy+loki rig by replaying 406 REAL captured lines:
   0 parse errors, both branches confirmed — Yealink lines get
   mac=80:5e:0c:b2:44:d3, the switch line passes through with no mac label.

2. Dashboard: identity by MAC, comparative panels, no invented thresholds.

   - Variable is now label_values(mac) with allValue ".+" (not ".*") so "All"
     matches only streams that HAVE a mac — excluding switches automatically.
   - Removed the absolute colour thresholds. They were derived from a ~4-hour
     sample of TWO atypically verbose handsets and were already wrong at 2
     phones (actionable errors read 537 against a red line of 800; scheduler
     timeouts 193 against yellow at 150). They were fleet-wide SUMS, so they
     scale with phone count rather than with health. The fleet onboarded today,
     so no representative 24h baseline exists to replace them with — panels are
     comparative (worst offenders) until a week of real data exists. The one
     surviving threshold is "phones failing provisioning" >= 1, which is a
     qualitative fault rather than a magnitude.
   - Stats now count AFFECTED DEVICES, not fleet-wide line totals.
   - "Phones reporting" -> "Phones heard from", documented as NOT a health
     count: most of the fleet runs at Yealink default log level 3 (Error) and
     is silent unless something breaks. A silent phone is a healthy phone.
   - Registration/DHCP panels labelled level-6-only: they key on info/notice
     lines that level-3 phones never send, so an absent phone there means
     "quiet", not "unregistered".
   - Added "loudest phones by log volume" to surface handsets left at level 6
     (two of them are ~80% of all volume).
   - refresh 1m -> 5m.

   Verified by firing all 14 panel queries CONCURRENTLY against a sink loaded
   with replayed real data: 14/14 return data, 0 rejected. Last time these were
   verified sequentially, which is what let the 429 bug ship.

Note: data predating the MAC rollout has no mac label and won't appear.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 19:55:09 +03:00
pipistrelloandClaude Opus 4.8 13b1887a16 Add provisioned Grafana dashboard for Yealink IP phones
Phones are ~99% of this Loki's log volume and nothing surfaced them. Adds a
file-provisioned dashboard (uid yealink-phones, folder "Network") plus the
provider config and the two grafana bind mounts it needs.

Built against a 24h baseline of real handset traffic; all 14 panel queries were
executed against live Loki before commit. Three findings shaped it, each of
which contradicts the obvious reading of the data:

- severity=emergency is NOT an emergency. All 24 emergency lines are the phone
  printing its own log-level table at boot ("sys log :type=1,E=3,W=4,N=5,I=6,D=7")
  — Yealink emits its logging config at level 0. Alerting on it would be 100%
  false positive. Filtered out.
- "LSYS<3+error> rtpcap get len not enough" is 467 of 830 error lines (56%) — a
  noise floor, not a fault. Excluded from every "actionable" panel.
- The hostname label cannot identify a device: its values are Yealink subsystems
  (sua, GUI, sys, cfg, ipp, dev, WEB, ATP), because the phones put the subsystem
  in the RFC3164 HOSTNAME field. Everything keys on source_ip. Yealink lines are
  separated from switch/router traffic on the same Loki by the |~ "<[0-7][+]"
  module marker (26,450 Yealink vs 33 non-Yealink lines in the baseline).

Thresholds are seeded from measured 24h counts: check passwd err 21 (auth),
data_task schedule time out 117 (handset scheduler overrunning its 30s
threshold — the best "phone is unwell" proxy), tftp/provisioning failures 3,
Register: update server 175 (SIP beat, 120s period), DHCP lease 9.

Caveat: only two handsets report today, so thresholds are baselined on a
2-phone sample and want revisiting once the fleet is onboarded.

Dashboards are file-provisioned, so UI edits are not persisted — change the
JSON here and redeploy.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 15:53:11 +03:00