Commit Graph

664 Commits

Author SHA1 Message Date
MarekWo 0d69947bfd docs: cover the My Repeaters (repeater administration) feature
user-guide.md gains a full My Repeaters section (adding repeaters,
saved passwords + the wrong-password/unreachable ambiguity, path
editing, and the six management tools) and fixes the stale Quick
Access placement list (now 12 items, real labels). architecture.md
documents the /api/repeaters/* endpoint family with its shared error
mapping, the repeaters DB table, and a DeviceManager subsection on
serialization, login sessions, CLI reply correlation, settings
batches and the actions whitelist. whatsnew.md adds three New
features bullets, a Reliability note (console login reports role),
and Deploy notes (meshcore>=2.3.7 rebuild, plaintext password
trade-off). rpt-mgmt.md now points to the built-in panel first.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 19:33:27 +02:00
MarekWo a036d6af6c feat(repeaters): Actions tool (stage 8)
- POST /api/repeaters/<pk>/action {action}: whitelisted one-shot commands
  via repeater_cmd_wait — zerohop_advert (advert.zerohop), flood_advert
  (advert), clock_sync (clock sync), reboot; admin-gated
- Replies surfaced verbatim with an ok flag (reply starts with OK), so
  firmware refusals like "ERR: clock cannot go backwards" show as-is
- reboot special case: the firmware restarts without ever replying, so a
  clean send followed by silence reports success ("Reboot command sent")
- Actions pane: action rows with inline result lines, flood advert
  styled as warning ("Not recommended - high network load"), Danger zone
  card with confirm()-guarded Reboot and a muted note that erase is only
  available on the USB serial console (firmware restriction)

Verified live against PL-KRA Wegrzce 2: zero-hop advert replied
"OK - zerohop advert sent" and the advert arrived back at our device
(log: Advert from 'PL-KRA Wegrzce 2' type=2); clock sync surfaced the
firmware refusal; reboot confirm-cancel path checked in Playwright
(reboot itself intentionally not executed).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 19:22:48 +02:00
MarekWo 92f5a7edca feat(repeaters): Settings tool (stage 7)
- GET/POST /api/repeaters/<pk>/settings?section=X: sequential text-CLI
  get/set batches via repeater_cmd_wait; per-field values/errors on read,
  per-field ok|failed|reboot_required classification on write (reply
  starting ok/password now = ok, mention of reboot = reboot_required);
  admin-gated via new _require_repeater_admin helper (also reused by /cli)
- Settings pane: accordion with 8 sections (Basic, Radio, Location,
  Features, Network health, Advertisement, Operator info, Advanced),
  lazy-loaded on first expand (each field is one mesh round-trip),
  per-section Refresh/Apply, dirty tracking with header badges
- Field types: on/off switches, 0/1 switches, select (loop.detect),
  composite radio (freq,bw,sf,cr) with reboot warning + confirm dialog,
  write-only admin password (updates the saved password on success),
  owner.info textarea with pipe-to-newline mapping
- Failed writes keep the field dirty and surface the firmware reply
  (e.g. "Error: interval range is 60-240 minutes") under the input
- Strip "%" suffix when loading number fields (get dutycycle -> "70.0%")

Verified live against PL-KRA Wegrzce 2: all 8 sections read cleanly,
flood.max 64->63->64 write round-trip, firmware range rejection surfaced
per-field, radio confirm-cancel path, dutycycle % parsing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 14:43:16 +02:00
MarekWo 4fb6b2a847 fix(repeaters): CLI terminal fills the viewport vertically
height: max(260px, calc(100vh - 430px)) instead of a fixed 340px, so
the scrollback grows with the window while staying usable on small
screens.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 13:57:26 +02:00
MarekWo 3e80ea0c85 feat(repeaters): CLI tool (stage 6)
Remote text console for the managed repeater. Core mechanism:
repeater_cmd_wait() sends the command and synchronously waits for the
reply — CLI replies arrive as CONTACT_MSG_RECV txt_type=1 with no
protocol-level correlation, so correlation = repeater lock (single
command in flight) + sender-prefix match. A single-slot waiter is
checked at the top of _on_dm_received: matched CLI replies are
consumed there and never stored as chat DMs; unmatched ones (e.g.
console fire-and-forget cmd) keep the legacy DM behavior. Wait time
derives from the device-suggested timeout (10-45 s clamp).

POST /api/repeaters/<pk>/cli is admin-gated (403 for guest sessions:
firmware silently drops guest text commands, which would look like a
timeout). Pane: dark terminal styled after the Console module, quick-
command chips, Enter to send, per-repeater arrow-key history in
localStorage, elapsed-time line, inline timeout errors (lost replies
happen over radio — a manual retry typically succeeds).

Settings (stage 7) and Actions (stage 8) will reuse repeater_cmd_wait.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 13:45:51 +02:00
MarekWo 2b6d82a11c feat(repeaters): Neighbors tool with map view (stage 5)
List view: all zero-hop neighbours (fetch_all_neighbours paginates the
firmware's ~14-entry pages), rows show resolved contact name (or
[pubkey prefix] for unknown repeaters), heard-ago and SNR; count in
the toolbar. Names/positions are enriched server-side by prefix-
matching device contacts with the DB contact cache as fallback.

Map view (List/Map toggle, shown only when something is mappable):
Leaflet with the managed repeater as a red marker, positioned
neighbours in green, dashed connection lines labeled with permanent
SNR tooltips, and a footnote counting neighbours without a known
position. Toggle hidden entirely when no coordinates exist.

Verified live on PL-KRA Wegrzce 2: 37/37 neighbours fetched, 20
mappable, SNR labels rendered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 13:23:43 +02:00
MarekWo 995cb49f4c feat(repeaters): Telemetry tool (stage 4)
All Cayenne LPP channels shown at once (no channel dropdown like the
standard app): one card per channel with typed rows — icon, name,
value and unit (voltage V, temperature °C, current A, power W,
humidity %, GPS lat/lon+alt, etc.). Channel 1 is labeled 'device'
(repeater's own vitals). Refresh + updated-ago label; timeout renders
inline with a retry button (exercised for real on a 2-hop repeater —
first attempt timed out, retry succeeded with 2 channels).

Backend: repeater_req_telemetry in device_manager (serialized under
the repeater lock; the old name-only request_telemetry stays for
sensor nodes) + login-gated GET /api/repeaters/<pk>/telemetry.
GPS values arrive as {latitude, longitude, altitude} objects from
meshcore 2.3.7 — formatter handles both object and array forms.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 08:36:02 +02:00
MarekWo 707fd20426 feat(repeaters): Status tool (stage 3)
First management tool. The Status tile auto-fetches on open and shows
a compact three-section table (not MO's oversized tiles):
- System: battery %/V (linear 3.3-4.2 V estimate), uptime, repeater
  clock, queue length, debug/error events
- Radio: last RSSI/SNR, noise floor, TX/RX airtime
- Packets: sent + received (flood/direct split), duplicates, RX
  errors, channel utilization computed client-side as
  (tx_air+rx_air)/uptime (matches the firmware's own 10.09% reading)
Refresh button + "updated Ns ago"; errors render inline with retry.

Backend: GET /api/repeaters/<pk>/status and /clock, both gated on an
existing login session (401 need_login) to fail fast instead of a
2-min timeout. Clock loads as a follow-up request so the table
appears immediately, then the clock row fills in.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 08:16:58 +02:00
MarekWo d825f4dc29 fix(repeaters): remove 8px modal margin on narrow viewports
The fullscreen-modal override (@media max-width:576px) that zeroes the
default .modal-dialog 0.5rem margin enumerates modal IDs by hand and
omitted #repeatersModal, so on narrow windows the My Repeaters modal
kept an 8px top/left gap that Contact Management did not. Added
#repeatersModal to the three override selector lists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 08:12:04 +02:00
MarekWo f288586ea6 feat(repeaters): Repeater Management panel (stage 2)
/repeaters/manage?pubkey=... — per-repeater management panel opened
automatically after login from the My Repeaters list:
- header card: name, shortened pubkey with copy, current path,
  location, ADMIN/GUEST badge from the captured login session
- tools grid (Status / Telemetry / Neighbors / CLI / Settings /
  Actions) with pane placeholders; CLI+Settings+Actions are locked
  for guest logins (firmware accepts text CLI from admins only)
- auto-login with the saved password when the in-memory session is
  gone (e.g. after app restart), password-modal fallback with retry
- REST: GET /api/repeaters/<pk> (merged entry + session state),
  GET .../session, POST .../logout

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 07:24:52 +02:00
MarekWo ca2a6bacf9 feat(repeaters): My Repeaters panel (stage 1)
New full-screen panel (main menu / FAB) for repeater administration:
- repeaters table (saved per-pubkey admin password, login metadata);
  plaintext per observer_brokers precedent (single-user LAN app)
- REST /api/repeaters: list merged with device contact truth, add
  (device REP contacts only), set/clear password, remove, login
- device_manager: _repeater_lock serializes all repeater ops (companion
  firmware has a single pending-request slot); repeater_login now
  filters LOGIN_SUCCESS by pubkey_prefix and captures is_admin/
  permissions into an in-memory session store
- login timeout UX: wrong password and unreachable repeater are
  indistinguishable (firmware stays silent) - error message names both
- panel UI: add-picker, password modal (remembered password), per-
  repeater path editor reusing /api/contacts/<pk>/paths + Leaflet
  repeater map picker (adapted from the DM panel)
- meshcore pin bumped to >=2.3.7 (send_login_sync era API, fixed
  LOGIN_SUCCESS/LOGIN_FAILED parsing; container already ships 2.3.7)

Stage 2 (management panel with Status/Telemetry/Neighbors/CLI/
Settings/Actions tools) follows after user review.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 22:55:44 +02:00
MarekWo 87b30f63b3 docs: Update Meshcore flasher URL 2026-07-14 20:41:03 +02:00
MarekWo e93961166c docs: cover the Observer (MQTT packet capture) feature
user-guide.md gains the Observer tab section (master switch, IATA code,
scheduled flood adverts, broker management, live badges) and the updated
Settings tab list. architecture.md documents the ObserverManager design
(hot path, per-broker paho clients, wire-format fidelity, reload without
restart), the observer_brokers table, the /api/observer/* endpoints and
the observer_status socket event. whatsnew.md announces the feature and
notes the paho-mqtt dependency (image rebuild via standard mcupdate) and
plaintext broker credentials.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 20:36:19 +02:00
MarekWo ca44acab7f fix(observer): don't log deliberate client teardown as a lost connection
Reload/stop tears old MQTT clients down on purpose; the disconnect
callback was logging 'lost connection ... auto-reconnecting' and
recording last_error for every config save. Mark clients as stopping
before teardown and return early in the callback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:48:28 +02:00
MarekWo 79ee58d91a feat(observer): scheduled flood adverts
Observer-owned daemon thread checks every 10 minutes whether
advert_interval_hours has elapsed since the persisted last-advert
timestamp and sends a flood advert via the existing DeviceManager
wrapper. Timestamp persists across restarts so redeploys don't
re-advertise early; failures and disconnected device skip the send
without touching the timestamp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:08:37 +02:00
MarekWo 668d8d5c79 feat(observer): Settings > Observer tab UI
New Observer tab: master switch, IATA location code, flood advert
interval, live status line with packet counters, broker list with
connection badges (green/red + error tooltip), enable switches and a
stacked add/edit modal (blank password on edit keeps the stored one).
Live badge/counter updates via the observer_status socket event.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:07:09 +02:00
MarekWo a60ff138ff feat(observer): REST API for brokers, settings and status
/api/observer/settings (GET/POST, partial updates, IATA + advert
interval validation), /api/observer/brokers CRUD (passwords never
returned; omitted password stays, empty string clears) and
/api/observer/status. Every mutating endpoint hot-reloads the
ObserverManager so config changes apply without a restart.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:03:32 +02:00
MarekWo 91af4324de feat(observer): wire packet capture into DeviceManager and app startup
DeviceManager hands device identity (name/pubkey/self_info/device_info)
to the observer after connect and feeds every RX_LOG_DATA packet into
handle_raw_packet() before the GRP_TXT echo gate, wrapped in its own
try/except so observer failures can never affect chat. create_app()
builds the ObserverManager and links it with the DeviceManager.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:01:54 +02:00
MarekWo d3590f9228 feat(observer): MQTT publishing module (meshcore-packet-capture compatible)
New app/observer.py: ObserverManager owns per-broker paho-mqtt clients
(own network threads, QoS 0, LWT + retained status) and the packet
fan-out. Packet decode/hash/format functions are ports of upstream
meshcore-packet-capture including its legacy path fallbacks and
unknown-version quirk; validated against 10 live packets from the
reference observer network (hash/route/type/len all byte-identical).
Adds paho-mqtt==2.1.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 19:00:37 +02:00
MarekWo 038ab47190 feat(observer): observer_brokers table and DB CRUD
New observer_brokers table (name/host/port/credentials/TLS flags) plus
Database CRUD mirroring the analyzers pattern. Groundwork for the MQTT
packet-capture observer; no behavior change yet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 08:15:55 +02:00
MarekWo 529b114104 docs: cover raw resend, touch-send, analyzer-as-row, and reliability fixes
Brings the docs up to date with the two undocumented batches on dev since
the last docs commit (d16093b). whatsnew.md is split by ship date: a
2026-06-26 section (already live on main) gains the raw channel-message
resend feature, the Edit-message button rename, and the log-loop /
connection-badge / failed-reconnect fixes; a new 2026-07-02 section covers
touch-device explicit send, DM line-break preservation, and the Letsmesh
analyzer becoming a normal editable row. A fresh Unreleased(since debb711)
opens on top.

architecture.md documents the channel_messages.raw_packet column, the
POST /api/messages/<id>/resend endpoint, the new /api/status fw/resend/
path-hash fields, the DeviceManager raw-resend + liveness-watcher
behavior, the werkzeug log-filter, and the analyzer seeding change.
user-guide.md adds a Message Actions section (Edit/Resend), the touch
Enter-vs-Send note in channels and DM, and rewrites the Analyzer tab to
reflect Letsmesh-as-normal-row.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 21:14:03 +02:00
MarekWo debb711b71 feat(analyzer): seed Letsmesh as a normal DB row and flip switch UX
Letsmesh Analyzer is now a regular analyzers row, seeded once via an
app_settings flag on first startup. It can be renamed, disabled, or
deleted like any user entry. The frontend drops the special built-in
pseudo-row and the built-in slot in the chooser.

Click resolution updated: 0 enabled → toast "No analyzer configured";
1 enabled → open directly; multiple with a default → open default;
multiple without default → chooser.

The row switch and the edit modal now show "Enabled" (checked = active),
which matches intuition. The DB column remains is_disabled so no data
migration is needed — the UI simply inverts the value on read/write.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-01 21:56:55 +02:00
MarekWo ca4c77da36 feat(chat): require explicit send button on touch devices
Enter on a phone's virtual keyboard now inserts a newline instead of
sending, in both group chat and DM. Detected via the pointer:coarse
media query - no user-facing setting, since accidental sends from a
mistapped Enter were the actual complaint. Desktop keeps Enter-to-send.
2026-06-29 09:31:50 +02:00
MarekWo 7be3e26161 fix(dm): preserve line breaks in sender's own DM message bubbles
.dm-content was missing white-space: pre-wrap, unlike .message-content
in the group chat, so multi-line DMs collapsed to one line on the
sender's own screen even though the recipient saw them correctly.
2026-06-29 09:31:35 +02:00
MarekWo 0eb7c772a7 fix(status): stop the connection badge from lying about device state
loadMessages() called updateStatus('connected') on every successful
/api/messages fetch -- a DB read that succeeds regardless of whether
the meshcore device itself is connected. Same conflation in dm.js,
which marked 'connected' on Socket.IO transport connect (browser <->
server), unrelated to device <-> server connectivity. Together these
silently overwrote the real status (e.g. an actual device disconnect)
back to "Connected" on the very next channel switch or auto-refresh.

Also fixed the 'device_status' socket handler in app.js, which built
its own markup against a `#connectionStatus` element that doesn't
exist in any template (dead code since introduction) instead of
calling the shared updateStatus() against the real `#statusText`
element. Added a 60s polling fallback via loadStatus() in case a
push is ever missed on a long-lived tab.

Verified locally: ran the app against no real device (always
disconnected) and confirmed via Playwright that the badge now stays
correctly "Disconnected" through repeated loadMessages() calls,
instead of flipping back to "Connected" as before the fix.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-26 19:04:41 +02:00
MarekWo 5b776bc927 fix(device): liveness watcher must keep retrying once fully disconnected
A forced reconnect can fail without raising (e.g. self_info comes back
empty), leaving _connected False. The watcher loop only re-checked
staleness while _connected was True, so a single failed reconnect
attempt silently stopped all further healing forever — nothing else
retries non-BLE transports. Combined with /health/strict treating
"not connected" as OK (so the device stays connected forever), the
container needed a manual restart to recover.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-26 18:08:40 +02:00
MarekWo 5202665c1a fix(logs): break werkzeug ↔ socket.io polling feedback loop
The MemoryLogHandler broadcasts every record to the /logs namespace, but
with async_mode='threading' Socket.IO falls back to long-polling. Each
polling request is logged by werkzeug, the broadcast wakes the pending
poll, the client immediately re-polls, and the cycle repeats — producing
10+ requests/sec per open System Log tab. Filter werkzeug access logs
for /socket.io/ and /api/logs/ paths so neither the buffer nor the
broadcast trigger the loop.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-10 20:25:23 +02:00
MarekWo f57b9474bb fix(channels): skip self-echoes of raw resends in _on_channel_message
CMD_SEND_RAW_PACKET bypasses sendFlood()'s _tables->hasSeen() self-mark.
The firmware seen-table is 160 entries in RAM and rolls over in minutes
on a busy mesh, so a resend issued after the original hash got evicted
comes back via repeater echo, fails hasSeen (entry gone), and is
delivered as RESP_CODE_CHANNEL_MSG_RECV. Result: the user's own resend
appears as an "incoming" message from themselves a few minutes later.

Detection: in _on_channel_message compute the expected pkt_payload from
the event's (channel_secret, sender_timestamp, txt_type, raw_text) and
ask the DB if any own row already has that exact hash. If yes, treat as
self-echo and return — no DB insert, no SocketIO emit. Index idx_cm_pkt
keeps the lookup cheap. Guarded with try/except so any detection failure
falls through to normal handling — we never want a check bug to drop
real inbound messages.

Documented as a known limitation in reference_meshcore_raw_packet.md;
this lifts that caveat.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-10 19:45:39 +02:00
MarekWo e14c4fab06 fix(channels): include message id in /api/messages response
/api/messages was dropping channel_messages.id when building the response
dict. The frontend therefore had msg.id=undefined for history-loaded rows,
so createMessageElement's "typeof msg.id === 'number'" guard suppressed
the raw-resend button, and refreshMessagesMeta couldn't find the wrappers
either (data-msg-id only gets set when msg.id is truthy). End result: the
raw-resend button only appeared on messages sent in the current session
and was lost on every channel switch or page reload.

Surfacing id costs nothing — the field already exists on the row — and
also unblocks any future per-message UI that needs to address the row
directly. createMessageElement already handles undefined id correctly for
older clients, so this is a one-way frontend benefit.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 19:06:08 +02:00
MarekWo 2f8f765af5 fix(channels): backfill raw-resend buttons when displayMessages races loadStatus
loadStatus and loadMessages fire in parallel at page init. Whichever lost
the race left own-message bubbles without the raw-resend button — and once
the row was rendered, neither displayMessages (no re-render on switch
without going back through createMessageElement, which is gated on
window.deviceCaps at the time of call) nor refreshMessagesMeta (skips rows
that already have route info) would patch it back in. So the button only
ever appeared for messages sent in the current session.

Walk visible own bubbles in two places — at the end of loadStatus and at
the end of displayMessages — and inject the button on any row missing it.
Idempotent (skips bubbles that already have .btn-raw-resend), cheap (no
network), and covers both the page-reload race and the channel-switch /
archive-view paths.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 19:01:07 +02:00
MarekWo e3211dd536 fix(channels): don't render raw-resend button against a _pending_<ts> id
The optimistic-send path renders the bubble with msg.id = "_pending_<ts>"
before the API confirms. The previous PR baked that string straight into
onclick="resendChannelMessageRaw(_pending_<ts>, this)" — an undefined
identifier — so the first click after sending threw ReferenceError and
nothing happened.

Skip the raw-resend button entirely while msg.id is non-numeric, then run
a one-shot refreshMessagesMeta([real_id]) right after the optimistic id
swap so the button shows up immediately even on channels where no echoes
arrive (so the existing echo-driven inject path never fires).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 18:07:04 +02:00
MarekWo 00f0544a47 feat(channels): UI for raw resend + clarify the edit-message button
PR #5 of 5. Wires the user-facing controls for the raw-resend feature.

Channel messages (own):
- The existing arrow-repeat button only pasted content into the composer
  for hand-edits, which the "Resend" tooltip mis-named as a true resend.
  Rename it to "Edit message" with a pencil-square icon.
- Add a new arrow-repeat button that POSTs to /api/messages/<id>/resend.
  Tooltip explains the actual semantics ("rebroadcast same packet so
  unreached repeaters can pick it up"). Spins .btn icon while in flight,
  shows a toast on result. Rendered only when the cached
  /api/status.supports_raw_resend is true (firmware ≥1.16).
- Inject the same buttons in updateMessageMetaDOM so history items loaded
  before window.deviceCaps was populated still get the new button on the
  next echo-driven meta refresh.

DM messages: rename the equivalent paste-button to "Edit message" with
the same pencil icon for UI consistency. The protocol-level retry stays
unchanged — there's no per-DM raw resend button (auto-retry covers it).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 15:02:25 +02:00
MarekWo 201bc137e5 fix(channels): also pass path_hash_size on the echo-driven raw_packet rebuild
Earlier path_hash_mode fix updated the send-time build but the matching
edit to _refresh_raw_packet_if_drifted didn't make it into commit 10df846.
For channels where the secret isn't available at send time, guess_pkt_payload
stays None and raw_packet is created for the first time in this fallback
path (triggered when echo correlation matches via the channel-hash branch).
Without the path_hash_size argument the build defaulted to 1-byte hashes,
producing the same mixed-size badge the prior fix was meant to eliminate.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 14:54:42 +02:00
MarekWo 7bdccac60a chore(status): expose path_hash_mode/size for debugging resend snapshots
A test resend was producing 1-byte hash entries despite the device being
configured with path_hash_mode=1 (2-byte). Surfacing the cached values
in /api/status makes it possible to verify from the client whether the
device_info fetch at connect actually populated the cache, without
shelling into the container.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 14:48:20 +02:00
MarekWo 10df8464b7 fix(channels): honor device path_hash_mode when building raw_packet
Resends were building raw_packet with the default 1-byte path-hash size,
ignoring the device's actual path_hash_mode. When path_hash_mode=1 (2-byte
hashes) the original send produced 2-byte path entries in repeater echoes,
but the resend's path_len byte said "1-byte" — so post-resend echoes
appended 1-byte hashes, mixing into the badge as inconsistent tokens
(e.g. "44D8, D103, E7" — the trailing E7 was a 1-byte fragment).

Cache path_hash_mode from DEVICE_INFO at connect (fw_ver_code >= 10) and
expose path_hash_size = max(1, mode+1). Pass it through to
_build_grp_txt_raw_packet in send_channel_message and the clock-drift
refresh path. Keep cache in sync with set_param('path_hash_mode', N).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 14:40:21 +02:00
MarekWo d23e865f35 feat(channels): merge post-resend echoes into existing repeater badge
PR #4 of 5. After a successful resend, re-arm _pending_echo with the
original msg_id and known pkt_payload so echoes from previously-unreached
repeaters that pick up the rebroadcast are classified as 'sent' and carry
msg_id in the SocketIO emit.

The frontend echo handler now collects forced msg_ids and passes them to
refreshMessagesMeta(forceIds), which bypasses the "already has route info,
skip" guard for those ids. End result: clicking resend extends the
repeater list on the existing message's badge in place — no duplicate row,
no stale count.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 14:32:44 +02:00
MarekWo 4729055900 feat(channels): firmware version gate for raw resend (requires ≥1.16)
CMD_SEND_RAW_PACKET (0x41) was introduced in companion-v1.16.0
(FIRMWARE_VER_CODE bump 11 → 13). Older firmware returns
ERR_CODE_UNSUPPORTED_CMD with no useful context for the user.

Capture fw_ver_code from the DEVICE_INFO event at connect (re-using the
existing send_device_query call) and expose a supports_raw_resend
property. The resend endpoint now refuses early with a clear message
("Firmware too old for raw resend, need ≥1.16, device reports
fw_ver_code=N") and /api/status surfaces both fw_ver_code and the
supports_raw_resend flag so the UI can hide or disable the button on
older firmware.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 13:03:48 +02:00
MarekWo 9c48518771 fix(channels): surface meshcore lib's error_code/code_string on resend failure
The lib's reader.py wraps device ERROR frames as {error_code, code_string},
not {reason, error}. The previous extraction collapsed every device error to
"unknown error", hiding the actual ERR_CODE_* the firmware sent back. Check
code_string/reason/error in order, then fall back to a raw error_code, then
"unknown error".

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 12:47:44 +02:00
MarekWo 67c59cc341 feat(channels): backend resend endpoint via CMD_SEND_RAW_PACKET
PR #3 of 5. Adds POST /api/messages/<msg_id>/resend, which re-broadcasts an
own channel message verbatim using the raw_packet bytes captured at send
time. Pushes the wire bytes directly through companion command 0x41
(CMD_SEND_RAW_PACKET), bypassing the higher-level send paths so repeaters
dedupe by packet hash via Mesh::hasSeen — only previously-unreached nodes
will pick up the resend.

Returns 404 for unknown msg_id, 400 for not-own / missing snapshot /
disconnected device, 500 for unexpected device errors.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 12:39:34 +02:00
MarekWo fa0b1c9109 feat(channels): capture raw_packet at send time for raw resend
PR #2 of 5. Builds the full GRP_TXT wire bytes (header + transport_codes if
scoped + path_len + encrypted payload) from the ts+0 pkt_payload guess and
stores it in channel_messages.raw_packet right after the send. When echo
correlation later identifies the actual pkt_payload (potentially using a
different ±dt candidate due to host/firmware clock drift), the raw_packet is
rebuilt from the actual one so a future resend matches the original packet
hash and dedupes at the repeaters.

Transport-scope codes are computed in Python via HMAC-SHA256(scope_key,
payload_type||payload)[:2], mirroring TransportKey::calcTransportCode in
MeshCore Core (including the 0x0000/0xFFFF reservations).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 12:27:17 +02:00
MarekWo bed838c365 feat(db): add channel_messages.raw_packet column for raw resend
Stores the full hex packet snapshot (header+transport_codes+path_len+payload)
captured at send time, enabling future "Resend (same packet hash)" feature
that lets repeaters deduplicate via Packet::calculatePacketHash while still
forwarding to nodes that didn't hear the original.

NULL for received and pre-migration rows (resend button stays disabled there).
PR #1 of 5; subsequent PRs wire up send-time capture, backend endpoint, echo
list merge, and UI.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 12:15:44 +02:00
MarekWo d16093b459 docs: add whatsnew.md release notes for the upcoming main merge
User-facing summary of changes since fd2b3d0 grouped into New features
(custom analyzers, apply-path button, DB Optimize button, auto retention,
sluggish-device watchdog), Reliability & polish (polling-only fix for
load freezes, channel cache, scope-key refresh, multi-byte path
rendering, TCP self-heal, region scope on slots >7, console parser),
and a deploy note about restarting the host watchdog. Also linked from
user-guide.md's Getting Help section. File is meant to be refreshed
before each merge to main.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-08 12:04:19 +02:00
MarekWo 53ef2759d5 docs: cover analyzer settings, vacuum/optimize, path apply, watchdog soft patterns
User-guide: new Settings > Analyzer tab (custom analyzer services with default/disabled
toggles and {packetHash} placeholder), apply-path upload button in DM Path Management,
Backup modal's Optimize button + live size label, console change_path now accepts
arrow/whitespace separators with consistent multi-byte chunk length and "path" output
shows hop count and byte size.

Architecture: new /api/analyzers CRUD + default endpoints, /api/db/size and the split
/api/db/vacuum kickoff + /api/db/vacuum/status polling (worker-thread VACUUM to survive
proxy idle timeouts), /api/contacts/<key>/paths/<id>/apply, /health and /health/strict
top-level routes, analyzers table and direct_messages.delivery_path_hash_size column,
recombined path_len byte storage. DeviceManager: per-send channel-secret refresh,
liveness telemetry (_last_rx_at + _consecutive_stats_failures), TCP self-heal via
_liveness_watcher_loop + in-place reconnect. Retention scheduler: on-by-default
90/90/60/30, post-cleanup VACUUM at >=1000 deletions, app-context wrapping, archiver
emoji-name fallback. Socket.IO clients forced to polling transport.

Watchdog: documented hard- vs soft-pattern detection (5 hits in 2 min for sluggish
get_stats / get_battery failures), pointer to /health/strict, and the systemd-restart
deploy note for scripts/watchdog/ changes.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-08 11:53:41 +02:00
MarekWo 7e9ff2e3aa fix(channels): drop hardcoded 0-7 limit on set-scope endpoint
The /api/channels/<idx>/scope route rejected idx>=8 with "Channel index
out of range (0-7)" even though current firmwares expose up to 40 slots
(send_device_query reports max_channels=40 in our logs). Users couldn't
set a region scope on channels like #ubot (idx 8) or #swietokrzyskie
(idx 15) — the UI showed the modal but Save returned 400.

Use device_manager._max_channels (set from send_device_query at connect)
as the upper bound, with 8 as a safe fallback if the DM is unreachable.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-08 11:06:30 +02:00
MarekWo fef6845c03 fix(connection): self-heal degraded long-lived TCP via in-place reconnect
Long-lived TCP against the meshcore-proxy can degrade in a way the socket
can't see: some commands (set_flood_scope_key with all-zero key) start
timing out while RX events and other commands keep working. The 5 s
execute() timeout fires with concurrent.futures.TimeoutError() — whose
str() is empty — so the UI showed "Could not set region scope (none):"
with no error text, and only channels with a mapped region could send
because their non-zero scope_key happened to keep working.

Two recovery paths:

- send_channel_message now detects the timeout case (set_flood_scope_key
  surfaces timed_out=True) and runs force_reconnect() + one retry before
  failing. The user sees a brief delay instead of a cryptic error and
  having to restart the container.

- A new _liveness_watcher_loop task runs on the DM event loop and forces
  a reconnect when no RX event has arrived for HEALTH_STRICT_MAX_RX_STALE_SEC
  (5 min). /health/strict now also reports rx_stale for TCP (previously
  serial/USB only), so an external watchdog could act on it too.

force_reconnect() runs on the DM loop via run_coroutine_threadsafe with
a 20 s cap, a 30 s cooldown to avoid churn under fire, and a
_reconnect_lock to prevent concurrent attempts. mc.disconnect() fires
DISCONNECTED — _intentional_disconnect tells _on_disconnected to skip
its own reconnect loop so the two don't race.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-07 21:10:03 +02:00
MarekWo f1477d84ac fix(db): run VACUUM in a worker thread to survive proxy timeouts
The reverse proxy fronting mc.wojtaszek.it closes idle HTTP responses
after ~30 s, so the manual Optimize endpoint timed out client-side
even though SQLite finished VACUUM and the Flask handler logged a 200.
The user saw "Optimize failed" while the DB had actually shrunk.

Split the endpoint into kickoff + polling: POST /api/db/vacuum spawns
a daemon worker thread, stores state in a module-level dict guarded by
a lock, and returns 202 immediately. GET /api/db/vacuum/status returns
{running, elapsed_seconds, ...} so the UI can poll every 2 s and show
the same "freed X bytes in Y s" toast once the worker is done. A
second POST while a VACUUM is in flight returns 409 instead of starting
a parallel rewrite.

Client polls for up to 10 minutes (300 × 2 s) before surrendering with
a "still running" warning — well past any real VACUUM duration we'd
expect, but bounded so a server-side crash can't leave the UI
spinning forever.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-07 12:30:09 +02:00
MarekWo 13a650bb6c fix(channels): read channels from DB instead of iterating device slots
The TimeoutError-based fallback added in 1d47c9c only fires when
mc.commands.get_channel() actually raises — but on a sluggish device the
call returns an empty/falsy event without raising, so the loop walks
all dm._max_channels slots (40 on the firmware in production), each
empty result returns None, and the API yields just Public (or whatever
slot 0 happened to succeed on). The DB fallback never triggered and the
user kept seeing just Public after refresh.

The channels table in the DB is already the authoritative cache:
- _load_channel_secrets() syncs it on every device connect and prunes
  stale rows,
- set_channel()/remove_channel() update it synchronously with the
  device,
- _refresh_channel_secret() refreshes individual rows on per-send
  refresh.

Drop the device-slot iteration in cli.get_channels() and read from the
DB. /api/channels response time becomes a single SELECT (<1 ms) and is
unaffected by device responsiveness — exactly what we wanted from the
fallback in the first place.

Also revert the TimeoutError re-raise in get_channel_info(): the
console `channels` and `add_channel` commands iterate slots and would
crash on the first slow one. Logging + None on failure is the right
behavior for slot iteration. The 3 s default timeout stays since it
still keeps individual slot probes cheap.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-07 12:21:34 +02:00
MarekWo f72f6d418a feat(db): VACUUM after retention and an Optimize button in Backup modal
SQLite DELETE marks pages free but doesn't shrink the file, so the
new retention job would keep DBs at their bloated size forever without
a follow-up VACUUM. Add db.vacuum() that runs PRAGMA-free VACUUM and
reports size_before/size_after/elapsed so callers can surface results.

The retention job now calls vacuum() automatically when it deleted at
least 1000 rows. Threshold avoids the multi-second VACUUM cost on quiet
days. Failure is logged, not raised — a missed VACUUM never crashes
the scheduler.

Power-user override: new "Optimize now" button in the Database Backup
modal triggers VACUUM on demand via POST /api/db/vacuum, alongside a
GET /api/db/size that drives the live "Current size" label. This way
users don't have to wait until 03:30 to reclaim space after the first
big retention pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-07 10:55:23 +02:00
MarekWo 422e7a3b34 feat(watchdog): catch sluggish-device failures via soft-pattern counting
The container watchdog only restarted on three legacy "device clearly dead"
log lines, so today's failure mode (firmware briefly stalls and get_stats_*
/ get_battery commands time out with an empty error while passive RX
keeps working) never tripped it — leaving the user with 10-15 s freezes
several times a day and no automatic recovery.

DeviceManager now tracks two liveness signals:
- _last_rx_at, bumped on every RX_LOG_DATA event
- _consecutive_stats_failures, incremented on get_stats_* / get_bat
  exceptions and cleared on success

New /health/strict endpoint exposes these to the watchdog. It returns 503
when the device is connected but has 5+ consecutive stats failures, or
when no RX event has been seen for over 5 minutes on a serial transport.
The cheap /health endpoint keeps its lenient behavior so Docker's
healthcheck doesn't suddenly start tripping.

The watchdog's check_device_unresponsive() gains a "soft" pattern class
with a count threshold of 5 in the last 2 minutes — matching against
get_stats_core/radio/packets failed:, Failed to get battery:, and
Failed to get channel. Hard patterns still trigger on a single hit.

Deploy note: the watchdog runs as a host-level systemd service and is
NOT restarted by mcupdate, so after deploy run:
  sudo systemctl restart mc-webui-watchdog.service

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-07 09:43:43 +02:00
MarekWo 63204fe08d fix(retention): enable retention by default and unblock the scheduler
The retention/cleanup scheduler was throwing "Working outside of application
context" on every boot because init_*_schedule() and the APScheduler jobs
themselves call api.get_*_settings(), which reads current_app.db. Pass the
Flask app into archiver.manager via set_flask_app() and decorate every
scheduled job with @_with_app_context so the context is active when the
worker thread runs. Wrap the init calls in main.py in app.app_context() too.

Extend cleanup_old_messages() to also delete from echoes, paths and acks
(the diagnostic tables — 191k echo rows account for the bulk of the 85 MB
DB). Each diagnostic table has its own retention window, defaulting to
min(days, 30) so debug data is purged on a tighter schedule than
user-visible messages.

Switch RETENTION_DEFAULTS to enabled=True with 90/90/60/30 days for
channel_messages/DMs/adverts/diagnostics; user opted in explicitly. First
run scheduled for 03:30 local. SQLite DELETE only marks pages free —
file size won't shrink until a manual VACUUM is run separately.

Archive job stopped working the day the device was renamed to include an
emoji ("MarWoj 💡"): the meshcore library strips non-ASCII when writing
the .msgs file, but our archiver still built /data/{device_name}.msgs and
hit ENOENT every night at 00:00. Add a tolerant fallback that globs the
data dir for a single non-archive .msgs file when the expected path is
missing — mirroring how migrate_v1 already handles this.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-07 09:41:07 +02:00