Both commands already existed in commands/device.py but had no entry in
the README's command reference table, so users had no way to discover
them (reported as "missing"). Add rows under Advanced Configuration
explaining that mode controls how many bytes of each hop's node ID are
stored per hop in advertised/logged paths (bytes per hop = mode + 1),
matching the plen >> 6 / hash_mode + 1 logic in reader.py and
commands/base.py.
Fixes#84
parsePacketPayload copies message/msg_hash/sender_timestamp/attempt/txt_type from a
prior channels_log entry matched by pkt_hash. When that prior entry was logged while
the channel could not be decrypted it has none of those keys, so a later duplicate of
the same transmission raises KeyError: 'message' and aborts handle_rx for the whole
packet, dropping its RX_LOG_DATA. It triggers when a channel-table refresh lands
between two copies of a flooded channel message.
Copy with .get() so the fields are simply absent instead of raising. Adds a regression
test that reproduces the crash (fails on main, passes here).
Follow-up to review of the two preceding commits.
The write lock added in 00135bb prevented concurrent writes from dropping the
link, but bounded nothing. CommandHandler's own timeout does not cover the
write: send() awaits _sender_func() and only afterwards arms
asyncio.wait(futures, timeout=...). So a stalled write -- observed on hardware
running to minutes -- held the lock indefinitely while every other command
queued behind it, with nothing logged, no error raised and no DISCONNECTED
event. Before the lock only the stalled command hung; the others went out and
could trip the error-19 disconnect, which at least recovered. The lock turned
a bounded, self-healing failure into an unbounded silent one, and also blocked
the post-reconnect CMD_APP_START behind the dead connection's holder.
The write (lock acquisition included) is now bounded by
BLEConnection.WRITE_TIMEOUT. On expiry the link is torn down rather than the
lock merely released: the underlying CoreBluetooth write may still be in
flight, and a second write racing it re-creates the overlap the lock exists to
prevent. Tearing down hands over to the reconnect path, which is bounded and
self-healing.
The lock is now a lazily-created property, mirroring _mesh_request_lock in
commands/base.py, so an instance built without __init__ still works. The two
BLE tests previously assigned _write_lock themselves, which meant deleting the
__init__ line left them green; there is now a test that __init__ provides it.
Also in send_anon_req:
- A failed change_contact_path() no longer proceeds. The device still has the
contact as flood, so sendAnonReq() floods the request, and the server gates
REGIONS/OWNER/BASIC behind isRouteDirect() and drops it -- the caller then
waits out a full path-scaled timeout for a reply that cannot arrive. It now
returns ERROR path_reset_failed.
- encode_reply_path() clamps to the server's 64-byte reply_path buffer, which
MyMesh.cpp memcpys into with no length check, and rejects hash mode 3 (the
4-byte hops Packet::isValidPathLen refuses). Not a regression -- the previous
encoder overflowed identically -- but this function is the chokepoint and its
comment claimed to bound the read.
- The hop count saturates at 63 instead of being masked with & 63, which would
wrap a 64-hop path to zero hops, i.e. request a zero-hop reply from a distant
node. Not reachable from a device-sourced contact (the reader caps the field
at 63) but silent if it ever were.
- The suggested_timeout multiplier now scales by the hops actually emitted
rather than the contact's claimed out_path_len, which can differ once the
encoder clamps or truncates.
Correction to 30446ed's message: the claim that mode 0 is "byte-for-byte
identical" is wrong. Differentially, over 30000 randomised contact fields
restricted to what a device can actually emit, mode 0 diverges in 1177 of
10118 cases -- every one of them a path containing a 0x00 byte, and in every
one the old encoder was the wrong one. The accurate claim is "unchanged for
mode-0 paths containing no 0x00 byte".
An anon request tells the server how to route its answer back. The leading
byte of that reply path packs two fields, which the server unpacks as:
reply_path_len = byte & 63
reply_path_hash_size = (byte >> 6) + 1
Three defects in producing it:
1. The hash mode was never written into the top two bits, so the server always
read a hash size of 1 whatever the contact's real mode was.
2. The path was reversed byte-wise (out_path[::-1]) rather than hop-wise. A
return path visits the same hops in reverse order with each hop's
multi-byte hash intact.
3. reader.py built out_path by stripping every NUL from the fixed 64-byte
field. That trims the padding but also eats a legitimate 0x00 inside a hop
hash, shortening the path and shifting every hop after it. It now takes
out_path_len * hash_size bytes, as PATH_DISCOVERY_RESPONSE already did.
Worked example at hash mode 2 (3 bytes per hop), for a contact two hops away
via aabbcc then ddeeff:
before: lenbyte 0x02, path ffeeddccbbaa
-> server reads 2 hops of 1 byte, replies via ['ff', 'ee']
after: lenbyte 0x82, path ddeeffaabbcc
-> server reads 2 hops of 3 bytes, replies via ['ddeeff', 'aabbcc']
The old form routes the response to hops that do not exist, so it is dropped
and the request times out.
At hash mode 0 both encodings are byte-identical -- the mode contributes
nothing to the high bits and byte-wise reversal equals hop-wise reversal for
single-byte hops -- which is why this stayed latent: mode 0 is the default.
Confirmed by the mode-0 and zero-hop tests passing unchanged against the old
code while the mode-1/mode-2 tests fail.
Scope: only anon requests routed direct to a contact with a known multi-hop
path. Flood requests are unaffected (the server answers via createPathReturn
and ignores the supplied reply path), as is login (handleLoginReq never sets
reply_path_len, so its reply always goes out flood). The neighbors zero-hop
probe is unaffected: length 0 makes hash size irrelevant.
Encoding is extracted into encode_reply_path() so it can be tested directly.
Verified on hardware only for the zero-hop case, which still works; the
multi-hop paths are covered by unit tests, as the test radio has no multi-hop
contacts to exercise on air.
Two overlapping write_gatt_char() calls on the same characteristic drop the
BLE link outright. Observed on macOS/CoreBluetooth as "BLE write failed: 19",
after which the connection is gone and the pending command never completes.
Nothing above the transport guaranteed callers were sequential: schedulers,
health checks, periodic status queries and user commands all issue commands
independently, so any unlucky overlap could take the radio down. The existing
_mesh_request_lock only guards a few binary-request helpers, not the transport.
Reproduced on a companion radio over BLE by issuing send_device_query() and
send_node_discover_req() concurrently:
before: BLE write failed: 19, connected=False, command hung >45s
after: both complete in 0.12s, connected=True
Issued sequentially the same two commands take 0.09s each and are fine, so it
is specifically the overlap. Ruled out as causes beforehand: notification load
(three discovers under a 91-packet firehose kept writes at 0.06-0.17s with the
link stable) and the discover command itself.
The lock is created lazily so it binds to the running loop, and is released on
the failure path so one failed write cannot wedge every later command.
send_anon_req() refused to send whenever the destination pubkey was absent
from the client-side contact cache, returning ERROR contact_not_found.
The contact is consulted for one thing only: building the reply-path bytes
appended to the request. The companion firmware needs no contact of its own --
since FIRMWARE_VER_CODE 13 its CMD_SEND_ANON_REQ handler synthesises a
transient anon contact for an unknown pubkey with out_path_len = 0 (zero-hop
direct). Those entries live in a reserved slot ring, are hidden from
CMD_GET_CONTACTS and are never persisted, so nothing is polluted by them.
The client-side refusal therefore blocked a case the device supports, such as
asking a freshly discovered neighbour for its regions before it has ever been
added as a contact. Fall back to a zero-hop reply path instead.
Two smaller fixes in the same function:
- out_path_len is now read once into a local rather than re-read from the
contact dict after the await. That dict is a live reference other commands
mutate in place (send_msg_with_retry's flood fallback, reset_path); if it
flipped to -1 mid-send the suggested_timeout multiplier became 4000 * 0 = 0,
registering the binary request with a zero timeout so the response was
dropped the moment it arrived.
- The value is clamped at 0. update_contact() normally reflects the change
back onto the dict, but if it fails the dict stays -1 and the unsigned
to_bytes raises OverflowError -- which skipped the reset_path at the end of
the method and left the contact pinned to zero-hop on the device.
Verified against a companion radio on fw ver 13: without the change all five
discovered repeaters were refused client-side in 0.0s with no RF sent; with it
all three answered with their region scopes in ~1.1s.
test_send_anon_req_contact_not_found is replaced -- it codified the removed
limitation -- but its original regression (a TypeError on the NoneType
subscript) stays covered.
Why: PR #86 (44b21be) reworked send_raw_data to frame as
0x19 | path_len(1) | path | payload and to reject payloads under 4
bytes, but left test_send_raw_data_wrapper unchanged. The test sent a
2-byte payload (now raising ValueError before any assertion) and
asserted captured_data[1:] == payload, which no longer holds because of
the inserted path_len byte. The firmware drops payloads under 4 bytes,
so the guard is correct; this fixes the test, not the guard.
Bumps the payload to 4 bytes and asserts the zero path_len byte plus the
payload at offset 2, validating the wire format #86 introduced.
Tests: tests/unit/test_protocol_surface_gaps.py -- full suite 156 passed
(test_send_raw_data_wrapper was the lone failure on current main).
Wire-format parity fixes in reader.py for trailing fields the
companion-radio firmware emits but the SDK decoder was dropping, plus an
over-read guard on DEFAULT_FLOOD_SCOPE. Each fix uses the existing
BATT_AND_STORAGE defensive-read pattern (up-front minimum-length check
plus per-field cumulative-length gates), so older firmware that doesn't
emit the trailing fields keeps decoding without raising. Every fix ships
a legacy-frame and modern-frame unit test pair.
AUTOADD_CONFIG (PacketType 25): companion-v1.14.0 firmware emits a
trailing max_hops byte (firmware commit 00566741); the SDK dropped it.
Adds a defensive `if len(data) >= 3` read. Pre-v1.14.0 frames decode to
`{"config": ...}` with no max_hops key.
LOGIN_SUCCESS (PacketType 0x85): the new-style RESP_SERVER_LOGIN_OK path
emits three trailing fields the SDK was dropping — server_timestamp (4B,
firmware commit 0e90b731), acl_permissions (1B, 7947e8a2), and
fw_ver_level (1B, 418ae08b), all first shipped in companion-v1.10.0.
Adds per-field length gates (>=12, >=13, >=14). Legacy 8-byte "OK"-path
frames decode unchanged.
ACK (PacketType 0x82): firmware emits a trailing 4-byte trip_time
(round-trip latency in ms) since companion-v1.0.0a (firmware commit
d9dc76f1, Jan 2025) — on the wire ~16 months but never surfaced by the
SDK. Adds an `if len(data) >= 9` read. Legacy 5-byte ACK frames decode
unchanged.
DEFAULT_FLOOD_SCOPE (PacketType 28): firmware emits a 48-byte populated
frame or a 1-byte sentinel when no scope is set. The SDK unconditionally
read 31+16 bytes, over-reading 47 bytes past the end of the sentinel
frame and dispatching `{"scope_name": "", "scope_key": ""}`. Adds an
`if len(data) >= 48` guard so the sentinel dispatches `{}`. Consumers
detecting "no scope" via `payload["scope_name"] == ""` should switch to
a key-presence check.
RAW_DATA: the payload-framing fix this branch originally carried landed
upstream first in PR #86 (commit 44b21be), with identical decode logic
(discard the reserved 0xFF byte, then read the remaining variable-length
payload). That redundant change is dropped here; this commit retains only
its regression test (test_raw_data_realistic_frame), which exercises the
upstream fix end-to-end — the upstream change shipped without a test.
CHANGELOG:
- Decode max_hops trailing byte in AUTOADD_CONFIG (companion-v1.14.0+).
- Decode server_timestamp / acl_permissions / fw_ver_level trailing
fields in LOGIN_SUCCESS (companion-v1.10.0+).
- Decode trip_time trailing field in ACK (companion-v1.0.0a+).
- DEFAULT_FLOOD_SCOPE dispatches `{}` on the 1-byte sentinel frame
instead of `{scope_name: "", scope_key: ""}`. Detect "no scope" via
key presence.
Tests: 10 tests in tests/unit/test_protocol_surface_gaps.py (legacy +
modern frame pair per finding, including a RAW_DATA regression test for
the upstream fix); all 10 pass on the rebased upstream base.
Why: companion-radio firmware has been emitting these fields on the wire
for between 2 and 16 months; SDK consumers (integrations, bots,
dashboards) lose access to data that is on the wire today.
Why: companion-radio firmware emits RESP_CODE_CHANNEL_DATA_RECV (27) for
group-channel binary data (PAYLOAD_TYPE_GRP_DATA), shipped in
companion-v1.15.0. The SDK had no PacketType value 27, no EventType, and
no reader handler, so these frames hit the unknown-packet-type
fallthrough and the payload was silently dropped.
This adds PacketType.CHANNEL_DATA_RECV = 27 (the previously-skipped enum
slot), EventType.CHANNEL_DATA_RECV, and a reader handler. The fixed
9-byte header (snr, reserved, channel_idx, path_len with the
sentinel/hash-mode encoding) reuses CHANNEL_MSG_RECV_V3's framing; the
typed tail decodes data_type (uint16 little-endian, widened from uint8
in firmware), data_len, and payload (hex string). An up-front length
gate mirrors the defensive-read pattern used by the other handlers.
data_type is exposed as an int (matching the txt_type convention) and
payload as a hex string (matching RAW_DATA's convention for binary data
of unknown encoding); attributes surface channel_idx and data_type so
subscribers can filter without unpacking the payload.
Tests: tests/unit/test_protocol_surface_gaps.py — 5 new tests (enum
slot present, direct-path frame, route-flood path_len bit-split,
under-minimum frame dropped, widened data_type round-trip). Full unit
suite: 146 passed.
Rename test_g6_protocol_surface_gaps.py to test_protocol_surface_gaps.py.
Strip G6 from module docstring, and finding IDs (N01, N02, N03, N05,
N09, R04) from docstrings and section comments.
Rename test_g5_asyncio_lifecycle.py to test_asyncio_lifecycle.py.
Strip G5 from module docstring, finding IDs (F05, F07, F08, F19)
from class names and docstrings.
Rename test_g4_transport_symmetry.py to test_transport_symmetry.py.
Strip G4 from module docstring, _g4_ from function names, and finding
IDs (F04, NEW-A, F18, M06, N04) from docstrings and section comments.
Rename test_g2_error_handling.py to test_error_handling.py. Strip G2
prefix from module docstring, _g2_ from function names, and finding
IDs (F22, F21/M01, M02, M04, N06, F14) from docstrings and section
comments. Proposal cross-references removed.
Strip internal forensics finding references (F06, N07, NEW-C, R02)
from docstrings, comments, and assertion messages. Descriptive text
is preserved — only the ID prefixes are removed.
Strip internal forensics finding references (F01, F02, F03, N11)
from docstrings and section comments. The descriptive text is
preserved — only the ID prefixes are removed.
Strip G1/ prefix from docstrings, _g1_ from function names,
and G1 from section comments. Finding IDs (F06, N07, etc.)
are preserved as they are meaningful standalone references.