First Boot After OTA Is Still a Transaction
A successful flash write does not prove the new image can initialize hardware, load state and stay healthy.
TOPIC
Lessons, explainers, experiments, and implementation notes.
A successful flash write does not prove the new image can initialize hardware, load state and stay healthy.
A release could look administratively complete before any device proved it was actually running.
Recovery changes how safely a fleet can ship updates and how clearly operators can diagnose failures.
Operator intent and device-reported reality can disagree legitimately during rollout, reboot, outage or rollback.
Hardware and software eligibility rules should be machine-readable where rollout decisions are made.
Some changes alter the substrate that makes normal A/B OTA safe and therefore cannot be treated like another application image.
Enrollment establishes trusted identity; heartbeat reports what an already known device is doing.
A device sending telemetry could look alive even when its identity or enrollment state was not valid for production operations.
A terminal-looking last_ota_result could survive a later reassignment and incorrectly poison or complete the new assignment.
A previous binary is not a valid rollback target if the new firmware transformed NVS into a format the old firmware cannot read.
A temporary network, DNS, backend or PBX outage could otherwise make a healthy image look defective during first boot.
A successful happy-path download could not prove rollback, credential rejection, compatibility gates or interrupted writes behaved safely.
Heartbeat snapshots alone could not reconstruct the sequence of download, validation, rollback and operator actions during a failed rollout.
A release needed gates between registration and fleet-wide use rather than one published flag.
A reboot into the new partition was too weak to count as a successful update.
The server needed to know what a device should run without pretending the device had already installed it.
A correct hash could prove that downloaded bytes matched the manifest but not that an authorized release process created that manifest and artifact.
If the OTA server stored the production signing private key, compromise of the delivery plane could become authority to mint trusted firmware.
Version text alone could not uniquely identify artifact lineage or distinguish reissued builds.
Rolling all boards at once would maximize blast radius before the first device produced field evidence.
A board needed a stable fleet identity without turning a public hardware identifier into an authentication secret.
Factory-programmed boards needed a controlled transition into an enrolled state without shipping a reusable fleet credential.
Simply placing a newer firmware file on the server risked turning storage into rollout policy.
Power can disappear during download, inactive-slot write, after boot-partition selection or during first boot.
A firmware update competing with a live voice call could damage the product function the OTA system exists to maintain.
An update should be able to fail during download or write without destroying the last working firmware.
A device can change state after assignment, so one compatibility decision made earlier may become stale before download.
Operators needed to stop a device temporarily without confusing that action with permanent credential invalidation.
One rollback mechanism could not cover both immediate boot failure and defects discovered after a release had already been accepted.
Provisioning and OTA control were new dependencies, but a temporary backend outage should not break an already configured voice device.