Production Firmware: Golden Artifacts & Identity · deep-dive
A Known-Good Binary Is a Release Asset, Not a Temporary Build
The clearest audio build could have disappeared under the next firmware experiment if it remained only a file in a working build directory.
A Known-Good Binary Is a Release Asset, Not a Temporary Build
I used to think a firmware release was the binary produced after a successful build. The clearest audio build could have disappeared under the next firmware experiment if it remained only a file in a working build directory.
The LOUP firmware path had a particularly unforgiving constraint: SIP/audio behavior already worked well enough to be valuable, so release work could not casually erase NVS, rewrite the flash layout, merge unrelated features or replace the rollback image. The goal was to make delivery safer without sacrificing the only known-good behavior.
For this article, the key evidence is specific: The V132A handoff treated the exact phone-read ESP32-S3 application image as an authoritative rollback artifact, identified by size, SHA-256, app metadata, flash offset and image validation rather than by the V132A nickname alone. The retained result was equally specific: The working binary became immutable rollback evidence while source recovery and later feature work continued independently. I treat both as observations from the documented release state, not as universal ESP32 rules.
Case notebook
| Question | Recorded conclusion |
|---|---|
| Release problem | The clearest audio build could have disappeared under the next firmware experiment if it remained only a file in a working build directory. |
| Evidence | The V132A handoff treated the exact phone-read ESP32-S3 application image as an authoritative rollback artifact, identified by size, SHA-256, app metadata, flash offset and image validation rather than by the V132A nickname alone. |
| Mechanism | A firmware release has at least three identities: the executable bytes, the source state intended to produce them, and the hardware/test context that proved the behavior. Those identities can diverge. |
| Rejected shortcut | Assuming that a version label or branch name is enough to recover the behavior later. |
| Retained result | The working binary became immutable rollback evidence while source recovery and later feature work continued independently. |
| Rule carried forward | Preserve the behavior first; improve reproducibility without overwriting the only artifact already proven on hardware. |
The table is intentionally stricter than a normal release note. A release note usually tells a reader what changed. This notebook also records what was not proven and which shortcut would have produced a misleading green status. That distinction mattered repeatedly in the LOUP work, especially while the golden V132A artifact, reconstructed source tree, V133A feature branch and later OTA control plane existed at different maturity levels.
Artifact identity is a multi-key problem
For a production firmware artifact I want enough information to distinguish four questions that are often collapsed: What bytes are on the device? What source state was intended to produce them? Which toolchain/dependencies produced the candidate? Which physical test proved the behavior? A binary SHA answers only the first. A Git commit answers only part of the second. A successful call answers the last but does not reconstruct source provenance.
That is why the golden-image workflow preserved exact application bytes separately from source cleanup. The golden artifact could remain the rollback authority even while the repository was being repaired. This is not an admission that reproducibility does not matter; it is a way to avoid destroying the only trusted executable while reproducibility is being established.
In a larger release system I would store this relationship explicitly rather than in filenames: artifact digest, source commit, build environment identity, hardware compatibility, signer key ID and acceptance record. The operational rule is that none of those fields silently substitutes for another.
A concrete scenario I use to test the rule
Imagine I have two files both named V132A.bin. One came from the proven phone, the other from a later rebuild. If I cannot answer which hash belongs to the real call test, rollback has already become ambiguous. My acceptance record therefore points to the digest, not the basename.
For this article, the scenario is useful because it targets the rejected shortcut directly: Assuming that a version label or branch name is enough to recover the behavior later.. I want the system to make that shortcut either impossible or obviously non-compliant with the release gate.
The falsification question is equally important. If a future implementation can demonstrate the same safety property with a simpler mechanism, I would change the mechanism. What I would not change casually is the invariant: Preserve the behavior first; improve reproducibility without overwriting the only artifact already proven on hardware.. The release process exists to preserve that invariant while the implementation evolves.
The mechanism underneath the release decision
The first release problem was preservation. A voice build that users trusted was more valuable than a prettier repository if the repository could not yet prove it produced the same behavior. I therefore separate the golden binary, source provenance, hardware context and test evidence. That lets source recovery move forward without rewriting the operational truth that already exists on the device.
For A Known-Good Binary Is a Release Asset, Not a Temporary Build, the important mechanism is this: A firmware release has at least three identities: the executable bytes, the source state intended to produce them, and the hardware/test context that proved the behavior. Those identities can diverge.
That mechanism tells me which evidence is relevant. If the question is artifact identity, a call test alone is not enough; I need a hash and source identity. If the question is rollback, a signature alone is not enough; I need partition and persistent-schema compatibility. If the question is promotion, a CI pass is not enough; I need observed device state tied to the same release ID.
release_identity = {
binary_sha256,
git_commit,
hardware_revision,
flash_offset,
toolchain_identity,
acceptance_evidence
}
A version string is metadata inside this record, not the primary key for truth.
I use this model to stop release engineering from becoming a sequence of shell commands. The commands are implementation. The release contract is the set of invariants that must still be true when the commands finish.
The controls I would require before accepting this state
- read back and hash the known-good application image
- record flash offset and image metadata
- pin the source commit separately from the binary hash
- keep golden tags immutable
- store acceptance notes beside the artifact
The point is not to maximize checklist length. Each control closes a different ambiguity that appeared in the real work. For this case, the shortcut I reject is Assuming that a version label or branch name is enough to recover the behavior later.. If that shortcut is allowed, the release can look successful while the underlying recovery or provenance guarantee is false.
I prefer a release gate that fails loudly and leaves the old artifact usable. That is why full-flash erasure, force-pushing golden tags, disabling certificate verification, or widening a rollout to compensate for unclear state are all wrong directions. They destroy evidence or increase blast radius exactly when uncertainty is highest.
How I would try to break this before trusting it
Release safety is difficult to prove with only the happy path. For this class of change I want at least one test that intentionally violates the assumption the release depends on.
For A Known-Good Binary Is a Release Asset, Not a Temporary Build, I would construct a negative test around the mechanism: A firmware release has at least three identities: the executable bytes, the source state intended to produce them, and the hardware/test context that proved the behavior. Those identities can diverge. That might mean removing a required source file from the clean checkout, presenting a wrong signature, assigning a release to the wrong hardware revision, forcing first-boot self-test failure, rolling back after a schema migration, or replaying a terminal OTA result for an older release ID.
The expected behavior should be boring: reject the artifact or assignment, keep/restore the previous accepted image, preserve persistent state where promised, and surface an attributable failure state. A test is especially valuable when it proves the system does not accept a dangerous shortcut.
This is why I distinguish recoverability tests from build tests. A compiler can prove syntax and linking. It cannot prove that a power interruption during slot write, a bad first boot or a stale heartbeat result leaves the fleet in a state the operator can understand.
Keep chronology honest
The V132A/V133A recovery documents are useful because they did not retroactively turn incomplete work into completed work. At the August handoff, the golden V132A rollback binary was proven, while the V133A feature branch had source work that still needed clean-build, flash, physical power and clear-audio regression evidence. Later fleet/OTA work added a different layer of production acceptance.
That chronology matters for this topic because The working binary became immutable rollback evidence while source recovery and later feature work continued independently. should be read as the result of the evidence chain that actually existed at that stage. I do not use a later production capability to rewrite an earlier handoff as if it had already passed.
This is also how I want release dashboards and articles to behave. A state such as built, signed, assigned, booted, accepted, active or stable should mean one thing and be backed by the evidence required for that state. The system becomes hard to operate when success words float free of their acceptance gates.
What this costs
The stricter release model adds work. Hashes, manifests, detached signatures, isolated builds, compatibility metadata, dual slots, schema versions and promotion gates all create operational surface area. On a small product team that overhead can feel disproportionate to one ESP32-S3 binary.
The alternative cost is hidden. Without these controls, a good audio artifact can be overwritten, a feature branch can silently become the new baseline, an old image can be unable to read migrated NVS, a runtime server compromise can become signing compromise, or the fleet can report “failed” for the wrong release because one stale result had no causal identity.
For this case the retained rule is Preserve the behavior first; improve reproducibility without overwriting the only artifact already proven on hardware.. I accept the additional release machinery when it closes a failure mode that would otherwise require physical recovery or make the operator unable to state which firmware is really running.
The lesson I keep
The result I keep from this case is: The working binary became immutable rollback evidence while source recovery and later feature work continued independently.
The deeper lesson is Preserve the behavior first; improve reproducibility without overwriting the only artifact already proven on hardware.
That is how I now define production firmware work. The release is not the moment a .bin file appears. It is the chain that connects source, artifact, signature, compatibility, flash topology, persistent state, real-device acceptance and observable running state—with a rollback path whose assumptions have actually been tested.