iuna

iuna

iuna - experimental mainnet-candidate protocol
git clone https://getiuna.org/git/iuna.git
Log | Files | Refs | README | LICENSE

commit 4231d8648b77b407743a186b3cf8bbfed360ba2b
parent 4946004c2617a4c25b767af4a200dde5212f35e6
Author: Joris Hartog <jorishartog@hotmail.com>
Date:   Wed, 19 Aug 2026 15:39:49 +0200

Document operator failure playbooks

Diffstat:
MREADME.md | 5+++++
MROADMAP.md | 4++--
Adocs/operator-playbooks.md | 228+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 235 insertions(+), 2 deletions(-)

diff --git a/README.md b/README.md @@ -119,6 +119,11 @@ iuna can expose a Stratum V1 endpoint for SHA-256 ASIC miners such as a Bitaxe: Use your iuna wallet address as the worker username. Accepted shares become PoW mine actions in the node mempool and are gossiped to peers. +## Operator Docs + +- [Starting the genesis node](docs/genesis.md) +- [Operator failure playbooks](docs/operator-playbooks.md) + ## Contributing We are open to PRs and help running, testing, and improving the devnet. diff --git a/ROADMAP.md b/ROADMAP.md @@ -50,7 +50,7 @@ These items are not protocol rules. They are the attack and reliability checks t - [x] HTTP/auth abuse tests cover CSRF same-origin behavior, lockout/backoff behavior, forwarded-header spoofing from untrusted peers, and session expiry. - [x] Multi-node in-memory simulation covers delayed gossip, withheld burn bundles, bundle equivocation, partitions, restarts, persistence reload, and convergence. - [x] Long-running release-mode soak test runs with automatic burn/finalization, P2P sync, Stratum-disabled and Stratum-enabled nodes, and periodic node restarts. -- [ ] Operator failure playbooks exist for stalled height, divergent tips, old snapshots, no burn committee signatures, recovery blocks, and corrupted local persistence. +- [x] Operator failure playbooks exist for stalled height, divergent tips, old snapshots, no burn committee signatures, recovery blocks, and corrupted local persistence. - [ ] Mainnet-candidate release rehearsal includes fresh genesis, published bootnodes, checksums, backup/restore instructions, and a no-reset stability window. ## Milestones @@ -72,7 +72,7 @@ Focus: make failure modes boring and observable. - [x] Run long-running tests unconditionally during deployment. - [ ] Run a multi-day public testnet without manual chain resets. - [ ] Add or improve operator-facing health metrics. -- [ ] Document common testnet failure/recovery playbooks. +- [x] Document common testnet failure/recovery playbooks. ### M2: Mainnet Candidate diff --git a/docs/operator-playbooks.md b/docs/operator-playbooks.md @@ -0,0 +1,228 @@ +# Operator Failure Playbooks + +These playbooks are for testnet and mainnet-candidate operators. They are intentionally practical: identify the failure, preserve useful evidence, recover the node, and avoid making a temporary network issue worse. + +Assume the management UI is bound to `127.0.0.1` and the public P2P listener is the only internet-facing service. The default data directory is `~/.iuna`; the default chain database is `~/.iuna/chain.sqlite3`; the default config file is `~/.iuna/config.json`. + +## First Checks + +Before changing state, capture the local view: + +```sh +curl -s http://127.0.0.1:18661/api/status +curl -s http://127.0.0.1:18661/api/network/health +curl -s http://127.0.0.1:18661/api/peers +curl -s http://127.0.0.1:18661/api/p2p/metrics +``` + +These API routes require the same authenticated management session as the UI. If `curl` returns an auth response, collect the same data from the local UI or include the browser session cookie in the request. + +Record: + +- local height and tip hash; +- best known peer height; +- peer count, stale peers, banned peers, and last peer error; +- current leader and `last_auto_finalization_status`; +- whether the wallet is locked; +- P2P listener, configured announce address, and whether inbound P2P is active. + +If the node runs with non-default paths, get them from startup logs or flags: + +```sh +iuna --help +``` + +The process prints the wallet file, config file, chain database, management UI, and P2P listener on startup. + +## Stalled Height + +Symptoms: + +- local height does not increase for longer than the target block time; +- `/api/network/health` reports `isolated`, `stale`, `peer errors`, or `syncing`; +- `last_auto_finalization_status` repeatedly says the wallet is locked, automatic mining is off, collecting burns, waiting for another finalizer, or VDF work is running too long. + +Checks: + +1. Compare local height with peers in `/api/network/health`. +2. Check `/api/peers` for recent successful contact and peer heights. +3. Check `/api/status` for wallet lock state, current leader, automatic mining, burn amount, VDF rounds, and recovery VDF threshold. +4. Check host clock synchronization. Blocks too far in the future are rejected and peer clock warnings appear in network health. + +Recovery: + +- If the node has no healthy peers, add a known bootnode in the Peers screen or restart with `--join <addr:port>` only when there is no local chain database. +- If the wallet is locked and this node is expected to finalize, unlock it in the UI. +- If automatic mining is off and this node should participate, enable Mine settings and set a fee-paying burn policy. +- If every healthy peer is at the same height, wait through fallback and recovery windows before resetting anything. A recovery block is the expected liveness escape hatch when ticket finalizers fail. +- If peers are ahead, keep the node online and let range sync catch up. Use local chain reset only if the node remains stuck after it has healthy peers. + +Avoid: + +- deleting the wallet file; +- repeatedly resetting multiple public nodes at once; +- starting a new genesis while the network still has a valid chain. + +## Divergent Tips + +Symptoms: + +- local height matches peers but tip hash differs; +- peers report different tips at the same or nearby heights; +- logs mention fork, snapshot, orphan, or block validation errors. + +Checks: + +1. Compare `/api/status` tip hash and height across at least two healthy peers. +2. Inspect recent blocks in the UI or `/api/blocks`. +3. Check `/api/p2p/metrics` for parse errors, session failures, and data envelope counts. +4. Look for repeated snapshot requests in logs. Nodes request snapshots when a peer appears to be on a better valid fork. + +Recovery: + +- If the divergence is inside the finality window, keep nodes connected. The protocol can reorg to a taller valid fork or to a same-height fork with better leader quality. +- If your node is behind a healthy majority, add direct peers to that majority and let snapshot/range sync resolve it. +- If your node is alone on an old tip and does not converge, use Delete local chain in Settings. That clears local chain/UI data, keeps wallet and settings, and broadcasts a snapshot request to peers. +- If multiple public nodes disagree beyond the finality window, stop automated restarts and preserve chain databases from both sides for analysis. + +Avoid: + +- forcing a manual database replacement from an untrusted peer; +- resetting the apparent majority tip before confirming it validates on more than one node. + +## Old Snapshots Or Long Catch-Up + +Symptoms: + +- a node starts from an old persisted height; +- `/api/network/health` reports `syncing` with peer height ahead; +- startup logs say the node resumed a chain database at an old height. + +Checks: + +1. Confirm the node has at least one healthy outbound or inbound peer. +2. Confirm peers know a higher height. +3. Check `/api/p2p/metrics` for outbound connect failures or queue pressure. +4. Confirm the configured P2P announce address is reachable from outside if the node is meant to be public. + +Recovery: + +- Leave the node online while it requests block ranges and snapshots from peers. +- Add one or two stable peers manually if peer exchange has not found healthy peers. +- Restart the process once if sessions are stale but the chain database loads successfully. +- Use Delete local chain only if the node cannot extend its local snapshot and healthy peers are available. This does not remove wallet or settings. + +Avoid: + +- deleting the chain database before capturing the startup error; +- using `--genesis` to recover an old node. `--genesis` is only for creating a fresh network. + +## No Burn Committee Signatures + +Symptoms: + +- `last_auto_finalization_status` stays on collecting burns or says a block has too few burn bundle signatures; +- there are fee-paying burns in mempools but rank `0` blocks do not finalize; +- fallback blocks start appearing more often than expected. + +Checks: + +1. Check `/api/mempool` for fee-paying burn transactions. +2. Check `/api/peers` for enough healthy peers. Burn bundles are gossip, so isolated nodes see fewer attestations. +3. Check whether committee members are online and unlocked if they are expected to sign bundles. +4. Inspect recent blocks: rank `0` needs the strictest burn-list attestation threshold; fallback ranks relax it for liveness. + +Recovery: + +- Improve peer connectivity first. Add stable peers and keep committee-capable nodes online. +- Make sure automatic mining is enabled and wallets that should finalize/sign are unlocked. +- Wait through fallback windows. Missing committee signatures should not permanently halt the chain; fallback ranks need fewer attestations. +- If the network reaches recovery blocks, treat it as a warning that normal ticket finalization or committee gossip is unhealthy. + +Avoid: + +- treating local mempool contents as proof that a remote block is invalid. Block validation must be independent of each validator's mempool. +- resetting a node merely because it did not personally see a burn bundle. + +## Recovery Blocks + +Symptoms: + +- recent blocks show finalizer mode `Recovery`; +- normal ticket blocks stopped for about the recovery delay; +- `last_auto_finalization_status` mentions recovery VDF. + +Checks: + +1. Confirm whether recovery blocks are accepted by multiple peers. +2. Check if selected ticket finalizers were offline, locked, not burning, or missing committee attestations. +3. Check clock warnings. Large clock skew can cause valid-looking local candidates to be rejected by peers. +4. Check VDF rounds and host performance if VDF work consistently finishes late. + +Recovery: + +- If recovery blocks are converging across peers, do not reset. Recovery is part of the liveness design. +- Bring ticket finalizer nodes back online and unlocked. +- Keep recovery VDF threshold conservative on low-power hosts if too many nodes are wasting work on fallback/recovery paths. +- After recovery, watch the next normal ticket blocks. Persistent recovery means the normal ticket path needs investigation. + +Avoid: + +- banning peers only because they produced a recovery block; +- assuming recovery is an emergency fork. It is valid, weaker-for-fairness liveness behavior. + +## Corrupted Local Persistence + +Symptoms: + +- startup fails with `failed to load chain database`; +- SQLite reports it cannot open or parse `chain.sqlite3`; +- the UI loads but chain views are empty or inconsistent while peers are healthy. + +Checks: + +1. Preserve the files before recovery: + +```sh +cp ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.bak 2>/dev/null || true +cp ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.bak 2>/dev/null || true +``` + +2. Keep `wallet.json` and `config.json`; they are not chain state. +3. Confirm at least one trusted peer is available before deleting local chain state. + +Recovery: + +- Preferred: use Settings, Delete local chain. This clears local chain/UI data, keeps wallet and settings, and asks peers for a snapshot. +- CLI fallback: stop the node, move only the chain/UI data files aside, then restart with the same wallet/config and healthy peers: + +```sh +mv ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.corrupt 2>/dev/null || true +mv ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.corrupt 2>/dev/null || true +iuna +``` + +- If starting from an empty chain database, join a known peer: + +```sh +iuna --join <trusted-peer-addr:port> +``` + +Avoid: + +- deleting `wallet.json`; +- editing SQLite files in place; +- accepting a snapshot from an unknown or isolated peer as the only recovery source. + +## Escalation Bundle + +When a problem needs developer investigation, collect: + +- command line and version; +- startup log lines showing wallet/config/chain paths and listeners; +- `/api/status`, `/api/network/health`, `/api/peers`, and `/api/p2p/metrics`; +- the last 20 blocks from `/api/blocks?limit=20`; +- whether the wallet was locked and whether automatic mining, P2P inbound, and Stratum were enabled; +- copies of corrupted chain/UI databases if persistence failed. + +Do not include recovery phrases, wallet passwords, or private keys in reports.