operator-playbooks.md (15330B)
1 # Operator Failure Playbooks 2 3 These playbooks are for testnet and mainnet-candidate operators. They are intentionally practical: identify the failure, preserve useful evidence, recover the node, and avoid making a temporary network issue worse. 4 5 Assume the management UI is bound to 127.0.0.1 and the public P2P listener is the only internet-facing service. The default data directory is ~/.iuna; the default chain database is ~/.iuna/chain.sqlite3; the default config file is ~/.iuna/config.json. 6 7 ## First Checks 8 9 Before changing state, capture the local view: 10 11 ```sh 12 curl -s http://127.0.0.1:18661/api/status 13 curl -s http://127.0.0.1:18661/api/network/health 14 curl -s http://127.0.0.1:18661/api/peers 15 curl -s http://127.0.0.1:18661/api/p2p/metrics 16 ``` 17 18 These API routes require the same authenticated management session as the UI. If curl returns an auth response, collect the same data from the local UI or include the browser session cookie in the request. 19 20 Record: 21 22 - local height and tip hash; 23 - best known peer height; 24 - peer count, stale peers, banned peers, and last peer error; 25 - current leader and last_auto_finalization_status; 26 - whether the wallet is locked; 27 - P2P listener, configured announce address, and whether inbound P2P is active. 28 29 If the node runs with non-default paths, get them from startup logs or flags: 30 31 ```sh 32 iuna --help 33 ``` 34 35 The process prints the wallet file, config file, chain database, management UI, and P2P listener on startup. 36 37 ## Legacy address-book migration 38 39 The upgraded wallet shows checksummed Bech32m receive addresses: iuna1... for 40 mainnet/mainnet-candidate and tiuna1... for local testnet. Transfer, 41 address-book, and Stratum input no longer accepts a 64-character hex public key. 42 This is intentional: silently accepting hex would bypass typo and 43 wrong-network protection. 44 45 Before genesis or before reusing an old contact, ask the recipient to copy a new 46 address from their upgraded wallet and replace the old address-book entry. 47 Never change iuna to tiuna (or the reverse) by hand; the checksum covers the 48 prefix. Existing wallet files and chain databases retain their internal hex 49 keys and must not be rewritten. Controlled migration tools may call the domain 50 migrate_legacy_address function only when the target network is explicitly 51 known. 52 53 No block-height activation is needed because on-chain and P2P address bytes do 54 not change. Treat this as a client compatibility rollout instead: deploy the 55 wallet UI and backend together, have recipients publish their new display 56 address, and update Stratum worker usernames before restarting miners. 57 58 ## Stalled Height 59 60 Symptoms: 61 62 - local height does not increase for longer than the target block time; 63 - /api/network/health reports isolated, stale, peer errors, or syncing; 64 - last_auto_finalization_status repeatedly says the wallet is locked, automatic mining is off, collecting burns, waiting for another finalizer, or VDF work is running too long. 65 66 Checks: 67 68 1. Compare local height with peers in /api/network/health. 69 2. Check /api/peers for recent successful contact and peer heights. 70 3. Check /api/status for wallet lock state, current leader, automatic mining, burn amount, VDF rounds, and recovery VDF threshold. 71 4. Check host clock synchronization. Blocks too far in the future are rejected and peer clock warnings appear in network health. 72 73 Recovery: 74 75 - If the node has no healthy peers, add a known bootnode in the Peers screen or restart with --join <addr:port> only when there is no local chain database. 76 - If the wallet is locked and this node is expected to finalize, unlock it in the UI. 77 - If automatic mining is off and this node should participate, enable Mine settings and set a fee-paying burn policy. 78 - If every healthy peer is at the same height, wait through fallback and recovery windows before resetting anything. A recovery block is the expected liveness escape hatch when ticket finalizers fail. 79 - If peers are ahead, keep the node online and let range sync catch up. Use local chain reset only if the node remains stuck after it has healthy peers. 80 81 Avoid: 82 83 - deleting the wallet file; 84 - repeatedly resetting multiple public nodes at once; 85 - starting a new genesis while the network still has a valid chain. 86 87 ## Divergent Tips 88 89 Symptoms: 90 91 - local height matches peers but tip hash differs; 92 - peers report different tips at the same or nearby heights; 93 - logs mention fork, snapshot, orphan, or block validation errors. 94 95 Checks: 96 97 1. Compare /api/status tip hash and height across at least two healthy peers. 98 2. Inspect recent blocks in the UI or /api/blocks. 99 3. Check /api/p2p/metrics for parse errors, session failures, and data envelope counts. 100 4. Look for repeated snapshot requests in logs. Nodes request snapshots when a peer appears to be on a better valid fork. 101 102 Recovery: 103 104 - Before height 1000, if the divergence is inside the six-block legacy finality window, keep nodes connected. The protocol can reorg to a taller valid fork or to a same-height fork with better leader quality. 105 - From height 1000, compare finalized_height and finalized_hash in /api/status or /api/network/health. Keep nodes connected when one valid chain has a higher checkpoint: fork choice adopts it even if its tip is temporarily shorter. 106 - If your node is behind a healthy majority, add direct peers to that majority and let snapshot/range sync resolve it. 107 - If your node is alone on an old tip and does not converge, use Delete local chain in Settings. That clears local chain/UI data, keeps wallet and settings, and broadcasts a snapshot request to peers. 108 - If multiple public nodes report conflicting hashes at the same finalized height, preserve both databases and signing evidence immediately. Nodes deterministically choose the lower checkpoint hash, so operators do not need to trust the first peer seen, but the conflict is a quorum safety incident and must be investigated before resuming release activity. 109 110 Avoid: 111 112 - forcing a manual database replacement from an untrusted peer; 113 - resetting the apparent majority tip before confirming it validates on more than one node. 114 115 ## Old Snapshots Or Long Catch-Up 116 117 Symptoms: 118 119 - a node starts from an old persisted height; 120 - /api/network/health reports syncing with peer height ahead; 121 - startup logs say the node resumed a chain database at an old height. 122 123 Checks: 124 125 1. Confirm the node has at least one healthy outbound or inbound peer. 126 2. Confirm peers know a higher height. 127 3. Check /api/p2p/metrics for outbound connect failures or queue pressure. 128 4. Confirm the configured P2P announce address is reachable from outside if the node is meant to be public. 129 130 Recovery: 131 132 - Leave the node online while it requests block ranges and snapshots from peers. 133 - Add one or two stable peers manually if peer exchange has not found healthy peers. 134 - Restart the process once if sessions are stale but the chain database loads successfully. 135 - Use Delete local chain only if the node cannot extend its local snapshot and healthy peers are available. This does not remove wallet or settings. 136 137 Avoid: 138 139 - deleting the chain database before capturing the startup error; 140 - using --genesis to recover an old node. --genesis is only for creating a fresh network. 141 142 ## Legacy Pre-v6 Reset 143 144 The live candidate release writes compact snapshot v7 and accepts snapshot 145 versions v6 and v7. At startup, the node detects the legacy JSON schema and 146 compact snapshot versions older than v6, checkpoints the database, archives it 147 as chain.sqlite3.pre-v6 (or the next available numbered suffix), and creates a 148 fresh database. The UI cache is then cleared normally. Wallet and configuration 149 files are left untouched. 150 151 The coordinated reset that created the live mainnet-candidate genesis is 152 complete. A current operator encountering a pre-v6 archive should join the 153 published candidate bootnode without --genesis; creating another genesis would 154 create a separate, incompatible network. Preserve the generated .pre-v6 155 archive if the old chain is needed as historical evidence. 156 157 Manual equivalent for operators who want to choose the archive names before starting: 158 159 ```sh 160 mv ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.pre-v6 2>/dev/null || true 161 mv ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.pre-v6 2>/dev/null || true 162 iuna --join <trusted-peer-addr:port> 163 ``` 164 165 For the disposable Compose testnet, docker compose down -v removes all volumes and the configured bootstrap creates the fresh genesis on the next start. 166 167 Avoid: 168 169 - running multiple independent --genesis nodes; 170 - using --genesis to join or recover the live mainnet candidate; 171 - copying an old snapshot blob into the current chain database; 172 - deleting or replacing wallets as part of the chain reset; 173 - joining a peer before checking the published network identity and release checksum. 174 175 ## Height 1000 Consensus Activation 176 177 Height 1000 is a coordinated consensus activation. At that height, VDF seeds 178 start committing to block content and ticket draws stop using the final block 179 hash. Burn-committee lineage draws make the same switch, closing the timestamp/ 180 block-hash grinding path for both leader and committee selection. Transaction 181 signatures and native and Stratum mine proofs also switch to 182 chain-bound binary format v1. Rank 0 blocks also start requiring a strict 183 two-thirds committee quorum so their next rank 0 child can objectively certify 184 them. This preserves blocks, snapshots, and UTXOs below 185 1000, but nodes running the earlier rule will reject the upgraded chain or 186 build an incompatible fork at activation. No database reset, new genesis, or 187 migration command is needed for this height activation. 188 189 Required activation procedure: 190 191 1. Publish a tagged release, commit, checksums, and the activation height. 192 2. Upgrade every known public peer, finalizer, and bootstrap node. 193 3. Verify the reported package version on each managed node and compare tips. 194 4. Confirm every independent node agrees on the block hash at height 999; 195 upgraded nodes freeze pre-activation history once they reach 1000. 196 5. Stop or isolate nodes that cannot be upgraded before activation. 197 6. Keep chain database backups from immediately before the activation window. 198 199 At and after height 1000, compare height and tip hash across at least three 200 independent nodes. From height 1001, also compare finalized height and hash. 201 If one chain has the higher valid checkpoint, normal sync should converge to it. 202 If equal-height checkpoint hashes conflict, preserve both histories and the 203 committee signatures, stop automated restarts, and investigate the quorum 204 failure; the deterministic lower-hash rule is the recovery decision and does 205 not by itself make the safety breach harmless. 206 207 ## No Burn Committee Signatures 208 209 Symptoms: 210 211 - last_auto_finalization_status stays on collecting burns or says a block has too few burn bundle signatures; 212 - there are fee-paying burns in mempools but rank 0 blocks do not finalize; 213 - fallback blocks start appearing more often than expected. 214 215 Checks: 216 217 1. Check /api/mempool for fee-paying burn transactions. 218 2. Check /api/peers for enough healthy peers. Burn bundles are gossip; finalizers request missing burn-bundle slots during collection, but isolated nodes still cannot receive responses. 219 3. Check whether committee members are online and unlocked if they are expected to sign bundles. 220 4. Inspect recent blocks: rank 0 needs the strictest burn-list attestation threshold; fallback ranks relax it for liveness. 221 222 Recovery: 223 224 - Improve peer connectivity first. Add stable peers and keep committee-capable nodes online. 225 - Make sure automatic mining is enabled and wallets that should finalize/sign are unlocked. 226 - Wait through fallback windows. Missing committee signatures should not permanently halt the chain; fallback ranks need fewer attestations. 227 - If the network reaches recovery blocks, treat it as a warning that normal ticket finalization or committee gossip is unhealthy. 228 229 Avoid: 230 231 - treating local mempool contents as proof that a remote block is invalid. Block validation must be independent of each validator's mempool. 232 - resetting a node merely because it did not personally see a burn bundle. 233 234 ## Recovery Blocks 235 236 Symptoms: 237 238 - recent blocks show finalizer mode Recovery; 239 - normal ticket blocks stopped for about the recovery delay; 240 - last_auto_finalization_status mentions recovery VDF. 241 242 Checks: 243 244 1. Confirm whether recovery blocks are accepted by multiple peers. 245 2. Check if selected ticket finalizers were offline, locked, not burning, or missing committee attestations. 246 3. Check clock warnings. Large clock skew can cause valid-looking local candidates to be rejected by peers. 247 4. Check VDF rounds and host performance if VDF work consistently finishes late. 248 249 Recovery: 250 251 - If recovery blocks are converging across peers, do not reset. Recovery is part of the liveness design. 252 - Bring ticket finalizer nodes back online and unlocked. 253 - Keep the fallback VDF threshold conservative on low-power hosts if too many nodes are wasting work on lower-ranked ticket paths. Every automatic node remains eligible for recovery once the recovery delay has elapsed. 254 - After recovery, watch the next normal ticket blocks. Persistent recovery means the normal ticket path needs investigation. 255 256 Avoid: 257 258 - banning peers only because they produced a recovery block; 259 - assuming recovery is an emergency fork. It is valid, weaker-for-fairness liveness behavior. 260 261 ## Corrupted Local Persistence 262 263 Symptoms: 264 265 - startup fails with failed to load chain database; 266 - SQLite reports it cannot open or parse chain.sqlite3; 267 - the UI loads but chain views are empty or inconsistent while peers are healthy. 268 269 Checks: 270 271 1. Preserve the files before recovery: 272 273 ```sh 274 cp ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.bak 2>/dev/null || true 275 cp ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.bak 2>/dev/null || true 276 ``` 277 278 2. Keep wallet.json and config.json; they are not chain state. 279 3. Confirm at least one trusted peer is available before deleting local chain state. 280 281 Recovery: 282 283 - Preferred: use Settings, Delete local chain. This clears local chain/UI data, keeps wallet and settings, and asks peers for a snapshot. 284 - CLI fallback: stop the node, move only the chain/UI data files aside, then restart with the same wallet/config and healthy peers: 285 286 ```sh 287 mv ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.corrupt 2>/dev/null || true 288 mv ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.corrupt 2>/dev/null || true 289 iuna 290 ``` 291 292 - If starting from an empty chain database, join a known peer: 293 294 ```sh 295 iuna --join <trusted-peer-addr:port> 296 ``` 297 298 Avoid: 299 300 - deleting wallet.json; 301 - editing SQLite files in place; 302 - accepting a snapshot from an unknown or isolated peer as the only recovery source. 303 304 ## Escalation Bundle 305 306 When a problem needs developer investigation, collect: 307 308 - command line and version; 309 - startup log lines showing wallet/config/chain paths and listeners; 310 - /api/status, /api/network/health, /api/peers, and /api/p2p/metrics; 311 - the last 20 blocks from /api/blocks?limit=20; 312 - whether the wallet was locked and whether automatic mining, P2P inbound, and Stratum were enabled; 313 - copies of corrupted chain/UI databases if persistence failed. 314 315 Do not include recovery phrases, wallet passwords, or private keys in reports.