iuna

iuna

iuna - experimental mainnet-candidate protocol
git clone https://getiuna.org/git/iuna.git
Log | Files | Refs | README | LICENSE

operator-playbooks.md (15330B)


      1 # Operator Failure Playbooks
      2 
      3 These playbooks are for testnet and mainnet-candidate operators. They are intentionally practical: identify the failure, preserve useful evidence, recover the node, and avoid making a temporary network issue worse.
      4 
      5 Assume the management UI is bound to 127.0.0.1 and the public P2P listener is the only internet-facing service. The default data directory is ~/.iuna; the default chain database is ~/.iuna/chain.sqlite3; the default config file is ~/.iuna/config.json.
      6 
      7 ## First Checks
      8 
      9 Before changing state, capture the local view:
     10 
     11 ```sh
     12 curl -s http://127.0.0.1:18661/api/status
     13 curl -s http://127.0.0.1:18661/api/network/health
     14 curl -s http://127.0.0.1:18661/api/peers
     15 curl -s http://127.0.0.1:18661/api/p2p/metrics
     16 ```
     17 
     18 These API routes require the same authenticated management session as the UI. If curl returns an auth response, collect the same data from the local UI or include the browser session cookie in the request.
     19 
     20 Record:
     21 
     22 - local height and tip hash;
     23 - best known peer height;
     24 - peer count, stale peers, banned peers, and last peer error;
     25 - current leader and last_auto_finalization_status;
     26 - whether the wallet is locked;
     27 - P2P listener, configured announce address, and whether inbound P2P is active.
     28 
     29 If the node runs with non-default paths, get them from startup logs or flags:
     30 
     31 ```sh
     32 iuna --help
     33 ```
     34 
     35 The process prints the wallet file, config file, chain database, management UI, and P2P listener on startup.
     36 
     37 ## Legacy address-book migration
     38 
     39 The upgraded wallet shows checksummed Bech32m receive addresses: iuna1... for
     40 mainnet/mainnet-candidate and tiuna1... for local testnet. Transfer,
     41 address-book, and Stratum input no longer accepts a 64-character hex public key.
     42 This is intentional: silently accepting hex would bypass typo and
     43 wrong-network protection.
     44 
     45 Before genesis or before reusing an old contact, ask the recipient to copy a new
     46 address from their upgraded wallet and replace the old address-book entry.
     47 Never change iuna to tiuna (or the reverse) by hand; the checksum covers the
     48 prefix. Existing wallet files and chain databases retain their internal hex
     49 keys and must not be rewritten. Controlled migration tools may call the domain
     50 migrate_legacy_address function only when the target network is explicitly
     51 known.
     52 
     53 No block-height activation is needed because on-chain and P2P address bytes do
     54 not change. Treat this as a client compatibility rollout instead: deploy the
     55 wallet UI and backend together, have recipients publish their new display
     56 address, and update Stratum worker usernames before restarting miners.
     57 
     58 ## Stalled Height
     59 
     60 Symptoms:
     61 
     62 - local height does not increase for longer than the target block time;
     63 - /api/network/health reports isolated, stale, peer errors, or syncing;
     64 - last_auto_finalization_status repeatedly says the wallet is locked, automatic mining is off, collecting burns, waiting for another finalizer, or VDF work is running too long.
     65 
     66 Checks:
     67 
     68 1. Compare local height with peers in /api/network/health.
     69 2. Check /api/peers for recent successful contact and peer heights.
     70 3. Check /api/status for wallet lock state, current leader, automatic mining, burn amount, VDF rounds, and recovery VDF threshold.
     71 4. Check host clock synchronization. Blocks too far in the future are rejected and peer clock warnings appear in network health.
     72 
     73 Recovery:
     74 
     75 - If the node has no healthy peers, add a known bootnode in the Peers screen or restart with --join <addr:port> only when there is no local chain database.
     76 - If the wallet is locked and this node is expected to finalize, unlock it in the UI.
     77 - If automatic mining is off and this node should participate, enable Mine settings and set a fee-paying burn policy.
     78 - If every healthy peer is at the same height, wait through fallback and recovery windows before resetting anything. A recovery block is the expected liveness escape hatch when ticket finalizers fail.
     79 - If peers are ahead, keep the node online and let range sync catch up. Use local chain reset only if the node remains stuck after it has healthy peers.
     80 
     81 Avoid:
     82 
     83 - deleting the wallet file;
     84 - repeatedly resetting multiple public nodes at once;
     85 - starting a new genesis while the network still has a valid chain.
     86 
     87 ## Divergent Tips
     88 
     89 Symptoms:
     90 
     91 - local height matches peers but tip hash differs;
     92 - peers report different tips at the same or nearby heights;
     93 - logs mention fork, snapshot, orphan, or block validation errors.
     94 
     95 Checks:
     96 
     97 1. Compare /api/status tip hash and height across at least two healthy peers.
     98 2. Inspect recent blocks in the UI or /api/blocks.
     99 3. Check /api/p2p/metrics for parse errors, session failures, and data envelope counts.
    100 4. Look for repeated snapshot requests in logs. Nodes request snapshots when a peer appears to be on a better valid fork.
    101 
    102 Recovery:
    103 
    104 - Before height 1000, if the divergence is inside the six-block legacy finality window, keep nodes connected. The protocol can reorg to a taller valid fork or to a same-height fork with better leader quality.
    105 - From height 1000, compare finalized_height and finalized_hash in /api/status or /api/network/health. Keep nodes connected when one valid chain has a higher checkpoint: fork choice adopts it even if its tip is temporarily shorter.
    106 - If your node is behind a healthy majority, add direct peers to that majority and let snapshot/range sync resolve it.
    107 - If your node is alone on an old tip and does not converge, use Delete local chain in Settings. That clears local chain/UI data, keeps wallet and settings, and broadcasts a snapshot request to peers.
    108 - If multiple public nodes report conflicting hashes at the same finalized height, preserve both databases and signing evidence immediately. Nodes deterministically choose the lower checkpoint hash, so operators do not need to trust the first peer seen, but the conflict is a quorum safety incident and must be investigated before resuming release activity.
    109 
    110 Avoid:
    111 
    112 - forcing a manual database replacement from an untrusted peer;
    113 - resetting the apparent majority tip before confirming it validates on more than one node.
    114 
    115 ## Old Snapshots Or Long Catch-Up
    116 
    117 Symptoms:
    118 
    119 - a node starts from an old persisted height;
    120 - /api/network/health reports syncing with peer height ahead;
    121 - startup logs say the node resumed a chain database at an old height.
    122 
    123 Checks:
    124 
    125 1. Confirm the node has at least one healthy outbound or inbound peer.
    126 2. Confirm peers know a higher height.
    127 3. Check /api/p2p/metrics for outbound connect failures or queue pressure.
    128 4. Confirm the configured P2P announce address is reachable from outside if the node is meant to be public.
    129 
    130 Recovery:
    131 
    132 - Leave the node online while it requests block ranges and snapshots from peers.
    133 - Add one or two stable peers manually if peer exchange has not found healthy peers.
    134 - Restart the process once if sessions are stale but the chain database loads successfully.
    135 - Use Delete local chain only if the node cannot extend its local snapshot and healthy peers are available. This does not remove wallet or settings.
    136 
    137 Avoid:
    138 
    139 - deleting the chain database before capturing the startup error;
    140 - using --genesis to recover an old node. --genesis is only for creating a fresh network.
    141 
    142 ## Legacy Pre-v6 Reset
    143 
    144 The live candidate release writes compact snapshot v7 and accepts snapshot
    145 versions v6 and v7. At startup, the node detects the legacy JSON schema and
    146 compact snapshot versions older than v6, checkpoints the database, archives it
    147 as chain.sqlite3.pre-v6 (or the next available numbered suffix), and creates a
    148 fresh database. The UI cache is then cleared normally. Wallet and configuration
    149 files are left untouched.
    150 
    151 The coordinated reset that created the live mainnet-candidate genesis is
    152 complete. A current operator encountering a pre-v6 archive should join the
    153 published candidate bootnode without --genesis; creating another genesis would
    154 create a separate, incompatible network. Preserve the generated .pre-v6
    155 archive if the old chain is needed as historical evidence.
    156 
    157 Manual equivalent for operators who want to choose the archive names before starting:
    158 
    159 ```sh
    160 mv ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.pre-v6 2>/dev/null || true
    161 mv ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.pre-v6 2>/dev/null || true
    162 iuna --join <trusted-peer-addr:port>
    163 ```
    164 
    165 For the disposable Compose testnet, docker compose down -v removes all volumes and the configured bootstrap creates the fresh genesis on the next start.
    166 
    167 Avoid:
    168 
    169 - running multiple independent --genesis nodes;
    170 - using --genesis to join or recover the live mainnet candidate;
    171 - copying an old snapshot blob into the current chain database;
    172 - deleting or replacing wallets as part of the chain reset;
    173 - joining a peer before checking the published network identity and release checksum.
    174 
    175 ## Height 1000 Consensus Activation
    176 
    177 Height 1000 is a coordinated consensus activation. At that height, VDF seeds
    178 start committing to block content and ticket draws stop using the final block
    179 hash. Burn-committee lineage draws make the same switch, closing the timestamp/
    180 block-hash grinding path for both leader and committee selection. Transaction
    181 signatures and native and Stratum mine proofs also switch to
    182 chain-bound binary format v1. Rank 0 blocks also start requiring a strict
    183 two-thirds committee quorum so their next rank 0 child can objectively certify
    184 them. This preserves blocks, snapshots, and UTXOs below
    185 1000, but nodes running the earlier rule will reject the upgraded chain or
    186 build an incompatible fork at activation. No database reset, new genesis, or
    187 migration command is needed for this height activation.
    188 
    189 Required activation procedure:
    190 
    191 1. Publish a tagged release, commit, checksums, and the activation height.
    192 2. Upgrade every known public peer, finalizer, and bootstrap node.
    193 3. Verify the reported package version on each managed node and compare tips.
    194 4. Confirm every independent node agrees on the block hash at height 999;
    195    upgraded nodes freeze pre-activation history once they reach 1000.
    196 5. Stop or isolate nodes that cannot be upgraded before activation.
    197 6. Keep chain database backups from immediately before the activation window.
    198 
    199 At and after height 1000, compare height and tip hash across at least three
    200 independent nodes. From height 1001, also compare finalized height and hash.
    201 If one chain has the higher valid checkpoint, normal sync should converge to it.
    202 If equal-height checkpoint hashes conflict, preserve both histories and the
    203 committee signatures, stop automated restarts, and investigate the quorum
    204 failure; the deterministic lower-hash rule is the recovery decision and does
    205 not by itself make the safety breach harmless.
    206 
    207 ## No Burn Committee Signatures
    208 
    209 Symptoms:
    210 
    211 - last_auto_finalization_status stays on collecting burns or says a block has too few burn bundle signatures;
    212 - there are fee-paying burns in mempools but rank 0 blocks do not finalize;
    213 - fallback blocks start appearing more often than expected.
    214 
    215 Checks:
    216 
    217 1. Check /api/mempool for fee-paying burn transactions.
    218 2. Check /api/peers for enough healthy peers. Burn bundles are gossip; finalizers request missing burn-bundle slots during collection, but isolated nodes still cannot receive responses.
    219 3. Check whether committee members are online and unlocked if they are expected to sign bundles.
    220 4. Inspect recent blocks: rank 0 needs the strictest burn-list attestation threshold; fallback ranks relax it for liveness.
    221 
    222 Recovery:
    223 
    224 - Improve peer connectivity first. Add stable peers and keep committee-capable nodes online.
    225 - Make sure automatic mining is enabled and wallets that should finalize/sign are unlocked.
    226 - Wait through fallback windows. Missing committee signatures should not permanently halt the chain; fallback ranks need fewer attestations.
    227 - If the network reaches recovery blocks, treat it as a warning that normal ticket finalization or committee gossip is unhealthy.
    228 
    229 Avoid:
    230 
    231 - treating local mempool contents as proof that a remote block is invalid. Block validation must be independent of each validator's mempool.
    232 - resetting a node merely because it did not personally see a burn bundle.
    233 
    234 ## Recovery Blocks
    235 
    236 Symptoms:
    237 
    238 - recent blocks show finalizer mode Recovery;
    239 - normal ticket blocks stopped for about the recovery delay;
    240 - last_auto_finalization_status mentions recovery VDF.
    241 
    242 Checks:
    243 
    244 1. Confirm whether recovery blocks are accepted by multiple peers.
    245 2. Check if selected ticket finalizers were offline, locked, not burning, or missing committee attestations.
    246 3. Check clock warnings. Large clock skew can cause valid-looking local candidates to be rejected by peers.
    247 4. Check VDF rounds and host performance if VDF work consistently finishes late.
    248 
    249 Recovery:
    250 
    251 - If recovery blocks are converging across peers, do not reset. Recovery is part of the liveness design.
    252 - Bring ticket finalizer nodes back online and unlocked.
    253 - Keep the fallback VDF threshold conservative on low-power hosts if too many nodes are wasting work on lower-ranked ticket paths. Every automatic node remains eligible for recovery once the recovery delay has elapsed.
    254 - After recovery, watch the next normal ticket blocks. Persistent recovery means the normal ticket path needs investigation.
    255 
    256 Avoid:
    257 
    258 - banning peers only because they produced a recovery block;
    259 - assuming recovery is an emergency fork. It is valid, weaker-for-fairness liveness behavior.
    260 
    261 ## Corrupted Local Persistence
    262 
    263 Symptoms:
    264 
    265 - startup fails with failed to load chain database;
    266 - SQLite reports it cannot open or parse chain.sqlite3;
    267 - the UI loads but chain views are empty or inconsistent while peers are healthy.
    268 
    269 Checks:
    270 
    271 1. Preserve the files before recovery:
    272 
    273 ```sh
    274 cp ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.bak 2>/dev/null || true
    275 cp ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.bak 2>/dev/null || true
    276 ```
    277 
    278 2. Keep wallet.json and config.json; they are not chain state.
    279 3. Confirm at least one trusted peer is available before deleting local chain state.
    280 
    281 Recovery:
    282 
    283 - Preferred: use Settings, Delete local chain. This clears local chain/UI data, keeps wallet and settings, and asks peers for a snapshot.
    284 - CLI fallback: stop the node, move only the chain/UI data files aside, then restart with the same wallet/config and healthy peers:
    285 
    286 ```sh
    287 mv ~/.iuna/chain.sqlite3 ~/.iuna/chain.sqlite3.corrupt 2>/dev/null || true
    288 mv ~/.iuna/ui_data.sqlite3 ~/.iuna/ui_data.sqlite3.corrupt 2>/dev/null || true
    289 iuna
    290 ```
    291 
    292 - If starting from an empty chain database, join a known peer:
    293 
    294 ```sh
    295 iuna --join <trusted-peer-addr:port>
    296 ```
    297 
    298 Avoid:
    299 
    300 - deleting wallet.json;
    301 - editing SQLite files in place;
    302 - accepting a snapshot from an unknown or isolated peer as the only recovery source.
    303 
    304 ## Escalation Bundle
    305 
    306 When a problem needs developer investigation, collect:
    307 
    308 - command line and version;
    309 - startup log lines showing wallet/config/chain paths and listeners;
    310 - /api/status, /api/network/health, /api/peers, and /api/p2p/metrics;
    311 - the last 20 blocks from /api/blocks?limit=20;
    312 - whether the wallet was locked and whether automatic mining, P2P inbound, and Stratum were enabled;
    313 - copies of corrupted chain/UI databases if persistence failed.
    314 
    315 Do not include recovery phrases, wallet passwords, or private keys in reports.