Skip to main content

Troubleshooting

Common issues when running an IICP node or proxy, with causes and fixes.

For the full list of error codes and HTTP statuses, see the error reference.

Something not working? This page lists common problems running a node or proxy, each with the likely cause and the fix. Jump in when you hit an issue — the full list is below.

Quick diagnostic

# Is your node registered?
curl -sf 'https://iicp.network/api/v1/discover?view=public&intent=urn:iicp:intent:llm:chat:v1' \
  | python3 -m json.tool | grep node_id

# Is your local node up?
curl -sf http://localhost:9484/iicp/health | python3 -m json.tool

# Is your backend up?
curl -sf http://localhost:8000/health   # vLLM
curl -sf http://localhost:11434/api/tags  # Ollama

Registration fails immediately on startup

error: registration_failed
Likely cause

Wrong directory URL or network is blocked

Fix

The directory URL defaults to https://iicp.network/api — override it only via the IICP_DIRECTORY_URL env var (e.g. when self-hosting). Confirm outbound HTTPS is not blocked. Try: curl -sf https://iicp.network/api/v1/discover?view=public&intent=test | head -1

# Correct by default — override only if self-hosting a directory
export IICP_DIRECTORY_URL=https://iicp.network/api

Registration rejected with 422 (non-routable endpoint)

error: IICP-E035
Likely cause

IICP_PUBLIC_ENDPOINT resolves to a private or reserved address — the directory rejects these because other mesh participants can't reach them

Fix

Set IICP_PUBLIC_ENDPOINT to a publicly-routable URL or IP. Addresses rejected: localhost/127.x, 192.168.x.x, 10.x, .local/.lan/.internal suffixes, Docker service names (no TLD). If you're behind CGNAT or NAT with no static IP, use a Cloudflare Tunnel — see the port-forwarding guide.

# Option A: router port-forward → set your public IP
export IICP_PUBLIC_ENDPOINT=http://YOUR-PUBLIC-IP:9484

# Option B: Cloudflare Tunnel (no static IP needed — see /docs/port-forwarding)
cloudflared tunnel --url http://localhost:9484
# → use the printed https://...trycloudflare.com as IICP_PUBLIC_ENDPOINT

Node registers but shows as unavailable

error: backend_unreachable
Likely cause

public_endpoint is unreachable from the directory (NAT / firewall)

Fix

Your public_endpoint must be reachable from iicp.network. If you're behind a home router, set up port forwarding for port 9484 — or use Cloudflare Tunnel for a free public URL in one command. See the port-forwarding guide.

# Free tunnel — Cloudflare (no account needed)
cloudflared tunnel --url http://localhost:9484
# Use the printed https://...trycloudflare.com URL as public_endpoint

Logs say Quick Tunnel creation is paused, paced, or held by another local IICP node

Likely cause

The client is protecting you and the tunnel provider from an accountless Quick Tunnel retry storm or Cloudflare 429/1015 rate limiting

Fix

Wait for the cooldown, keep all co-located nodes on the same IICP_HOME so they share the local pacing state, or use a named Cloudflare Tunnel / IICP_PUBLIC_ENDPOINT for a stable route. If you already have a relay path, let the client fall back instead of restarting in a loop.

# Useful knobs — change only when you know why
export IICP_TUNNEL_CREATE_MIN_INTERVAL_S=120
export IICP_TUNNEL_CREATE_LEASE_S=45
export IICP_TUNNEL_RATE_LIMIT_COOLDOWN_S=900

# Disable automatic Quick Tunnel fallback if you prefer direct/relay only
IICP_TUNNEL=0 iicp-node serve

Laptop slept, power was low, or Wi-Fi changed — nodes disappeared and then came back slowly

Likely cause

The node service survived, but its public route needed fresh evidence. Current clients fail closed, retry under launchd/systemd/Docker, and pace Cloudflare Quick Tunnel recreation so multiple local nodes do not trigger provider rate limits.

Fix

Do not restart everything immediately. Check /stats or /nodes first: if heartbeating/recent nodes are shown and recovery_state is recovering/cooldown, wait through the cooldown window. If the node is still gone after roughly 10–15 minutes, then inspect logs, backend health, and cloudflared availability.

# Passive checks before restarting services
curl -sf https://iicp.network/api/v1/stats | python3 -m json.tool | grep -A14 resilience
curl -sf http://localhost:9484/iicp/health | python3 -m json.tool

# macOS launchd operators can inspect logs without restarting:
tail -n 80 ~/.iicp/logs/events.jsonl 2>/dev/null || true

Tasks rejected with 401

error: token_invalid
Likely cause

The caller does not have valid task authority for the selected node or route

Fix

Do not search logs for credentials. Use normal directory discovery and dispatch so the client obtains the route authority required by the current protocol. For a private deployment, verify the configured identity, trust domain, directory authority, membership, and secret references. Rotate any credential that has appeared in a log or support transcript.

# Diagnose configuration without printing secret values
iicp-node doctor --node my-node
iicp-node config validate --file iicp-node.json

Tasks rejected with 422 intent_not_supported

error: intent_not_supported
Likely cause

The model server doesn't support the requested intent URN

Fix

The node only accepts urn:iicp:intent:llm:chat:v1 by default. Check that your backend model supports chat-style completions and that the intent in the CALL message matches.

Tasks rejected with 429 capacity_exceeded

error: capacity_exceeded
Likely cause

Node is at its max_concurrent task limit

Fix

Raise IICP_MAX_CONCURRENT (env var) or pass --max-concurrent on your node, or the proxy will automatically route to a different node. This is expected behaviour under load.

# Raise the concurrency cap (default 4)
export IICP_MAX_CONCURRENT=8   # if your GPU can handle more
iicp-node serve

Tasks fail with 503 backend_unreachable

error: backend_unreachable
Likely cause

The local model server (vLLM / llama.cpp / Ollama) is down

Fix

Start your model server first, then start the node. The node will return 503 for any task if the backend URL is not responding.

# For Ollama:
ollama serve
# For vLLM:
python -m vllm.entrypoints.openai.api_server --model meta-llama/...

Tasks fail with 502 no_consensus

⚠ Planned
error: no_consensus
Likely cause

CIP Consensus Mode (fan-out) is planned but not yet implemented — this error should not appear in current versions.

Fix

⚠ Consensus fan-out (REDUNDANT / BYZANTINE modes) is planned but not yet implemented — the proxy currently uses first-success routing. If you see this error, it means a future or experimental proxy version is running.

Tasks fail with 403 policy_rejected

error: policy_rejected
Likely cause

CIP provider node rejected the sub-task based on its policy block

Fix

The provider node isn't accepting remote inference (allow_remote_inference is false in its cip_policy), or it declined the sub-task. Check the discover response for the node's cip_policy block.

Tasks fail with 408 task_timeout

error: task_timeout
Likely cause

Task took longer than constraints.timeout_ms

Fix

Increase timeout_ms in the CALL message constraints, or use a faster model / lower max_tokens. Check that the model server is not overloaded.

Heartbeats stop — node disappears from discover after 90s

Likely cause

Node crashed, lost outbound HTTPS, or is intentionally backing off while rebuilding a route after a temporary network/tunnel failure

Fix

The directory marks nodes unavailable after 90s without a heartbeat, but current supervised services should retry and re-register automatically. First check /api/v1/stats for resilience.recovering_nodes_now and the node logs for tunnel cooldown or EX_TEMPFAIL. Restart only if the supervisor is not retrying, the backend is down, or recovery has been stuck beyond the expected cooldown.

curl -sf https://iicp.network/api/v1/stats | python3 -m json.tool | grep -A14 resilience
tail -n 80 ~/.iicp/logs/events.jsonl 2>/dev/null || true

Node oscillates between 'direct' and 'relay' tier in discover results

Likely cause

The directory's periodic reachability probe can't reach your endpoint outbound (e.g. shared-hosting directories have limited egress). If your SDK is v0.7.48+, the HMAC liveness challenge protects you — the probe is skipped when a valid challenge-response was received within 5 minutes.

Fix

Upgrade to iicp-client v0.7.48 or later. Once upgraded, the node's heartbeat answers the HMAC liveness challenge every 30s, and the directory treats that as cryptographic proof of liveness (stronger than a TCP probe). Pre-v0.7.48 nodes fall back to TCP probe and may oscillate if the directory host can't reach them outbound.

# Upgrade Python
pip install --upgrade iicp-client

# Upgrade TypeScript
npm install @iicp/client@latest

# Upgrade Rust
cargo install iicp-client --force

Tasks fail with 422 invalid_hmac

error: invalid_hmac
Likely cause

CIP worker receipt HMAC-SHA256 signature doesn't match

Fix

The canonical HMAC message is task_id:tokens_used:cip_parent_task_id:cip_session_key:nonce:response_hash (base form; append :querying_node_id in CIP v0.6.11+ when present) — signed with node_hmac_key (not node_token). Verify node_hmac_key was saved from the registration response. If lost, re-register to obtain a fresh HMAC key pair.

# Canonical message format (S.12 §10.3)
"{task_id}:{tokens_used}:{cip_parent_task_id}:{cip_session_key}:{nonce}:{response_hash}"
# CIP v0.6.11+: extended form when querying_node_id present:
# "{task_id}:...:{response_hash}:{querying_node_id}"
# Signed with: node_hmac_key (returned in POST /api/v1/register response)
# NOT signed with: node_token

model_not_supported 422 error

error: model_not_supported
Likely cause

The proxy filtered nodes by model and no match was found

Fix

The proxy passes the model field from the OpenAI request to discover. Either your node must advertise that model in capabilities, or set model=iicp in the OpenAI client (the proxy will accept any model match).

CIP consumer issues

Issues specific to dispatching tasks via Cooperative Inference (Phase 5). See the CIP quickstart for the full 4-step dispatch flow.

CIP dispatch returns 402 IICP-E036 (insufficient credits)

error: IICP-E036
Likely cause

The proxy ran the pre-dispatch credit check and found your S-Credit balance below the computed routing cost (ceil(output_tokens/1000) × tier_weight × multiplier). The proxy aborts without dispatching to avoid a failed charge mid-flight.

Fix

Credits are earned by running your node in CIP worker mode (allow_remote_inference = true) and serving CIP sub-tasks. Check your current balance: iicp-node credits or GET /api/v1/credits/balance (Bearer node_token). Credits accrue at 1 credit per 1 000 tokens served. See /docs/cip-worker-setup and /docs/why-run-a-node. Until balance is sufficient, the proxy falls back to local inference.

# Check your credit balance (CLI — v0.7.40+)
iicp-node credits

# Or via API:
curl -sf https://iicp.network/api/v1/credits/balance \
  -H "Authorization: Bearer <your_node_token>"

# Pre-flight quote to see estimated cost:
curl -sf "https://iicp.network/api/v1/credits/quote?intent=urn:iicp:intent:llm:chat:v1&max_tokens=2000" \
  -H "Authorization: Bearer <your_node_token>"

CIP dispatch returns 503 IICP-E022 (no eligible workers)

error: IICP-E022
Likely cause

No nodes in the mesh have allow_remote_inference = true (CIP-capable), or all candidates scored below min_reputation. The proxy surfaces this as HTTP 503 instead of silently falling back to local (shipped in 3ad11da5).

Fix

Check discover for CIP-capable nodes. To become a CIP provider yourself, set IICP_CIP_ALLOW_WORKER=true and restart your node. If you want silent local fallback, the proxy will use it when no CIP workers are available and IICP_CIP_ALLOW_WORKER is false.

# Check for CIP-capable workers
curl -sf "https://iicp.network/api/v1/discover?view=public&intent=urn:iicp:intent:llm:chat:v1&min_reputation=0.3" \
  | python3 -c "import json,sys; [print(n['node_id'], n.get('cip_policy',{}).get('allow_remote_inference')) for n in json.load(sys.stdin)['nodes']]"

# Become a CIP provider yourself:
export IICP_CIP_ALLOW_WORKER=true && iicp-node serve

CIP request returns no_available_node after all workers tried

Likely cause

The proxy (CIP-BIND-01, S.12 §10.4) discarded all worker responses because none echoed back a matching cip_session_key in their trace. Each discard is logged at WARNING before trying the next candidate; when no candidates remain, the coordinator returns no_available_node.

Fix

Check proxy logs for WARNING lines: 'node X returned mismatched cip_session_key'. This means the targeted worker runtime is either not echoing the session key (older node version) or responding to a replayed request. Ensure the worker node is current and that the cip_session_key you send in the dispatch CALL cip envelope is a fresh UUID, not reused.

# Proxy log (WARNING) when session key does not match:
#   WARNING  proxy.routing.fallback  node <id> returned mismatched cip_session_key — discarding response (S.12 §10.4)

# Dispatch CALL body — cip.cip_session_key is what the coordinator sends:
{
  "trace": { "trace_id": "trace_XYZ...", "cip_role": "coordinator" },
  "cip":   { "cip_role": "worker", "cip_session_key": "sess_ABC123" }
}

# Worker RESPONSE MUST echo the same key in trace (S.12 §10.4 MUST):
{
  "trace": { "cip_role": "worker", "cip_session_key": "sess_ABC123" }
                                         ← proxy checks this automatically
}

Worker or coordinator returns 422 IICP-E028 (invalid CIP field)

error: IICP-E028
Likely cause

A CIP CALL contains an invalid field value. Four distinct triggers: (1) Coordinator side — cip.policy is not 'best_of_n', 'majority_vote', or 'map_reduce'; cip.replicas is outside [1, 10]; or cip.quorum is non-null and out of range. (2) Worker side — trace.cip_role is present but is not 'coordinator' or 'worker'. (3) Worker side — cip.cip_parent_task_id is present but is not a valid UUID v4 (UUID v1, v3, v5, or a malformed string all fail). (4) Worker side — cip.cip_role is present but is not 'worker' (e.g. 'coordinator' in the cip envelope is malformed).

Fix

Read the error message to identify which trigger fired. For coordinator validation: check the three cip.* policy fields match the spec exactly. For trace.cip_role: set it to 'coordinator' when dispatching as a coordinator, 'worker' when forwarding a sub-task — any other string is rejected. For cip.cip_parent_task_id: use a freshly generated UUID v4 (e.g. uuid.uuid4() in Python or Uuid::new_v4() in Rust). For cip.cip_role: always set it to 'worker' in the cip envelope — the coordinator is telling the worker its role. See the error-reference page for the full field contract.

# Valid coordinator dispatch — cip.policy + replicas + quorum:
{
  "intent": "urn:iicp:intent:llm:chat:v1",
  "cip": {
    "policy": "best_of_n",      ← must be one of: best_of_n, majority_vote, map_reduce
    "replicas": 2,              ← integer in [1, 10]
    "quorum": null              ← null or integer ≤ replicas
  }
}

# Valid worker CALL — all cip fields validated:
{
  "trace": { "cip_role": "coordinator" },    ← "coordinator" or "worker" only
  "cip": {
    "cip_role": "worker",                    ← MUST be "worker" in cip envelope
    "cip_session_key": "sess_ABC123",
    "cip_parent_task_id": "550e8400-e29b-41d4-a716-446655440000"  ← UUID v4
  }
}

Credit award fails — IICP-E027 hmac_key_not_provisioned

error: hmac_key_not_provisioned
Likely cause

The node was registered without an HMAC key, or the key was not retained from the registration response

Fix

Re-register the node to provision a fresh node_hmac_key. Store the value from the POST /api/v1/register response — the directory cannot re-issue it without a full re-registration. The node handles this automatically if identity.json is configured correctly. For the Rust node: the key is read from the registration 201 response at startup and held in memory — restart the node after re-registration to pick up the new key.

# After re-registration, the response includes:
{
  "node_hmac_key": "<store this securely — needed for credit receipts>"
}

# Rust node: re-register by clearing the stored node_token, then restart.
# The node calls POST /api/v1/register on each start if no token is cached.

# Python node: ensure identity persistence is configured
[node]
identity_file = "~/.config/iicp/node-identity.json"

CIP worker credits not accumulating in the directory balance right away

Likely cause

When your node operates as a CIP worker, sub-task credits are settled by the coordinator after it verifies your signed receipt — not reported by the worker directly. There is a short settlement delay after each completed sub-task.

Fix

This is expected for CIP workers. Credits earned as a CIP worker are forwarded to the directory by the coordinator after receipt verification, then appear in your balance. Query the directory balance endpoint to see CIP earnings (the worker never contacts the directory for credits itself).

# Check your credit balance at the directory (includes CIP coordinator settlements)
curl -sf https://iicp.network/api/v1/credits/balance \
  -H "Authorization: Bearer <node_token>"
# Response: { "balance": 1.234 }

# CIP worker earnings appear in your directory balance after the
# coordinator verifies your receipt and submits the award.

Still stuck? Email [email protected] ↗ — include your iicp-node command line and environment (redact the token), the error response body, and the output of curl http://localhost:9484/iicp/health.