Skip to main content
Public research references · IICP spec repo

Research & Validation

IICP should not ask people to trust vague promises. This page points to the public research record and explains, in plain language, what kind of evidence is needed before a claim becomes strong.

This page links public specification and research sources, then summarizes them in machine-readable-style detail blocks for implementers, operators and reviewers.

The short version

Use the public research record

The public research index explains why IICP makes certain design choices and what still needs proof.

Separate evidence levels

A simulation, a local test, a live network observation and a future design are different kinds of proof and should be labeled differently.

Check live status separately

This page explains the research rules. Current node health, adoption and mesh state belong on the live Nodes and Stats pages.

Bottom line: if a claim matters, it should be traceable to the public specification, a reproducible test, or live status on the Nodes/Stats pages.

Public sources to verify

Research index ↗

Track map for reputation, provider selection, routing, credit economy, adoption research and operational evidence discipline.

Validation methodology ↗

Rules for distinguishing simulation-backed feasibility estimates from empirical/live measurements and public claims.

Evidence model

The validation methodology requires public wording to say whether a claim is design rationale, simulation, controlled implementation evidence or live observation.

evidence_levels:
  design_rationale:
    meaning: "documented reasoning or architecture choice"
    claim_strength: "why this direction is plausible"
  simulation:
    meaning: "bounded model or harness result"
    claim_strength: "directional until validated in implementation"
  controlled_test:
    meaning: "reproducible local, CI or conformance evidence"
    claim_strength: "implementation behaviour under known conditions"
  live_observation:
    meaning: "measured in the running mesh"
    claim_strength: "current operational evidence, still time-bound"

public_rule:
  never_say: "simulation proves production performance"
  prefer: "simulation suggests; controlled tests validate; live data currently shows"

Strategic interoperability and taxonomy research

Current research supports a small, stable intent core with separate capability, policy, evidence and extension profiles. This keeps IICP focused on finding eligible execution capabilities under user-controlled constraints, while MCP and A2A remain complementary protocols.

Deterministic compatibility and receipt-boundary simulations support preserving existing URNs, using redacted directory metadata with correlated client/node receipts, and researching composite inverse-load selection next. These are pre-normative design decisions, not a new deployed protocol contract.

Read the public strategic research note ↗

Questions reviewers should ask

Are old performance numbers production measurements?

No. The validation methodology says the historic headline numbers are simulation-backed feasibility estimates, not production SLOs.

Public source: Validation methodology §§2–4

concern: "performance overclaiming"
public_source: "spec/v1.9/validation-methodology.md"
current_answer:
  historic_figures: "simulation-backed feasibility estimates"
  not_allowed: "cite as measured production performance"
  allowed: "cite as modelling evidence with the simulation caveat"
required_next_evidence:
  - reproducible load tests
  - live latency/availability measurements
  - clear separation of inference time vs routing overhead

How does IICP avoid bad provider selection or early lock-in?

Provider-selection research compares simple routing, exploration, reputation, multi-path modes and cold-start effects; the public rule is to keep best-effort routing honest until stronger evidence exists.

Public source: Research index — R1R7 and RS6 tracks

concern: "provider concentration and cold start"
public_source: "research/RESEARCH.md"
researched_dimensions:
  - reputation-aware selection
  - exploration vs exploitation
  - bootstrap floor for new capable nodes
  - high-assurance multi-path strategies
public_claim_limit:
  best_effort: "usable default routing"
  not_yet: "perfect provider choice or universal answer quality"

Can reputation, Silver/Gold/Platinum or trust scores be over-trusted?

Reputation is an evidence signal. The public research record documents penalty asymmetry, decay, identity-age gates and adversarial/collusion simulations, but it does not make trust absolute.

Public source: Research index — REP track

concern: "trust score becomes a magic badge"
public_source: "research/RESEARCH.md"
known_controls:
  penalty_asymmetry: "failure penalty larger than success bump"
  identity_age_gate: "used to reduce whitewash risk for high tiers"
  feedback_ema: "bounded feedback influence"
  bootstrap_floor: "lets capable new nodes gather evidence"
still_needed:
  - verified receipts
  - anti-self-dealing enforcement
  - broader independent operator data
public_claim_limit: "score is evidence, not a guarantee"

What makes reachability, relays and IPv6 hard?

A node can be alive but not independently reachable from every vantage point. Research must distinguish self-attested, directory-observed and externally signed probe evidence.

Public source: Validation methodology §7

concern: "reachable from one path, unreachable from another"
public_source: "spec/v1.9/validation-methodology.md §7"
evidence_basis:
  self_attested: "node says it is reachable"
  directory_observed: "directory origin reached it"
  external_probe_ipv6: "IPv6-capable worker signed a probe result"
relay_rule:
  direct_first: true
  quick_tunnel: "bootstrap/fallback, not production availability promise"
  stable_relay: "requires operator-controlled endpoint and relay-path controls"
dated_live_observation:
  date: "2026-07-05"
  result: "local sleep/low-power route loss recovered without manual restart"
  limit: "limited live sample; not a reliability guarantee"

Does encryption mean prompts are universally private?

No. Keyed payload routing protects the path to nodes that publish keys, but a provider executing a prompt can still read the prompt it runs.

Public source: Validation methodology §7 and core security rules

concern: "privacy overclaiming"
public_source: "spec/v1.9/validation-methodology.md §7"
accurate_statement:
  path: "clients can encrypt to nodes that publish keys"
  execution: "the executing provider sees the prompt it executes"
  transition: "claims must remain conservative while key coverage and encrypted round-trip evidence can change with live node churn"
not_the_same_as:
  - universal confidentiality
  - metadata privacy
  - permanent proof that every future live node is key-ready

Can a browser tab really become part of the mesh?

Browser execution is useful for adoption and testing, but browser provider mode depends on relay-backed serving today; direct WebRTC-style serving needs a signaling contract first.

Public source: Validation methodology §7 and research index

concern: "browser node sounds production-ready"
public_source: "spec/v1.9/validation-methodology.md §7"
current_shape:
  browser_consumer: "can test mesh use from a web page"
  browser_provider: "relay-backed experimental participation path"
not_yet_proven:
  - production-grade browser uptime
  - WebRTC self-addressing without signaling spec
  - relay confidentiality equivalent to direct encrypted path

Why does operator identity matter?

Without a stable operator signal, reputation and diversity controls can be gamed by repeatedly creating fresh nodes. Identity helps, but is not a trust badge by itself.

Public source: Research index and semantics/core specs

concern: "many node IDs, one hidden operator"
public_source: "research/RESEARCH.md"
identity_purpose:
  - bind nodes to a stable operator signal
  - support future diversity and anti-Sybil policy
  - make reputation harder to reset by relabeling
claim_limit:
  identity_present: "orientation signal"
  identity_not: "automatic proof of honesty"

Does GDPR/EU AI Act readiness mean IICP is legally certified?

No. It means the project has started turning privacy, rights, routing, retention and risk claims into documented controls and testable gates. Legal certification and production DPAs are separate work.

Public source: IICP.network public privacy notice baseline

concern: "legal compliance overclaiming"
public_source: "project/compliance and /legal/privacy"
audit_scope:
  included: "directory, clients, nodes, relay/browser paths, public docs and runtime claims"
  excluded: "agent-loop tooling and internal development automation"
current_controls:
  - beta privacy notice baseline
  - manual DSR workflow
  - retention matrix and breach/prompt-leak runbook
  - AI-risk classification and DPIA trigger checklist
  - remote-routing profiles: sensitive, eu-restricted, strict-policy
  - signed policy manifests with first operator-bound directory signal
still_needed:
  - self-service DSR/operator portal
  - cross-client known-operator transfer evidence
  - durable policy-key revocation records
  - redacted routing receipts
  - lawyer-reviewed policy, terms, DPAs and subprocessor handling
public_claim_limit: "compliance-readiness, not compliance certification"

Why not always ask many nodes and compare answers?

Multi-path can improve assurance for high-value tasks, but it costs more and can be wrong for latency-sensitive or low-value tasks. The research keeps this as a mode, not a universal default.

Public source: Research index — R6 multi-path execution

concern: "more providers always means better"
public_source: "research/RESEARCH.md"
routing_modes:
  best_effort: "single provider, default for normal tasks"
  redundant: "multiple providers when assurance is worth the cost"
  byzantine: "expensive mode for adversarial threat environments"
claim_limit:
  not_default: "quorum answering for every request"
  reason: "cost, latency and task value differ"

Is IICP already a fully decentralized federated control plane?

No. Federation groundwork exists, but Phase 6 claims depend on Phase 5 security, reachability, relay and privacy evidence becoming stable first.

Public source: Validation methodology §7 and research index

concern: "future federation presented as current reality"
public_source: "spec/v1.9/validation-methodology.md §7"
current_claim:
  allowed: "federation is researched and partially prepared"
  not_allowed: "the public mesh is already decentralized/federation-ready"
gates_before_phase6:
  - stable relay and reachability evidence
  - SDK/key adoption evidence
  - privacy and receipt hardening
  - governance and conformance maturity

How claims should be worded

AreaEvidence posturePublic wording guard
Discovery and routingBest-effort routing is implemented in the reference meshKeep perfect-provider-selection claims out of public wording.
ReputationResearch-backed evidence signalDo not present tiers as a guarantee of answer quality.
Node healthLive pages show operational signalsKeep status values tied to live Nodes/Stats instead of static claims here.
RelayUseful for reachability, still trust-sensitiveDifferentiate stable operator relay from bootstrap tunnel fallback.
IPv6 reachabilityRequires correct probe vantage pointDo not turn directory IPv6 limitations into false node failures.
Browser providerAdoption/test pathDo not imply production browser serving until the transport contract is stronger.
PrivacyKeyed payload path where keys existState provider-visibility and incomplete metadata/privacy work plainly.
Compliance readinessBaseline controls and docs existDo not present beta workflows as legal certification or production DPAs.
FederationGroundwork/researchDo not imply Phase 6 production decentralization before gates close.

How to challenge or improve the research

If a claim is unclear, overstated or contradicted by your own measurements, the most useful contribution is reproducible evidence: the spec section, client version, network conditions, command or test used, expected behaviour and observed behaviour.