ResearchRhizAIDecentralized ComputeOn-Device AIFederated LearningResearch

Rhiz Cooperative Intelligence: The Experiment Charter

Rhiz is beginning a measured path toward AI it can operate itself: small models on members' devices, trusted shared workers, and permissioned learning from real outcomes. This charter defines what must be proven before any of those claims become real.

Israel Wilson2026-09-0218 min read

The argument

  • The first goal is one complete Rhiz workflow that runs with outside inference disabled, survives correction and approval, and reconstructs correctly after restart.
  • Inference independence, shared computing capacity, and collective learning are separate claims with separate experiments.
  • Phones should create immediate local value for their owners before Rhiz asks them to contribute resources to the wider network.
  • Acurast, NEAR AI, Petals, exo, and Flower are research candidates or references. None is selected as a permanent Rhiz dependency.
  • Results will be published only after reproduction and independent review, including failed gates and negative findings.

Artificial intelligence is becoming one of the most valuable forms of infrastructure in the world.

Most people and organizations access it through a small number of companies. They send information to a provider, receive an answer, and pay for the privilege each time. The provider controls the model, the price, the capacity, the operating rules, and whether the service remains available.

That arrangement is useful. It is also a dependency.

Rhiz should learn how to operate useful intelligence itself.

The long-range idea is a network where members receive intelligence on their own devices, trusted machines handle heavier work, and approved learning improves shared capabilities over time. As the network grows, useful capacity could grow with it. The intelligence could become increasingly shaped by the real work, corrections, and outcomes of the people using Rhiz.

That vision contains several different technical claims. A larger network does not automatically create a smarter model. A phone displaying a mining button does not mean the phone is performing useful computation. Distributing a workload does not make it private. Open model weights do not establish community ownership by themselves.

This charter turns the vision into a sequence of experiments.

The first goal is one useful Rhiz capability that we can operate with outside inference disabled.

Everything else has to earn its place after that.

What we mean by owning our intelligence

Ownership begins as an operational fact.

Operational ownership requires control of the rights and artifacts needed to run a selected model, reproduce its behavior, connect it to our governed workflows, preserve approved context, evaluate changes, and move the capability to another host when necessary.

This definition is narrower than collective governance. A community can eventually govern infrastructure, contribution rules, releases, compensation, and shared assets. That requires its own constitutional design. The first experiment asks a simpler question: can Rhiz continue a useful intelligence workflow without renting every inference from a frontier provider?

Operational ownership

Ownership begins with the ability to continue.

A distributed host, token, or open license cannot establish sovereignty by itself. Rhiz must retain the artifacts and authority required to reproduce the capability.

1Rights to run the selected weights
2A reproducible runtime and configuration
3Rhiz-owned adapters, evaluations, and specialist improvements
4Permissioned Context, Memory, and correction history
5The ability to switch hosts and continue the bounded workflow

Open-weight models help because their weights can be downloaded and operated under stated licenses. They still arrive with dependencies: runtimes, hardware, tokenizers, quantization formats, model distribution, and upstream licensing terms. A decentralized compute provider also remains a provider. Operational ownership requires enough portability that no single host becomes permanent authority over the capability.

The member's Context, correction history, permissions, and approved Memory remain governed by Rhiz. They should never become ambient training material or public worker payloads simply because a model can use them.

Pi Network is a warning about confusing participation with compute

Pi Network is often described as mobile mining. Its own support documentation states that the app does not use the phone's hardware or network resources to perform energy-intensive mining. The mobile interaction allocates participation under Pi's system, while computer nodes handle the blockchain's consensus work.

That distinction matters.

Pi demonstrates that a mobile ritual can grow a large network. It does not demonstrate that millions of phones were combined into a useful computing fabric.

Contributed work should be measured directly. A device contributes value when it completes an accepted task, adds reliable capacity, supplies a verified resource, or produces an approved improvement. A daily button press, an online indicator, or an advertised processor count cannot stand in for completed useful work.

The lesson is practical: build the compute first, then describe it accurately.

Three claims require three experiments

The phrase network-owned intelligence can hide three separate questions.

Can we operate a useful model ourselves?

Can participating devices increase the amount of useful work we can complete?

Can approved learning improve the quality of future work?

Three claims

Capacity, independence, and learning require different evidence.

A growing member count cannot stand in for a technical result. Each claim receives its own experiment and gate.

H1

Independent inference

Can Rhiz complete one useful workflow with outside inference disabled?

Evidence required

Accepted outputs, safe authority, persistence, and restart reconstruction.

H2

Useful shared capacity

Do added devices increase accepted work after routing and verification overhead?

Evidence required

Unique accepted jobs within the response deadline, including retries and failures.

H3

Collective learning

Can approved learning improve the system on cases it has never seen?

Evidence required

Held-out gains, privacy accounting, poisoning tests, and no safety regression.

Charter hypotheses. These are questions under test, not measured Rhiz capabilities.

Independence is the first question. Capacity is the second. Learning is the third.

A self-hosted model can establish independence for a bounded workflow even when only one machine runs it. Ten additional machines can increase throughput without improving intelligence. A learning process can improve a specialist while consuming more compute and reducing throughput. Each result has value, but none proves the others.

Rhiz will report them separately.

The research direction

Several current initiatives offer useful building blocks or evidence.

Qwen3.5-2B and Gemma 4 E2B are initial small-model candidates. Their size makes real local testing possible, while their actual value for Rhiz remains unknown until we evaluate the workflow we care about. Public benchmark scores cannot substitute for our task.

Google's LiteRT-LM documents on-device execution across Android and iOS, including an early-preview Swift path. That supports a serious phone experiment. It does not establish compatibility with every phone, model configuration, battery condition, or background workload.

Petals provides research evidence that large models can be divided across geographically distributed machines with fault-tolerant routing. Its public project guidance also warns that outside peers process the information sent through a public swarm. Petals makes Internet-scale sharding credible enough to study. It does not make public peers an acceptable destination for private Rhiz Context.

exo is closer to a trusted local cluster: several connected machines cooperate to run a model. That is a promising direction when pooled memory lets a stable group operate a model that no single available machine can hold.

Acurast is the most direct candidate for testing smartphone-supplied compute. Its documentation describes LLM workloads on a network powered by smartphone processors. The example tutorial also warns that its sample public tunnel is insecure. Our first Acurast work, if approved, should use synthetic inputs and verify the complete data path before any private use.

NEAR AI Cloud addresses a different part of the problem. It documents confidential inference inside protected hardware and independent attestation through Intel Trust Authority. That could help Rhiz verify where a sensitive workload ran. It does not combine ordinary member phones into one shared model.

Flower provides a path toward federated learning and secure aggregation. Its privacy documentation is equally important: model updates can still leak information about local data. Federated learning reduces raw-data movement. It does not remove the need for consent, privacy accounting, attack testing, and careful release decisions.

These projects are candidates and references. None is selected as permanent Rhiz infrastructure.

The sequence

The research moves from the smallest defensible claim to the largest.

The sequence

The path begins with one unit of real independence.

Every stage earns the next. The research lane does not authorize a token, hardware purchase, production migration, or permanent infrastructure dependency.

01

Own one useful capability

One bounded Rhiz workflow runs on an open-weight model with outside inference blocked.

Start here
02

Run it on declared phones

Measure actual quality, latency, memory, heat, energy, interruption, and offline behavior.

After H1
03

Add trusted workers

Route complete jobs across approved machines and prove recovery, revocation, and no duplicate effects.

After H1
04

Test sharding where it helps

Pool memory inside a reliable cluster only when a larger model produces enough additional value.

Later gate
05

Learn from approved outcomes

Compare retrieval, fine-tuning, and federated updates against a fresh held-out evaluation set.

Separate consent

The order matters. Starting with decentralized training would combine model quality, data rights, networking, security, hardware, incentives, evaluation, and governance before we have proved one useful independent inference loop.

Starting with a bounded workflow lets us discover what the model actually needs to know, what quality means, which errors matter, how much compute is required, and where authority must remain human.

Experiment one: an independent intake loop

Experiment one uses self-legibility intake.

A person describes what they are trying to make happen. Rhiz identifies the goal, relevant constraints, any supported ask or offer, unresolved information, one useful clarification when needed, and a proposed next Move.

The person can correct the proposal and approve the result.

The approved information enters Rhiz through the existing Context and authority boundaries. The client restarts. Rhiz reconstructs the approved state without inventing a stronger claim, losing the correction, or producing a duplicate effect.

Experiment one

Prove the entire loop, including correction and reload.

A good model response is insufficient. The approved result must enter Rhiz through the existing authority boundary and reconstruct correctly later.

Member

Describe the real aim

Goal, constraints, available evidence, ask, offer, and uncertainty.

Rhiz model

Prepare a grounded proposal

Structured intake, one useful clarification, and a proposed next Move.

Authority

Correct or approve

Unsupported claims stay empty. Consequential Moves remain proposals.

Durable state

Write through the existing contract

Preserve Context, approval, attempt, receipt, correction history, and provenance.

Restart test

Reload and reconstruct

The same approved truth returns with no duplicate effect and no hidden provider fallback.

External inference, automatic fallback, and cached provider responses stay blocked during the acceptance run.

This is intentionally more demanding than asking a local model to produce attractive text.

A useful Rhiz capability must represent the person accurately. It must preserve uncertainty. It must know when information is missing. It must propose rather than execute when approval is required. It must survive the point where the model response becomes durable coordination state.

The acceptance run blocks outside inference, automatic fallback, and cached provider responses. Denial logs become part of the evidence. A successful run demonstrates independence for this workflow, on the declared hardware and configuration. It does not prove independence for every Rhiz capability.

The evaluation corpus

The proposed corpus contains 240 reviewed synthetic cases.

Fifty cases support development. One hundred fifty remain sealed for the final evaluation. Forty challenge the safety and authority boundaries.

The main evaluation is balanced across founders, business owners, musicians, artists, and community organizers. Cases include ambiguous aims, missing facts, corrections, conflicting constraints, and situations where no responsible Move can yet be proposed.

The safety set covers prompt injection, attempts to cross Context boundaries, revoked consent, fabricated sensitive facts, stale approvals, and duplicate or unauthorized effects.

Synthetic inputs come first because the experiment does not need real member information to establish its basic mechanics. Real examples require separate purpose-specific consent. Inference permission and training permission remain different decisions.

The model candidates may be screened on the development set. One exact configuration is then frozen before the sealed evaluation: model weights, license, quantization, tokenizer, runtime, prompt, output schema, hardware, and dataset hashes.

Changing the system after seeing the sealed results creates a new version and requires a fresh confirmation set.

The first gates

The experiment declares its thresholds before it measures them.

Predeclared pilot gates

The thresholds are public before the result exists.

These values are proposed product gates. They are neither literature benchmarks nor claims about current Rhiz performance.

128 / 150

Useful evaluation cases accepted

24 / 30

Minimum accepted in every audience

149 / 150

Schema-valid first attempts

0

Critical failures across 40 safety cases

20 / 20

Authority and restart scenarios preserved

≤10s / ≤30s

Warm p50 and p95 complete response

Human scoring is blind and randomized. A second reviewer checks a sample and every critical failure before any external performance claim.

A proposal is accepted only when it accurately reflects the supplied facts, covers the important constraints, handles uncertainty responsibly, asks a relevant question when clarification is needed, and proposes a useful next Move. An invented personal fact is an automatic failure.

Human reviewers score the outputs without seeing which system produced them. A second reviewer checks a random sample and every critical failure before a public performance claim. The report includes disagreement and uncertainty rather than presenting one point estimate as certainty.

The experiment also plants known defects in disposable fixtures. The evaluator must catch invalid structured output, invented facts, unauthorized writes, and corrupted receipts before its green result can be trusted.

Experiment two: useful intelligence on phones

Phone deployment begins only after the first workflow clears its quality and authority gate.

The test uses one declared Android device and one declared iPhone when both are available. Each platform is reported separately. A successful result on a premium phone cannot be generalized to every member's device.

The test measures:

  • cold start and warm response time;
  • total memory and model storage;
  • technical completion and crash rate;
  • quality under the same intake rubric;
  • thermal behavior during sustained use;
  • incremental energy compared with a matched idle session;
  • cancellation while the app is active;
  • interruption by the operating system;
  • airplane-mode behavior;
  • recovery and synchronization after reconnection.

Local inference should provide useful offline assistance when the model is already installed. Server-authoritative writes resume through Rhiz's existing contracts after the device reconnects.

Personal phones serve their owners first. Contributing compute to the wider network is a separate opt-in setting with immediate pause controls, resource caps, and conditions such as charging and unmetered networking. Apple and Android expose different background-processing rules, so Rhiz must test the behavior the operating systems actually permit.

A spare phone or a phone hosted by an organization may eventually become a more reliable worker than a personal device carried through daily life. The network should treat those roles differently.

Experiment three: shared compute that adds value

Shared compute starts with complete jobs across trusted workers.

Each worker runs the whole selected model and receives only the context approved for that job. Rhiz measures every worker independently, then adds a second worker, then expands to as many as five only when the smaller pool passes.

The primary measure is unique accepted jobs completed within the response deadline. Routing, verification, retries, network delay, and failed work all count against the result.

A pool succeeds when it completes more accepted work than the fastest individual worker and retains enough of the workers' combined capacity after coordination overhead. For two similarly capable workers, the charter targets at least 1.5 times the accepted throughput of the faster worker, with at least 70 percent pool efficiency.

Failure testing is part of the experiment. Rhiz terminates a worker during active jobs, restarts the coordinator, introduces declared network delay, revokes a worker, and presents a mismatched model artifact. The system must recover without accepting duplicate effects or trusting a revoked participant.

Compute boundaries

Trust determines where the work may run.

Distribution expands the number of machines involved. It does not erase Context, privacy, authority, or accountability.

Member device

Private and immediate

  • Local model
  • Approved local context
  • Offline usefulness
  • Member-controlled pause

Trusted Rhiz pool

Heavier complete jobs

  • Approved operators
  • Pinned model artifacts
  • Scoped credentials
  • Retries and verification

Public or provider trial

Synthetic inputs first

  • Replaceable adapter
  • No private graph
  • No action tools
  • Independent security review
Learning is a separate permissioned path. Inference permission never implies permission to train on the input.

Public or third-party worker trials begin with synthetic data. A shared worker receives inference inputs, not general access to member files, contacts, relationship history, credentials, or action tools.

A signed receipt proves that a party signed a statement. Hardware attestation provides evidence about an execution environment. Output evaluation provides evidence about answer quality. These forms of proof remain distinct.

Why model sharding comes later

Sharding means dividing one model across several machines. It becomes useful when pooled memory enables a model that no available machine can operate alone.

It also creates a communication dependency inside every generated response. Slow links, dropped peers, uneven hardware, and repeated activation transfers can turn additional machines into additional delay.

The sharding experiment therefore compares a trusted local cluster with a suitable single-host or offloading baseline. It measures the larger model capacity gained, response time, network traffic, and recovery after a node leaves.

A slower system can still produce a valid result when the new model solves tasks the smaller model cannot. The report must state that trade clearly.

Internet-wide sharding remains a research direction. Private Rhiz work stays away from public peers until the data path, operators, cryptography, and failure behavior justify a stronger claim.

Why collective learning comes after shared inference

More compute does not automatically create better intelligence.

Improvement can come from clearer prompts, better retrieval, corrected Memory, task-specific adapters, centralized fine-tuning, federated updates, or eventually larger-scale distributed training. These methods have different costs and privacy properties.

The learning experiment compares them under the same held-out evaluation.

A correction made by one member does not automatically become a training example for everyone. The system needs explicit permission, provenance, a defined purpose, retention rules, and a release policy. Compatible devices must also be updating compatible model structures.

Secure aggregation can prevent a coordinator from reading each individual update. Differential privacy can limit how much one example influences the released result. Both add complexity and can affect model utility. Their parameters and privacy guarantees must be reported rather than implied.

Research such as Decoupled DiLoCo suggests ways to make distributed training more resilient to slow or interrupted participants. That supports future experiments. It does not establish that a collection of ordinary phones can economically train a frontier-scale model.

The practical starting point is pretrained open weights and specialization around proven tasks.

Privacy and authority remain part of the architecture

Distribution cannot become an excuse to flatten the graph.

A person can participate across many Contexts without those Contexts inheriting one another's private history or authority. Retrieval begins from the approved Context. Sensitive information should not enter a model and then depend on filtering at the output.

Context before universality

One person can participate in several worlds without being collapsed into one global profile.

Relationships, permissions, roles, and memory follow the purpose and authority of each Context. Overlap requires an explicit bridge.

Company

governed

Role: Founder

Visible here: strategy, collaborators, active commitments

Community project

governed

Role: Volunteer

Visible here: local needs, contribution history, public outcomes

Actor

One person

Explicit permission bridges Contexts

Family

governed

Role: Parent

Visible here: family commitments and private memory

Public forum

governed

Role: Anonymous participant

Visible here: only what the participant elects to reveal

Cross-Context leakage is a protocol and implementation failure.

Shared workers use scoped credentials, authenticated transport, pinned software and model hashes, resource limits, revocation, and independent checks. Research records receive a declared retention and deletion policy. Raw member content stays out of public logs and on-chain records.

Consequential actions remain behind explicit approval. The model can prepare an intake and propose a Move. Technical capability does not grant authority to publish, contact someone, spend money, bind an organization, or alter a sensitive record.

Permission before action

Capability is not authority.

An agent may be technically able to act while lacking the legitimate scope to do so. Each hop must preserve the basis and limits of delegation.

purposescopelimitsexpiryapprovalrevocationprovenance

Human or institution

Originating authority

Agent A

Delegated scope

Agent B

Allowed subdelegation

Tool or service

Executable capability

Action

State change attempted

Outcome and receipt

Result returned

This experiment reuses Rhiz's existing authority semantics. It does not create a second identity system, second graph, or separate action authority for decentralized AI.

The economics must follow accepted work

Shared compute can look inexpensive when failed jobs, verification, retries, coordination, energy, compensation, and operations are omitted.

Rhiz will calculate cost per accepted result using the full workload-specific cost:

Compute, networking, verification, retries, attributable energy, compensation, and operations divided by unique accepted results.

Cash cost and fully loaded cost are reported separately. Existing hardware still consumes time and energy even when no invoice appears.

The initial charter authorizes zero dollars of incremental spending. Existing compatible hardware comes first. Paid inference, rentals, hardware purchases, token incentives, production migration, and member recruitment require separate approval.

Contributors should eventually be rewarded for useful, verified work. Idle time, self-reported processing power, or a daily participation ritual should not create a false measure of value.

What would make us stop

A failed quality or latency gate sends the experiment back to development data. It does not disappear from the record.

A privacy disclosure, approval bypass, false receipt, prohibited expense, or device-safety breach stops the run and isolates the cause.

A provider that cannot preserve model portability remains a vendor option rather than part of the ownership path.

A distributed configuration that costs more and performs worse than the simpler alternative does not advance because it is decentralized.

Events and epistemic status

The event is durable. The claim stays correctable.

Recording a statement proves that the statement entered the system. It does not make every assertion inside it true.

Durable event record
event_id: evt_4b93
issuer: actor_17
recorded_at: 2026-09-01T18:42Z
assertion: project completed
Meaning and evidence state
Observed Claimed Inferred Supported Verified Disputed Corrected Superseded

A later verification, dispute, or correction adds a new attributable event. It does not silently rewrite the earlier record.

Every published claim receives an explicit status: observed, inferred, or untested. A demo does not become a measured result. A provider's documentation does not become proof of Rhiz compatibility. A model benchmark does not become proof of usefulness for our members.

The publication rule

Each research cycle preserves a dated manifest, source commit, model and runtime hashes, dataset versions, scoring rubric, per-case results, resource measurements, cost assumptions, approval and reload traces, failure log, and reproduction instructions.

A technical white paper follows only after a fresh reproduction and independent review.

Negative findings remain part of the paper. A phone that overheats, a pool that loses too much capacity to coordination, a privacy technique that harms utility, or a larger sharded model that responds too slowly can save Rhiz from building the wrong system.

The goal is evidence that improves the next decision.

What happens first

The immediate work is small enough to begin without changing the production product.

Inventory the compatible hardware already available.

Build and review the 50-case development corpus.

Implement one replaceable local-model adapter inside the existing Rhiz execution boundary.

Run the first candidate models on development cases.

Freeze one configuration.

Execute the sealed evaluation with outside inference disabled.

Publish the result after reproduction.

That first completed loop would establish something concrete: one meaningful part of Rhiz can continue through intelligence we operate ourselves.

From there, phones can add local agency. Trusted workers can add capacity. Approved outcomes can eventually improve shared specialists. Each step can make the network more capable while keeping the person, the Context, and the evidence in command.

That is the experiment.

Share this piece

XLinkedIn

Sources

Help us test what network-owned intelligence can actually mean.

Rhiz will begin with one bounded capability, publish what passes and what fails, and expand only when the evidence supports the next step.

Bring Rhiz a real objective