Skip to the text version of this research instrument
SYSTEM / 000AGENT STATE RECOVEREDAGENT_0042
SUBSYSTEMS 08 / RELATIONS 14LEDGER CHAIN OKOBSERVED BEHAVIOUR ONLY
0.714ILLUSTRATIVE SIMULATION
+17.9%ILLUSTRATIVE SIMULATION
BUT
WHY?
CAUSALRSI
Causal Credit Assignment for Recursive Self-Improving Agents
Phases: AGENT_N at 2 percent; MODIFICATIONS at 17 percent; AGENT_N_PLUS_1 at 32 percent; CAUSAL_ATTRIBUTION at 47 percent; COUNTERFACTUALS at 62 percent; CREDIT_ASSIGNMENT at 77 percent; RECURSIVE_IMPROVEMENT at 92 percent

Current causal state

Active agent AGENT_0043, illustrative score 0.842.

Baseline AGENT_0042, illustrative score 0.714.

No intervention applied.

Factual world.

Illustrative attribution: PROMPT +0.01, MEMORY +0.07, PLANNER +0.04, RETRIEVER -0.01. Interaction MEMORY by PLANNER +0.05. Residual, unexplained, -0.032.

All figures are illustrative simulation values, not experimental results.

CausalRSI

Causal Credit Assignment for Recursive Self-Improving Agents

Research into which modification actually caused an improvement when a self-improving agent changes several components at once.

Research question

When a self-improving agent changes multiple components and then performs better, which modifications actually caused the improvement?

The problem

A self-improving agent rarely changes one thing. It rewrites a prompt, extends a memory, swaps a retriever and re-plans, then scores higher than it did before. Ordinary evaluation reports the difference and stops there. It cannot say whether the gain came from the memory, from the planner, from an interaction between the two, or from a change that merely correlated with the others. CausalRSI treats that gap as a causal credit-assignment problem: define the versioned subsystems as variables, intervene on them, and ask what the outcome would have been had one change not happened. Without that, a self-improvement loop is optimising against noise it cannot see.

Approach

Model the agent as versioned subsystems in a directed causal graph, record every change through a typed, hash-chained intervention ledger, and estimate per-subsystem and interaction effects against explicit counterfactuals rather than against a single aggregate score. Validity control and false-commit rate are treated as the floor the work has to clear, not as an afterthought.

Status

Early-stage research scaffold. Current results come from a toy simulation and are negative or unresolved: no tested history-versus-state comparison cleared the predeclared practical margin, and the specialised history mechanism under test was stopped at that scale. Research gates G1 and G2 are open. There is no real-LLM result, no benchmark, no deployment and no recursive-self-improvement claim.

Architecture

  • SUBSYSTEMS

    The unit of analysis is a versioned subsystem, not a whole agent. Prompt, memory, retriever, planner, tools, reasoning, policy, code and evaluator each carry a version.

  • LEDGER

    Changes are recorded in a typed event ledger with a SHA-256 record chain. Tamper-evident against a pinned root hash; not immutable against a rewrite.

  • ESTIMATION

    Effects are estimated per subsystem and per interaction against explicit counterfactuals, with the residual carried openly rather than absorbed.

  • INFRASTRUCTURE

    The Python package in this repository is research infrastructure and carries no empirical claims. It is dependency-free and guards its own import boundaries.

Demonstration values — ILLUSTRATIVE SIMULATION

The interactive experience uses the figures below. They are demonstration data for a deterministic toy model and are not experimental results.

Agent scores
AGENT_00420.714
AGENT_00430.842
Difference+0.128
Attribution
PROMPT+0.01
MEMORY+0.07
PLANNER+0.04
RETRIEVER-0.01
MEMORY x PLANNER+0.05
Residual / unexplained-0.032
Candidate lineage
AGENT_00420.714
AGENT_0043A0.842
AGENT_0043B0.769
AGENT_0043C0.688
AGENT_0044A0.871
AGENT_0044B0.804
AGENT_00450.903

Operators

  • Arshan Bhanage — Research and engineering
  • Parth Maradia — Research and engineering
  • Mohsen — Research

Source

SOURCE github ↗