CausalRSI
Causal Credit Assignment for Recursive Self-Improving Agents
Research into which modification actually caused an improvement when a self-improving agent changes several components at once.
Research question
When a self-improving agent changes multiple components and then performs better, which modifications actually caused the improvement?
The problem
A self-improving agent rarely changes one thing. It rewrites a prompt, extends a memory, swaps a retriever and re-plans, then scores higher than it did before. Ordinary evaluation reports the difference and stops there. It cannot say whether the gain came from the memory, from the planner, from an interaction between the two, or from a change that merely correlated with the others. CausalRSI treats that gap as a causal credit-assignment problem: define the versioned subsystems as variables, intervene on them, and ask what the outcome would have been had one change not happened. Without that, a self-improvement loop is optimising against noise it cannot see.
Approach
Model the agent as versioned subsystems in a directed causal graph, record every change through a typed, hash-chained intervention ledger, and estimate per-subsystem and interaction effects against explicit counterfactuals rather than against a single aggregate score. Validity control and false-commit rate are treated as the floor the work has to clear, not as an afterthought.
Status
Early-stage research scaffold. Current results come from a toy simulation and are negative or unresolved: no tested history-versus-state comparison cleared the predeclared practical margin, and the specialised history mechanism under test was stopped at that scale. Research gates G1 and G2 are open. There is no real-LLM result, no benchmark, no deployment and no recursive-self-improvement claim.
Architecture
SUBSYSTEMS
The unit of analysis is a versioned subsystem, not a whole agent. Prompt, memory, retriever, planner, tools, reasoning, policy, code and evaluator each carry a version.
LEDGER
Changes are recorded in a typed event ledger with a SHA-256 record chain. Tamper-evident against a pinned root hash; not immutable against a rewrite.
ESTIMATION
Effects are estimated per subsystem and per interaction against explicit counterfactuals, with the residual carried openly rather than absorbed.
INFRASTRUCTURE
The Python package in this repository is research infrastructure and carries no empirical claims. It is dependency-free and guards its own import boundaries.
Demonstration values — ILLUSTRATIVE SIMULATION
The interactive experience uses the figures below. They are demonstration data for a deterministic toy model and are not experimental results.
| AGENT_0042 | 0.714 |
|---|---|
| AGENT_0043 | 0.842 |
| Difference | +0.128 |
| PROMPT | +0.01 |
|---|---|
| MEMORY | +0.07 |
| PLANNER | +0.04 |
| RETRIEVER | -0.01 |
| MEMORY x PLANNER | +0.05 |
| Residual / unexplained | -0.032 |
| AGENT_0042 | 0.714 |
|---|---|
| AGENT_0043A | 0.842 |
| AGENT_0043B | 0.769 |
| AGENT_0043C | 0.688 |
| AGENT_0044A | 0.871 |
| AGENT_0044B | 0.804 |
| AGENT_0045 | 0.903 |
Operators
- Arshan Bhanage — Research and engineering
- Parth Maradia — Research and engineering
- Mohsen — Research