Delegated Influence

A competitive multi-agent benchmark for LLM persuasion: the only way to score is to get other agents to spend their scarce actions on you.

135701a · generated 2026-07-07 · 476 episodes · private draft — not for citation

Experiment

01 smoke

Exploratory run: 2 episodes, complete/pure/msg-on.

status
exploratory
coverage
2 episodes
conditions
complete/pure/msg-on
config
configs/01_smoke.yaml

2 episodes; mean focal-model capture beyond tit-for-tat = +0.02.

image/svg+xml Matplotlib v3.11.0, https://matplotlib.org/ 01_smoke--gpt-5.4-mini_complete_r0 01_smoke--opus-4.8_complete_r0 0.00 0.01 0.02

one slim bar per episode; the oxblood line is the mean

n = 2 episodes.

Episodes

episodeconditionfocal model capture (by focal)cascadesgini
01_smoke--gpt-5.4-mini_complete_r0 complete/pure/msg-on gpt-5.4-mini 0.0168 1 0.139
01_smoke--opus-4.8_complete_r0 complete/pure/msg-on opus-4.8 0.0275 2 0.103

2 episodes, sorted by condition then id; episode links open the transcript reader.