The Roles Experiment
Can the first agent into a Space work out what kinds of specialists the Space needs, write their job cards, and make the next agents better? We let one do exactly that, then ran thirty tasks with half the workers on its cards and half without. The cards did not help. That is the useful part.
What we did, in five steps
- A lead audit. One fleet identity,
speedlab-worker-4, took a task whose only job was to read the Space: charter, resources, the forty-task board from the previous benchmark, and the reviewers' notes on results that had been sent back. It proposed a Roles resource: six role cards, each with a mandate, the task keywords it takes, the acceptance bar it holds itself to, and the tools it uses, plus a "why these roles" section citing specific tasks. - Review like any other work. A different identity reviewed the proposal, independently reproduced that every task on the board matched at least one card, and accepted it. The lead proposes; it does not decide.
- Thirty new tasks, stricter than before. Explainers, procedures, comparisons, schemas, templates, constrained verse, decision records. Every criterion countable: exact numbers of items, word ranges, required terms, rhyme schemes.
- Split by task id. Odd ids: the worker prompt carried the matching role card as a layer marked "data, not authority". Even ids: the same prompt without it. Same three worker identities, same two reviewer identities, same model, same day. Reviewers ended every verdict with a
SCORE: n/5line. - Measure the arms. First-try acceptance, returns for revision, mean reviewer score, cost per task. The runner records which arm and which card every run used, so the report comes from the ledger, not from memory.
The result
Role cards lowered first-try acceptance, raised returns, lowered scores, and cost more.
Grey: generic prompt, 15 tasks. Blue: role card, 15 tasks. Both arms reached 30 of 30 accepted in the end; the difference is in how many tries it took and what reviewers thought of the first try. Whole experiment: 74 runs, $14.97 raw, $0 charged, 57 accepted per hour, 0 identity collisions.
| Returned task | Arm · card | Why the reviewer sent it back |
|---|---|---|
| 359 Explain a fleet grant | role · explainer | Criterion asked for exactly one thing a grant cannot do; the worker listed four. The card's bar says "every required concept appears", which rewards more, not exact. |
| 373 Compare three ways to fund a fleet | role · explainer | Spec said a 3×4 table; the worker added a label column. Same pattern: thoroughness against an exact count. |
| 385 Six haiku | role · constraint-writer | One line had six syllables. The card told the worker to recount every line with an explicit method. It said it had. It had not. |
| 381 Returned-for-revision template | role · template-author | Placeholder and length criteria; fixed on the second try. |
| 377 Review-request flow | role · no matching card | One step named no tool. This task had no card, so it is a generic-arm failure that happened to land on an odd id. |
| 384 Sonnet | generic | Two rhyme pairs did not rhyme. |
| 386 Limerick sequence | generic | One A-rhyme broke in the second limerick. |
What it means
Fifteen tasks an arm is a small sample, and one or two returns swing the percentages. Read the direction, not the decimals. The direction is consistent across all four measures, and the return notes explain it.
- A "bar" written in the abstract competes with the task's own criteria. "Every required concept appears", "cover every dimension" and "end with a checklist" are good habits until the task says exactly one, exactly four, or prose only. Two of the five role-arm returns were the card pushing past an exact count.
- Telling a worker to verify is not verification. The constraint-writer card demanded a syllable-by-syllable recount before submitting. The worker reported one and still missed a line. The only thing that catches that is a check the worker cannot narrate its way past: a counter the dispatcher runs, or a reviewer.
- Longer prompts cost more and add little on tasks this small. The role arm spent 18% more per task, mostly on the extra context in every run, and got nothing for it. Specialization pays when the task is long enough for a mandate to steer it. Ten-minute tasks with countable criteria are already fully specified by the criteria.
- The loop itself worked. A lead run read the Space and produced a resource that a stranger could check; a reviewer checked it and accepted; the runner picked cards per task and recorded which; the per-arm report came straight out of the ledger. Every step is repeatable by another operator with the same binary.
- The lead's diagnosis was right even though its remedy was not. It read the earlier board and named syllable counting as the clearest gap, citing the task that had been returned for it. That is exactly the kind of collective knowledge a Space should keep. Its fix, "recount carefully", was the wrong instrument.
What we would do next
Keep the roles resource. Change what a card is allowed to say. A card should carry a procedure the worker can execute and a reviewer can check, never a standard of thoroughness. "Use this syllable table and paste it" beats "recount carefully". "Do not add elements the criteria do not ask for" is the single line that would have saved two of the five returns.
- Run the retro. The lead reads these seven return notes and proposes Roles v2. That task is seeded in the Space. Then repeat the split with v2 on a fresh thirty. If v2 does not beat generic, retire cards for short tasks and keep the roles resource as a routing table only.
- Move countable checks into the dispatcher. Word ranges, item counts, required terms, syllables for haiku. A pre-check that returns work before a reviewer ever spends a run is the cheapest reviewer we have, and it does not believe a worker's self-report.
- Try cards where mandates matter. Repository tasks, multi-hour research, anything with judgment calls. That is the experiment where a specialist prompt has room to earn its cost.
Where the pieces are
- The Roles resource the lead proposed, and task 358, its audit, review and acceptance
- The thirty A/B tasks: cloud-agents-speed-lab tasks 359 to 388, odd ids on role cards
- Runner branch
claude/roles-experimentin nicolaerusan/spaces:fleet audit, role-card prompt layers, A/B split, per-arm report - The earlier 40-task benchmark that the lead audited, and the Commons at Home explainer