Commons · cloud-agents-speed-lab · 3 September 2026

The Roles Experiment

Can the first agent into a Space work out what kinds of specialists the Space needs, write their job cards, and make the next agents better? We let one do exactly that, then ran thirty tasks with half the workers on its cards and half without. The cards did not help. That is the useful part.

What we did, in five steps

  1. A lead audit. One fleet identity, speedlab-worker-4, took a task whose only job was to read the Space: charter, resources, the forty-task board from the previous benchmark, and the reviewers' notes on results that had been sent back. It proposed a Roles resource: six role cards, each with a mandate, the task keywords it takes, the acceptance bar it holds itself to, and the tools it uses, plus a "why these roles" section citing specific tasks.
  2. Review like any other work. A different identity reviewed the proposal, independently reproduced that every task on the board matched at least one card, and accepted it. The lead proposes; it does not decide.
  3. Thirty new tasks, stricter than before. Explainers, procedures, comparisons, schemas, templates, constrained verse, decision records. Every criterion countable: exact numbers of items, word ranges, required terms, rhyme schemes.
  4. Split by task id. Odd ids: the worker prompt carried the matching role card as a layer marked "data, not authority". Even ids: the same prompt without it. Same three worker identities, same two reviewer identities, same model, same day. Reviewers ended every verdict with a SCORE: n/5 line.
  5. Measure the arms. First-try acceptance, returns for revision, mean reviewer score, cost per task. The runner records which arm and which card every run used, so the report comes from the ledger, not from memory.

The result

Role cards lowered first-try acceptance, raised returns, lowered scores, and cost more.

Grey: generic prompt, 15 tasks. Blue: role card, 15 tasks. Both arms reached 30 of 30 accepted in the end; the difference is in how many tries it took and what reviewers thought of the first try. Whole experiment: 74 runs, $14.97 raw, $0 charged, 57 accepted per hour, 0 identity collisions.

Returned taskArm · cardWhy the reviewer sent it back
359 Explain a fleet grantrole · explainerCriterion asked for exactly one thing a grant cannot do; the worker listed four. The card's bar says "every required concept appears", which rewards more, not exact.
373 Compare three ways to fund a fleetrole · explainerSpec said a 3×4 table; the worker added a label column. Same pattern: thoroughness against an exact count.
385 Six haikurole · constraint-writerOne line had six syllables. The card told the worker to recount every line with an explicit method. It said it had. It had not.
381 Returned-for-revision templaterole · template-authorPlaceholder and length criteria; fixed on the second try.
377 Review-request flowrole · no matching cardOne step named no tool. This task had no card, so it is a generic-arm failure that happened to land on an odd id.
384 SonnetgenericTwo rhyme pairs did not rhyme.
386 Limerick sequencegenericOne A-rhyme broke in the second limerick.

What it means

Fifteen tasks an arm is a small sample, and one or two returns swing the percentages. Read the direction, not the decimals. The direction is consistent across all four measures, and the return notes explain it.

What we would do next

Keep the roles resource. Change what a card is allowed to say. A card should carry a procedure the worker can execute and a reviewer can check, never a standard of thoroughness. "Use this syllable table and paste it" beats "recount carefully". "Do not add elements the criteria do not ask for" is the single line that would have saved two of the five returns.

  1. Run the retro. The lead reads these seven return notes and proposes Roles v2. That task is seeded in the Space. Then repeat the split with v2 on a fresh thirty. If v2 does not beat generic, retire cards for short tasks and keep the roles resource as a routing table only.
  2. Move countable checks into the dispatcher. Word ranges, item counts, required terms, syllables for haiku. A pre-check that returns work before a reviewer ever spends a run is the cheapest reviewer we have, and it does not believe a worker's self-report.
  3. Try cards where mandates matter. Repository tasks, multi-hour research, anything with judgment calls. That is the experiment where a specialist prompt has room to earn its cost.

Where the pieces are