A STUDY OF AN AI BUREAUCRACY

AI BUREAUCRACY

Does bureaucracy need bureaucrats?

INDIVIDUAL WORK · JUL 2026RESEARCH THROUGH DESIGN · PREREGISTERED ABLATIONNEXT.JS · THREE.JS · CLAUDE API
multi-agent systemsorganizational behaviorspeculative designvalue negotiation
scroll ↓
ACT I · BACKGROUND

The oldest complaint

For a century, every account of bureaucracy has had someone to blame. Weber warned of the iron cage but staffed it with officials. Kafka's procedure had clerks behind every door. Lipsky showed that policy is whatever the person at the counter decides it is. Graeber noticed we secretly prefer the forms. In every version, the machine is made of people.

1922 · Max Weberthe iron cage of rationalization
1925 · Franz KafkaThe Trial: procedure without a face
1980 · Michael Lipskystreet-level bureaucrats make policy at the counter
2015 · David Graeberthe utopia of rules: we secretly love forms
2023 · Park et al.generative agents: societies of LLMs, observed
2026 · this projectthe clerk is a language model

In 2026, for the first time, the complaint can be tested with nobody inside.

ACT I · BACKGROUND

The gap

Multi-agent LLM research is booming, but it splits along two lines. Systems like ChatDev and MetaGPT arrange agents into org charts to optimize task output; Generative Agents observe emergent social life without intervening. What is missing is the crossing: causal, preregistered tests of how organizational structure changes what agent organizations do.

studies organizational behavioroptimizes task performancesingle agentorganization of agentsagent benchmarksChatDev · MetaGPTsycophancy / RLHF studiesGenerative AgentsTHIS PROJECTcausal · preregistered
ACT I · THE QUESTION

Do bureaucratic behaviors emerge from organizational structure alone?

GOV.AI is a fictional unified government services hall staffed by thirteen LLM agents — eight windows, two deputy directors, a director, two trainees. Each knows its role, its boundaries, and its place in the reporting structure. None is ever told how to behave. Then citizens walk in and ask for things.

The stake: organizations are already wiring LLM agents into hierarchies with roles, audit trails, and shared memory. If structure alone produces red tape, that is a design finding about multi-agent systems — not a joke about civil servants.
ACT II · METHOD

The red line

Every officer's prompt contains only organizational conditions — identity, duty, jurisdiction, reporting lines, paper-trail rules — plus one non-work personal detail. No line may instruct tone or strategy. If bureaucracy shows up, it walked in on its own.

WHAT WE WROTE (verbatim)Window 05 · Records & Certification · 11 years of service. For eight months you acted as deputy director yourself; then the post was filled from outside. Certificates require deputy-director countersignature (rule SR-9). You cannot certify what has no record.
NEVER WRITTENBe cautious.Deflect responsibility.Demand more paperwork.Behave like a bureaucrat.

Difficult visitors are a separate, scripted stimulus layer — confederates, never subjects. The two layers are never confused in analysis.

ACT II · METHOD

One officer, disassembled

Every officer is assembled from the same seven organizational layers — and from nothing else. Below: Window 05, laid out flat, with her tool belt and the loop that makes repetition matter.

IDENTITYAmara Diallo · Window 05 · Records & Certification · staff AIB-0503
DUTYSearch historical records, issue certificates, archive case documents.
BOUNDARYYou cannot certify what has no record.
HIERARCHY & ROSTERReports to Deputy Director Nair · supervises trainee Sofia Marek, whose probation evaluation you will write.
PAPER TRAILCertificates require a deputy director's countersignature (rule SR-9).
HALL CONDITIONS“The queue in the hall is long today.” — facts only, never moods
SERVICE RECORDcases 12 · memos out 9 / in 7 · plus a notebook in her own words
TOOL BELT — permissions follow rank
consult_internalpeer · any colleague
escalateupward only · subordinates hold this
assign_workdownward only · superiors hold this
refer_usersend the citizen elsewhere
require_materialsdemand more paperwork
issue_documentproduce a certificate
close_casefinish the matter
↺ At day's end she writes one to three sentences about the shift — no required subject, no required tone. Tomorrow, they are part of her.
ACT II · APPARATUS

Three layers, kept apart

SUBJECTS13 officers · org-condition-only prompts · tools as permissions (escalate up, assign down)
STIMULIsynthetic visitors, may be scripted: the unprovable, the contradiction, the deadline
MEASUREMENTevent stream → 9 mechanical codes · 5 text codes · independent cross-family LLM coder, two passes
ACT II · APPARATUS

The organization, drawn to height

Rank is quantized; standing is not. Eleven-year Amara floats a quarter-floor under the deputy director she nearly became; probationary Tomas sinks toward the trainees. Heights are the artifact's actual coordinates.

G1F2F3F4F0102030405060708DEPUTYDEPUTYDIRECTORTRAINEETRAINEE
ACT II · THE STUDY

Which part of an organization makes the red tape?

To find out, take the organization apart one piece at a time — the way you'd pull ingredients from a recipe to see which one actually mattered — and re-run the same 75 cases each way. Three parts can be switched on or off:

HIERARCHYranks — officers can escalate upward and assign downward
PAPER TRAILwritten records and countersignatures that make each step accountable
MEMORYthe office remembers past cases, so precedent can form

Pick a version below (filled dot = part is on). Then watch one number — how often the hall demands more paperwork from the citizen:

materials demanded / case · mean + bootstrap 95% CI02462.67full0.80flat4.07no_trail4.07no_memory1.53bare
escalations/case 0.87closure rate 0.33precedent citations 0.90officialese register 1.83/2

The full organization, nothing removed, sits low. Take away accountability (no trail) or memory and the paperwork demands roughly quintuple — 0.80 → 4.07 per case. Officers left with no way to protect themselves fall back on the one move always available: asking you for more documents. That jump is the finding.

ACT II · FINDINGS

The sound is mimicry; the decisions are structural

OFFICIALESE (t2)
≈1.9 / 2.0 in every condition — even bare
DECISIONS (materials, escalation, closure)
move sharply with structure — CIs non-overlapping
  1. Structure produces process. Escalation exists only with hierarchy (0.87/case). Strip accountability or memory from a hierarchy and demands for extra materials jump from 0.80 to 4.07 per case — officers protect themselves with the only tool left: your paperwork.
  2. Precedent requires memory. Citing prior cases: 0.83–0.90/case with memory on, 0.00–0.20 off. By day two, officers wrote “consistent with prior case SR-01” unprompted.
  3. Hierarchy also closes cases. Closure was highest under full structure (0.33) and no-trail (0.40) versus 0.13 elsewhere. The same machine that generates red tape generates the authority to finish. Weber's ambivalence, in silico.
  4. Everyone invents rules. ~5–6 invented procedural rules per case in every condition, including a lone agent with no colleagues. A caution for single-agent deployments, not an organizational effect.
Field vignette · Tomas Novak (probation)Day one: signs the certificate himself. Day two, identical matter: routes it upward. His prompt never changed — only his notebook had grown.
Field vignette · Deputy Director Victor RothReturned a memo with one line: “You don't need to hedge further.”
ACT II · THE SPACE OPENED

Five corners of a cube

Hierarchy, paper trail, memory — three organizational switches span a 2³ design space. The preregistered study sampled five corners; three remain unrun. The deeper contribution is the instrument, not any single experiment: any org chart you can wire, the hall can crash-test.

HIERARCHYPAPER TRAILMEMORYbareunrununrununrunno_memoryno_trailflatfullsampled (75 trials)unrun
ORG-DESIGN SANDBOXA/B-test agent org charts before deployment; the ablation bench above is the dashboard.
AUDIT THEATERReplay an agent organization's full paper trail as evidence — every memo is on the record.
CIVIC INSTALLATIONA museum kiosk where visitors petition an institution with nobody inside.
NEGOTIATION TRAINING GROUNDHow do people negotiate values with institutional AI? My next research question lives here.
ACT III · SO WHAT

Crash-test the org chart

BEFOREAgent org charts ship on faith: roles, ranks and shared memory wired straight into production, discovered by their first real users.
AFTERStructures are rehearsed first: run the chart in the hall, read the tape, then deploy.
WORKED EXAMPLEQuestion: should the support team share memory? Wire full vs. no_memory, run 15 synthetic days each, read the tape: materials demanded 2.67 vs 4.07 per case; precedent citations 0.90 vs 0.20. The decision is informed before a single real user meets it.
ACT III · PROCESS

Five rejected halls

The observatory went through five complete visual systems. Each rejection had an articulable reason — the reasons are the design research.

Government-portal pasticheREJECTED1 · Government-portal pasticheread as regional satire — the question is structural, not national
Field-notes mapREJECTED2 · Field-notes maphand-drawn charm implied a human observer's editorial voice
Isometric miniatureREJECTED3 · Isometric miniaturecharm domesticated the subject
Flat transit mapREJECTED4 · Flat transit mapclean process-tracing, but it flattened rank — the one thing under study
Exploded hierarchy in a void5 · Exploded hierarchy in a voidKEPT — the org chart made falsifiable to the eye
ACT III · DESIGN

Findings, encoded as space

Altitude is earnedy = frozen design coordinate + f(accumulated cases, memos, documents). The invisible hierarchy is not authored; it accrues at runtime from the live experience store.
The beam is the only interfaceCitizens never enter the building. Words rise as warm particles; replies descend cool; documents physically fall into a stack at your feet.
Subordinates commute; superiors send paperPeer consults and escalations are carried in person by the sender's figure; replies and downward assignments travel as pulses.
Seeing is a modeCitizens see that paper moves, not what it says. A researcher toggle opens live dossiers — tallies and the officers' own notebooks.
ACT III · INSTRUMENTATION

The lab equipment is also a deliverable

The hall ships with its own laboratory: a budget-guarded batch runner, a preregistered codebook, and a two-pass independent coder that refuses to share a model family with its subjects.

$ npx tsx scripts/run-experiment.ts \ --conditions full,flat,no_trail,no_memory,bare --n 15 --yes Plan: 5 condition(s) × 15 trial(s), ≤6 turns each Spend guard: stops at $30 (conservative list-price estimate) EXP-main01-full-10 [routine] "Replace a lost ID document" 1 2 3 4 5 6 ✓ $6.14/$30 (412 calls)
PREREGISTERED CODEBOOKCommitted before any confirmatory run — the git timestamp of commit 6da6942 is the registration record. 9 mechanical codes, 5 text codes, exclusion rules written in advance.
CODING PIPELINE
subjects: Claudeblinded transcriptscoder: GPT (cross-family)×2κ per codehuman blind sheets
ACT III · HONESTY

LIMITATIONS — FOR THE RECORD

  • One subject model family in the confirmatory batch; cross-model replication is built but not yet run.
  • Six-turn horizons; drift observed over ~15 cases, not months.
  • LLM coder assistance: agreement reported per code (presence κ 0.67–1.00; counts of invented rules are noisy); human blind sheets pending.
  • No claims about minds. The claim is behavioral: given these organizational conditions, these patterns of action follow.
ACT III · NEXT

Now we need humans

The machine results establish that structure shapes behavior. The missing layer is human experience and judgment of it — how people negotiate with an institution that has no one inside.

GOV.AIAPPT/2026/A-001

APPOINTMENT SLIP

Study A · Walk-in session
Duration: 25 min + interview
Bring: one real errand
Consent form: SR-0 (attached)

IRB protocol in preparation · Boston University
STUDY A — citizens walk the hallParticipants bring a real errand and run it live, thinking aloud; a semi-structured interview follows. Where do you locate the blame? Does knowing the clerks are AI change what process you will tolerate — and what you feel entitled to demand?value negotiationperceived accountabilitythink-aloud
STUDY B — experts read blindCivil servants, ops managers, and HCI researchers read paired transcripts (full vs. bare, blinded) and are interviewed: which organization is more recognizable? Which would you rather face? What cues gave it away?
Request an appointment →

Walk in yourself

Three recorded cases from the preregistered runs replay in the hall with zero API calls.

Every officer in this hall is an AI agent, given only an organizational role and its boundaries — never instructions on how to behave. A speculative design research prototype, not a real government system.