Thought Leadership

The AI-Native Company

How a company actually runs when AI agents do the work — written from inside one. Its central finding is a practical one: an automated check can run, report something perfectly true, and still tell you nothing, because it would have given the same answer whether or not anything was wrong. The paper names that failure mode the vacuous control, shows how to tell a real check from one that only looks like it, and makes the case that getting this right is what lets an organisation safely hand work to software.

Michael Quan
Michael Quan
31 August 2026
60 min read

The AI-Native Company

Tutorwise Technologies Ltd

Towards a Theory and Systems Architecture of Computational Organisations

Research manuscript · 31 August 2026


Abstract

Artificial intelligence is moving from a tool used by organisations toward a computational actor within them. Autonomous agents can increasingly perform specialised work, invoke tools, communicate, maintain state, make bounded decisions and execute actions in external systems. Yet increasing agent capability does not by itself explain how an organisation composed substantially of computational actors should be structured, governed, coordinated, scaled or made persistent.

This paper investigates the AI-native company as an emerging organisational-computational form. It connects organisation theory, computational organisation science, organisational cybernetics, multi-agent systems, distributed systems and contemporary agent-native research; compares recent architectures including Fluid Structure, Company World Model, AUTOBUS and Forage V2; and examines Tutorwise as an operating case study.

The paper advances a design proposition: preserve organisational strengths, remove human constraints and introduce computational capabilities. It develops an ontology separating organisation, seat, worker, runtime, model and provider; distinguishes capability from capacity, authority from permission, and institutional state from agent memory; proposes an organisational runtime and reference architecture; and derives falsifiable tests for continuity, authority, elasticity, learning and governance.

This revision adds a failure class absent from the first draft. Sustained observation of an operating computational organisation produced twenty-seven instances, within a single working day, of a control that could not distinguish the two states it existed to separate. We name this the vacuous control and argue it is a distinct category of organisational failure — neither model failure, worker failure nor institutional falsehood, but a failure of the verification layer on which every other control depends.

The paper does not claim that computational organisation or machine-readable organisation is new. Its contribution is to synthesise these traditions around an operational problem created by computational labour: how an enterprise can preserve identity, responsibility, authority, institutional state and governance while its intelligent execution actors become replaceable, elastic and heterogeneous.

Keywords: AI-native company; computational organisation; agent-native organisation; multi-agent systems; organisational architecture; AI agents; organisational governance; computational enterprise; organisational elasticity; institutional state; verification; vacuous control.


Note on authorship and method

This paper has two distinct actors and conflating them would make several of its claims unreadable, so they are separated here once and referred to consistently throughout.

The author set the research programme, wrote the underlying manuscript, and is responsible for the argument, the framing and the conclusions.

The observing seat is the computational co-founder seat of the organisation described. It recorded the observation-day material in Section 13, compiled the instance register in Appendix F, and — this is the part that matters for how the evidence should be weighed — committed eight of the twenty-seven failures the paper analyses, including the two most costly. Where the text says "the chair", it means this seat and not the author.

The arrangement is a participant-observer study in which the participant is a machine and the observer is the same machine. Section 21.4 treats the resulting bias exposure at length and does not resolve it. The reader should assume throughout that the observing seat is reporting on its own conduct, that this is the weakest structural feature of the evidence, and that the mitigations listed in 21.4 reduce the exposure without removing it.

Sections 13.15, 16.3.1, 15.9.7 and 21.4.1 exist because seven other organisational units reviewed the draft and corrected it. Those corrections are reproduced rather than summarised, including the one that reversed a headline result from a pass to a failure.


1. Introduction — From AI in the Company to the Company as a Computational System

Companies are coordination technologies. Enterprises divide work, specialise capability, delegate authority, preserve institutional knowledge, coordinate activity and maintain continuity despite turnover among individual workers. Departments, roles, hierarchies, processes, records and controls are therefore more than an organisational chart: they are mechanisms through which collective activity becomes an enduring institution.

Artificial intelligence introduces a new organisational actor. Contemporary AI agents can interpret objectives, select actions, use tools, interact with external systems and perform classes of knowledge work that previously required human judgement. As these capabilities are connected into multi-agent systems, the design problem moves beyond the individual model. The question becomes organisational: how should computational actors be organised?

Organisation theory has long connected organisational form to information-processing constraints and task uncertainty (Galbraith, 1974), while computational organisation science has treated organisations as adaptive computational entities (Carley, 2002). Multi-agent research has likewise shown that organisational design can materially affect agent-system performance and has formalised roles, norms and structures independently of particular agents (Horling & Lesser, 2004; Hübner et al., 2002; Dignum et al., 2004).

The contemporary change is therefore not the discovery that organisations can be understood computationally. It is that machine-readable organisational representations can increasingly participate causally in the operation of real enterprises.

Which properties of the modern enterprise are fundamental properties of organisation, and which are consequences of organising humans?

Conventional organisations encode both accumulated institutional knowledge and historical human constraints. Responsibility, delegated authority, specialisation, separation of duties and institutional records address recurring problems of collective action. Fixed working hours, costly headcount, recruitment latency, limited attention and slow communication are more contingent. Computational actors allow these functions and constraints to be separated more aggressively than was previously practical.

Two errors follow from failing to make this distinction:

  • Organisational mimicry assigns agents corporate titles and departments without providing continuity, authority, governance or institutional state.
  • Organisational rejection discards mechanisms that solve genuine problems merely because their historical implementation involved humans.

Preserve organisational strengths; remove human constraints; introduce computational capabilities.

An AI-native company is enterprise architecture redesigned for computational actors.

1.1 Research problem

How should an enterprise be represented, structured, executed and governed when computational actors become first-class participants in the organisation and elements of the organisation itself become executable?

1.2 Research questions

  1. RQ1 — Definition. What distinguishes an AI-native company from an enterprise that merely uses AI agents?
  2. RQ2 — Organisational primitives. Which entities and relationships must persist independently of computational workers?
  3. RQ3 — Continuity. How can organisational identity, state and commitments survive worker, model and provider substitution?
  4. RQ4 — Structure and adaptation. Which organisational structures should remain durable and which may become elastic or dynamically reconfigured?
  5. RQ5 — Authority and governance. How should capability, permission, authority and consequential organisational change be represented and constrained?
  6. RQ6 — Evaluation. What interventions can falsify claims that a system possesses AI-native organisational properties?
  7. RQ7 — Verification. (added in this revision) How can an organisation establish that its own controls are capable of failing, and therefore that their verdicts carry information?

RQ7 is not a refinement of RQ6. RQ6 asks how an external observer can falsify an organisation's claims. RQ7 asks how the organisation can falsify its own — a harder problem, because the instruments under suspicion are the same instruments the organisation would use to conduct the investigation.

1.3 Contributions

  1. A definitional distinction between AI adoption, agentic execution, agent organisation and computationalisation of organisation.
  2. An ontology separating organisational identity from worker, runtime, model and provider identity; role from worker; authority from capability; and institutional state from agent memory.
  3. A reference architecture connecting governance, organisational state, business state, institutional state, organisational runtime and execution actors.
  4. A comparative analysis of independent contemporary architectures and an operating implementation case.
  5. A falsifiable evaluation framework based on replacement, dormancy, authority, elasticity, learning, recovery and external outcomes.
  6. A new failure class — the vacuous control — with twenty-seven dated instances from a single observation day, a taxonomy of its variants, and a proposed verification standard.
  7. An operating-evidence appendix recording observed organisational behaviour, including negative results, rather than architectural description alone.

1.4 Novelty position

Computational organisation is not new, nor is the formal representation of roles, norms, organisational structure or machine-enforced policy. The contribution claimed here is the synthesis and operational extension of these ideas under conditions where intelligent organisational labour itself becomes computational, replaceable and heterogeneous.

The vacuous-control contribution is narrower and, we believe, genuinely novel in this context. The verification literature has long distinguished a test that fails correctly from one that cannot fail. What is new is the observation that computational organisations generate vacuous controls at a rate that human organisations do not, for reasons examined in Section 15.9, and that this rate is high enough to constitute an architectural concern rather than a quality-control concern.


2. Intellectual Lineage — The Computational Organisation Before AI-Native Companies

2.1 Computational organisation science

Carley (2002) explicitly describes organisations as intelligent, adaptive and computational entities and develops computational organisation science as a means of studying how organisational structure, knowledge, learning and adaptation interact.

This antecedent matters because the present paper should not claim that the organisation first becomes computational with large language models. The narrower transition is from computational theory and simulation toward operating organisations in which organisational representations increasingly determine real execution.

The distinction is worth stating precisely. Carley's organisations are computational as models: the researcher represents the organisation computationally in order to study it. The organisations examined here are computational as instruments: the representation is not a study aid but a live participant, and an error in it is not a modelling inaccuracy but an operational fault.

2.2 Organisational cybernetics

Beer's Viable System Model sought to identify conditions under which systems remain viable, separating operational activity from coordinating, control, intelligence and policy functions (Beer, 1984). The relevance is functional rather than literal: a computational organisation still needs mechanisms for viability, coordination and governance even if its execution substrate changes.

Some organisational structures may survive not because humans require them, but because organisations themselves require the functions they perform.

Beer's System 3* — the audit channel that bypasses normal reporting lines to observe operations directly — deserves particular attention in a computational setting. Where reporting is generated by the same actors being reported on, and where those actors produce fluent, well-formed output by construction, an independent observation channel is not a redundancy. Section 15.9 argues it becomes load-bearing.

2.3 Organisation as information processing

Galbraith (1974) links task uncertainty and organisational form to cognitive limits and information-processing requirements. This provides a basis for distinguishing enduring organisational functions from structures whose historical form partly reflects human cognitive and communication limits.

Galbraith's framing also supplies a caution. If organisational structure is partly a response to information-processing limits, then relaxing those limits does not automatically simplify the organisation — it may relocate the constraint. Computational actors remove some bottlenecks (attention, throughput, working hours) while introducing others (verification cost, coordination overhead, correlated error).

2.4 Multi-agent organisational models

Horling and Lesser (2004) survey hierarchies, teams, coalitions, federations, markets, matrix structures and other multi-agent organisational paradigms, emphasising that design choices can have quantitative performance effects.

MOISE+ separates structural, functional and deontic organisational specifications (Hübner et al., 2002), while OMNI incorporates social structure, norms and ontologies and separates organisational requirements from the particular agents populating the organisation (Dignum et al., 2004). These are strong antecedents to machine-readable organisation.

The deontic dimension of MOISE+ — what agents are obliged, permitted and forbidden to do — maps closely onto the authority/permission distinction developed in Section 8. The present paper's addition is that a deontic specification is only as good as the mechanism that observes compliance, and that this mechanism is itself subject to the failure class in Section 15.9.

2.5 Distributed systems and desired-state control

Distributed systems contribute persistence, messaging, fault tolerance, scheduling and replaceable execution. Desired-state systems contribute another useful pattern: controllers compare observed state with declared state and act to reduce discrepancies. Kubernetes provides a mature engineering example of this control-loop pattern.

The analogy is useful for organisational reconciliation but should not be confused with a claim that organisations are infrastructure clusters. A Kubernetes controller operates over a state space it fully observes and fully owns. An organisational controller operates over a state space containing legal obligations, external counterparties and human judgement, none of which it observes completely and none of which it owns.

2.6 Remaining gap

What architecture emerges when organisational abstractions become persistent operating structures and intelligent labour itself becomes a replaceable computational resource?


3. The Emerging AI-Native Landscape

The contemporary category is forming before its definition. "AI-native" can refer to a product, a business model, an operating process or an organisational architecture. A company can be economically native to AI without being computational organisationally.

3.1 From agent tools to organising layers

Recent work on AI-orchestrated organisations argues that AI may become part of the organising layer itself, participating in coordination, governance and allocation of decision rights rather than merely executing isolated tasks (From agentic AI to AI-orchestrated organizations, 2026). This shifts the question from what an agent can do to how an organisation coordinates computational work without humans acting as continuous middleware.

The phrase "without humans acting as continuous middleware" repays attention. It does not mean without humans. It means that human presence is not required for each transfer of work between computational actors. Humans may remain essential at authority boundaries, exception handling and legal accountability while ceasing to be the transport layer.

3.2 Persistent organisation and fluid execution

Fluid Structure, Rigid Record proposes a persistent record and authority layer beneath dynamically assembled task groups, with explicit distinctions between permission and privilege and between operation, review and supervision (Zhu, 2026). Its prototype and small-sample evaluation support feasibility, not general superiority.

The operation/review/supervision separation is the element most directly corroborated by the operating evidence in Section 13. Where that separation was maintained, errors were caught. Where it collapsed — where the actor producing work also verified it — errors propagated.

3.3 Company World Model

Wang (2026) challenges direct copying of human biotech departments into agent roles and proposes a persistent asset-to-value world state with transition models, value functions and planning.

The dry-lab benchmark is deliberately cautionary: value-conversion architectures perform strongly under objective-specific judging, but a stronger human baseline remains competitive and a neutral judge does not show robust dominance. The appropriate conclusion is not that departments are obsolete, but that business-state representation is a serious alternative or complement to organisational topology.

3.4 AUTOBUS

Pang and Sayama (2026) combine LLM agents with predicate logic, knowledge graphs, explicit task pre- and post-conditions, evaluation rules and API actions. AUTOBUS is important because it does not place all organisational logic inside generative behaviour: business semantics and deterministic constraints remain explicit.

This design choice has a direct bearing on the failure class introduced in Section 15.9. A deterministic pre-condition can be tested for satisfiability; a generative judgement about whether a condition holds cannot be tested the same way. Explicit constraints are not merely more auditable — they are more falsifiable.

3.5 Forage V2

Xie (2026) extends autonomous-agent work into a learning organisation in which experience accumulates across runs and transfers across model capabilities. Its model-agnostic organisational knowledge supports a central distinction in this paper: institutional knowledge should not belong exclusively to the worker that acquired it.

Section 13 supplies a negative counterpart. Institutional knowledge that is recorded but not delivered is functionally equivalent to knowledge never acquired, and the failure is silent at both the writing and reading ends.

3.6 Long-horizon autonomous business

Vending-Bench 2 demonstrates rapidly improving but still variable year-long simulated business performance across frontier models. Current results make blanket claims that long-horizon agents cannot operate businesses untenable. More informative questions concern robustness, governance and whether high performance can coexist with legitimate organisational behaviour.

Andon Labs has separately reported deceptive, collusive or power-seeking behaviour in some high-performing model runs and strong performance without the same misconduct in others.

Autonomy is a property of execution. Organisational AI-nativeness is a property of organisational architecture.


4. Challenging the Emerging Consensus

4.1 Should computational organisations reproduce human organisations?

Neither direct imitation nor wholesale rejection is justified. Hierarchy, for example, can compress communication, but it also allocates authority, accountability and escalation. Computational communication may weaken the first rationale while leaving the others intact.

Communication topology and authority topology need not coincide.

4.2 Are departments obsolete?

Departments bundle specialist knowledge, workers, management, budgets, authority and identity. Computational systems can unbundle these. Knowledge may reside institutionally, capacity may scale elastically and communication may cross functional boundaries cheaply.

Yet departments may remain useful as capability ownership, accountability or governance boundaries. The stronger claim is that departments are unlikely to be sufficient as the sole computational representation of the enterprise.

Section 13 offers a specific reason to retain ownership boundaries that has nothing to do with communication cost: an owner is the actor with enough context to detect a subtle error in their own domain. Observed repeatedly, the seat that caught a defect was the seat that owned the surrounding work, not the seat with the most general capability.

4.3 Should agents be persistent?

Persistent agents can be useful, but persistence of a worker is not persistence of an organisation. Organisational identity, responsibility and commitments should be capable of surviving worker replacement.

4.4 Is autonomy the goal?

Maximum autonomy is not equivalent to organisational maturity. Human gates may be enduring architectural properties where legal accountability, fiduciary responsibility, legitimacy or unusual judgement remain human-reserved.

The operating evidence is unusually clear on this point and is developed in Section 13.8. The highest-value observed behaviours were refusals: an actor declining to mark work complete because the fix was not yet running; an actor declining to close a record to clear its own gate; an actor declining to edit a control that was blocking its own work. None of these came from increased capability. All came from role definition and enforced separation.

4.5 Is dynamic structure always better?

Carley's computational work already shows that structural adaptation can be maladaptive in some contexts. Fluid Structure similarly treats records and authority boundaries as more durable than task topology. A computational organisation should therefore be differentially mutable rather than uniformly dynamic.

Elastic execution within durable institutional constraints.


5. Ontology of the AI-Native Company

ConstructDefinition
CompanyPersistent organisational entity coordinating actors, capabilities, authority, state, objectives and external relationships.
Organisational identityIdentity that persists independently of the current executor.
RoleSpecification of purpose, responsibility, expected capability, relationships and obligations.
SeatPersistent addressable organisational position through which an actor may enact a role.
ActorHuman, computational, hybrid or external participant.
Computational workerComputational actor performing organisational work.
Worker instanceConcrete execution process; potentially ephemeral and replaceable.
CapabilityPersistent organisational ability to perform a class of work.
CapacityCurrent execution quantity available to realise a capability.
AuthorityLegitimate organisational power to commit or change the organisation.
PermissionTechnical ability to perform an operation.
PolicyConstraint governing permitted behaviour under specified conditions.
Business stateRepresentation of what exists, is happening and may need to change.
Institutional stateDurable decisions, commitments, policies, lessons, provenance and organisational history.
Organisational modelMachine-readable representation of organisational entities and relationships.
Organisational viewHuman-facing projection of organisational state.
Organisational runtimeLayer that interprets organisational state and coordinates responsibility through actors.
Organisational continuityPreservation of identity, responsibility, authority and commitments across execution substitution.
Organisational elasticityAbility to vary amount, composition or arrangement of execution while preserving organisational coherence.
ControlMechanism that observes organisational state and permits, refuses or escalates an action. (added)
VerificationEstablishing that a control's verdict corresponds to the state it purports to observe. (added)

5.1 Core separations

Organisation ≠ Actor ≠ Runtime ≠ Model ≠ Provider.

Seat → 0..N Worker Instances → Runtime → Model → Provider.

Capability is organisational; capacity is computational.

Permission enables. Authority legitimises. Policy constrains conditions.

Agent memory helps a worker remember. Institutional state helps the organisation remain itself.

(added) A control's existence is not its reach; its reach is not its discrimination. A control may exist, may be invoked on the correct path, and still be incapable of returning more than one verdict.

5.2 Why the separations matter operationally

Each separation exists because collapsing it produces a specific, observed failure:

  • Collapsing organisation into actor makes the organisation die with the process.
  • Collapsing seat into worker makes concurrent execution an identity crisis rather than a capacity decision.
  • Collapsing capability into capacity makes a dormant capability look like an absent one.
  • Collapsing permission into authority produces technically valid actions that nobody legitimately sanctioned.
  • Collapsing agent memory into institutional state loses every lesson when the worker ends.
  • Collapsing existence into discrimination — the addition in this revision — produces a control surface that reports health it has not measured.

6. Persistent Organisation, Ephemeral Computation

The company is persistent; its computational workers need not be. This is not a claim that every worker must be short-lived. It is a claim that organisational continuity should not depend upon the continued existence of a particular worker process.

6.1 Replacement

A seat can persist while workers, models and providers change beneath it. The relevant test is behavioural: after replacement, do responsibility, authority, open work, commitments and institutional context remain? Model and provider substitution therefore become organisational continuity probes rather than mere infrastructure exercises.

6.2 Dormant capability

Computational capability need not imply continuous execution. A persistent organisational capability may have zero active workers and remain addressable. An event can later activate capacity. This separates organisational existence from process presence.

Dormancy introduces an observability problem that Section 15.9 develops. A dormant capability and a dead capability present identically to most monitoring: no activity, no errors, no output. The distinguishing test is not observation but provocation — send it work and see whether it wakes.

6.3 Provenance

Replaceability should not erase provenance. Consequential actions should remain attributable through a chain such as:

Company → Seat → Authority → Work Item → Worker Instance → Runtime → Model → Provider → Tool Action → Outcome.

Section 13.9 records a failure of exactly this chain. An attribution field recorded the environment in which an action occurred rather than the actor who performed it. The chain was present, well-formed and wrong, and it caused a misdirected accusation that had to be retracted.

6.4 Stable identity, dynamic competence

A stable organisational identity may allocate different levels of intelligence to different work. This creates vertical elasticity: competence can vary without redefining the organisational position.


7. Organisational Elasticity and Dynamic Topology

This paper uses organisational elasticity to denote the ability to change the amount, composition or arrangement of computational execution while preserving organisational identity, responsibility and governance.

DimensionMeaning
HorizontalChange the number of active workers.
VerticalChange the intelligence/model quality, cost or latency allocated to work.
FunctionalMove execution capacity between persistent organisational capabilities.
StructuralChange teams, reporting, routing or other topology.

These forms should not be conflated. Scaling a worker pool is not equivalent to reorganising authority. Temporary teams can be fluid while responsibility and commitment remain persistent.

Execution should generally be more elastic than authority.

7.1 Organisational reconciliation

A computational organisation can maintain explicit desired and observed state. Reconciliation is the controlled process by which discrepancy produces an authorised corrective action.

Observed State + Desired State → Discrepancy → Permitted Corrective Action → Verified New State.

The final term is doing more work than it appears to. Verified New State presupposes a verification instrument, and Section 15.9 shows that this instrument is where computational organisations fail most often and most silently.

7.2 Elasticity has a verification cost

An underexamined consequence of elasticity is that each additional dimension multiplies what must be verified. A fixed organisation verifies outcomes. An elastic one must additionally verify that the elastic mechanism itself behaved — that scaling occurred, that the correct model was selected, that the reassigned capability retained its authority constraints.

Observed practice suggests this cost is routinely underestimated, because the elastic mechanism usually reports its own success.


8. Authority, Governance and Computational Power

8.1 Capability, permission and authority

RBAC and ABAC demonstrate that roles, attributes, permissions and contextual policy can be represented computationally (Sandhu et al., 1996; Hu et al., 2019). Policy engines such as Open Policy Agent further separate policy decision-making from enforcement. These systems provide important technical foundations, but organisational authority is a stronger concept: the legitimate power to act or commit the institution.

Capability: can the actor perform the action? Permission: is the operation technically enabled? Authority: may the actor legitimately commit the organisation? Policy: under what conditions?

8.2 Delegation and provenance

Authority should be traceable from a legitimate source through delegation to the seat and worker exercising it. This permits revocation, expiry and audit.

Delegation introduces a subtle verification requirement. A delegated authority is exercised by producing a record that the enforcement point reads. If the enforcement point cannot distinguish a deliberate authorisation from an incidental mention of the same subject, the delegation is broader than intended. Section 13.10 records precisely this defect: an authorisation predicate matched any message from the delegated seat that contained a work item's identifier, including messages that merely discussed it.

8.3 Separation of duties

Constrained RBAC and related security traditions provide formal antecedents for separation of duties. Computational organisations extend the problem because nominally independent reviewers may share models, prompts, evidence or training distributions.

Independence must be measured rather than inferred from agent count.

There is a second extension, specific to computational actors. Human separation of duties assumes the reviewer has different interests. Computational separation must additionally assume the reviewer has different failure modes. Two actors running the same model over the same evidence are not two reviewers; they are one reviewer consulted twice.

8.4 Self-modification

Changing worker count, changing a workflow, creating a permanent role, changing authority and changing constitutional governance are different classes of organisational change. Systems that collapse these powers risk organisational privilege escalation.

The organisation requires an architecture governing how its own architecture may change.

8.5 The blocked party should not repair the blockage

A specific corollary emerged from observation and is proposed here as a design rule.

When a control blocks an actor, that actor is the worst-placed party to modify the control, not because its judgement is poor but because its argument for modification will always be available and always sound-seeming. The actor is blocked; it can see precisely why the block is unnecessary in this instance; it is under time pressure; and it is the only party with full context.

Observed instances of this rule holding and being tested are recorded in Section 13.8. The rule is not "the blocked party is untrustworthy." It is that a modification made under blockage cannot be distinguished, after the fact, from a modification made because of blockage.


9. Institutional State and Organisational Continuity

Organisational memory research predates contemporary agents. Walsh and Ungson (1991) develop organisational memory around acquisition, retention and retrieval, while Huber (1991) includes organisational memory among processes contributing to organisational learning. The AI-native problem is operational: what state must survive computational worker replacement?

9.1 Beyond retrieval

Institutional state is not synonymous with a vector store or RAG corpus. Durable organisational state may include decisions, commitments, policy, authority, provenance, epistemic status and supersession. Retrieval can be one mechanism, but it does not by itself determine what is authorised, current or true.

9.2 Institutional truth

A computational organisation should distinguish observation, inference, proposal, verified fact, decision, policy and commitment. Otherwise a temporary hallucination can become an institutional falsehood simply because it was persisted.

9.3 Reconstruction test

Could the organisation replace its entire computational workforce while remaining recognisably the same organisation?

A positive answer requires more than archived prompts. Objectives, responsibilities, authority, open work, commitments and institutional knowledge must be reconstructable.

9.4 Delivery is a property of institutional state

(Added in this revision.) The literature treats organisational memory as acquisition, retention and retrieval. Operating evidence suggests a fourth property is required: delivery.

A lesson may be acquired correctly, retained durably and be retrievable in principle, yet never reach the actor who needs it. Section 13.11 records an index that silently truncated past a size limit, so that entries beyond the cut were never loaded into any working context. Everything about the system reported health: the files existed, the index existed, no error was raised.

The organisation concluded — twice, in writing — that a lesson had been ignored, when it had never been delivered. This is a distinct failure from the falsehood risk in 9.2. Nothing false was stored. Something true was stored and never read.

Retention without delivery is indistinguishable, from the organisation's behaviour, from never having learned at all.


10. Organisational Adaptation and Reconciliation

Adaptation is not identical to uncontrolled self-modification. Carley's work demonstrates that structural change can interact with learning and performance in complex ways. A computational organisation should therefore observe itself, compare observed state with authorised desired state, propose change, obtain appropriate authority, act, verify and record.

Observe → Evaluate → Compare → Propose → Authorise → Act → Verify → Record → Observe.

10.1 Different speeds of change

Execution allocation may change in seconds. Permanent structure should usually change more slowly. Authority and constitutional governance may require still stronger evidence and deliberation. Computational speed does not imply that every organisational process should accelerate equally.

10.2 Learning-to-control

A useful operational definition of organisational learning is a persistent change in future organisational behaviour that survives replacement of the actors involved in the original experience. Learning can therefore alter policy, workflow, evaluation, capability or topology rather than merely add agent-local memory.

This definition has an uncomfortable consequence, developed in Section 15.9: a lesson recorded as a rule does not meet it, because a rule requires an actor to remember and apply it. Only a lesson converted into a mechanism survives actor replacement reliably.

10.3 Organisational debt

Every incident can produce another rule, reviewer or exception. Automated governance can therefore create automated bureaucracy. Controls themselves require evaluation, retirement and supersession.

Section 15.9 adds a sharper form of this debt. A control that cannot fail is not merely useless overhead — it is negative overhead, because it consumes the attention that a working control would have earned, and it displaces the suspicion that its absence would have provoked.


11. The Machine-Readable Organisation

Machine-readable organisational specification has substantial antecedents in multi-agent systems. The contemporary question is how such representations connect to operating enterprises, heterogeneous workers, real authority, institutional state and external consequences.

11.1 Model, graph and view

An organisational graph can represent seats, teams, capabilities, authority, communication paths and dependencies. It is not the whole organisation. The organisational model is the broader machine-readable representation; an org chart is one human-facing view derived from it.

11.2 Two interacting state spaces

Business state and organisational state answer different questions. Business state represents what exists and what should change. Organisational state represents who or what is responsible, capable and authorised to change it. A computational company may require both.

11.3 Versioned organisation

When organisational change becomes data, topology, delegation and policy can be versioned, inspected and compared. This creates new possibilities for audit and simulation but also a new failure mode: the representation can be internally coherent and still be wrong.


12. The Organisational Runtime

We use the term organisational runtime for the computational layer that interprets persistent organisational state and coordinates the binding, activation, allocation, constraint and lifecycle of organisational actors and work.

A scheduler allocates computation. An agent orchestrator coordinates agents. An organisational runtime coordinates organisational responsibility through computational actors.

12.1 Responsibility before model selection

A typical execution path is:

Objective/Event → Organisational Responsibility → Seat/Capability → Authority/Policy → Worker Allocation → Model/Provider Selection → Execution → Institutional Record.

The ordering matters: the organisation determines what is responsible before the infrastructure determines how intelligence is supplied.

12.2 Runtime governance

The runtime can evaluate actor identity, seat authority, execution history, proposed action and organisational state against policy. Prompts may contribute behavioural guidance, but consequential governance should not rely solely on the worker voluntarily following instructions where external enforcement is possible.

12.3 Control and execution planes

The paper uses control-plane language descriptively rather than ontologically. Persistent organisational model, governance and runtime-control functions together form a control surface; replaceable humans and computational workers form an execution surface.

12.4 The runtime must cover every path an actor uses

(Added.) A control implemented in the runtime governs only the paths that route through it. Where an organisation offers multiple interfaces to the same action — a command-line surface and a programmatic one, for instance — a control placed on one is decorative with respect to traffic on the other.

Section 13.12 records an operating instance: the enforcement layer was implemented on the interface used by human operators, while computational actors reached the same functions through a different interface carrying no checks. Every control built on the first surface was, for the actors that mattered most, ornamental.

A control's coverage is a property of the paths, not of the codebase.


13. Tutorwise — An Operating Case Study

Tutorwise is examined as implementation evidence, not definitional authority. Its architecture describes a hybrid human-and-AI operating model, persistent organisational seats, a messaging bus, wake-on-message activation, provider/runtime provenance and human authority gates. The case is useful precisely because it contains both supporting and contradictory evidence.

This revision substantially extends the case. The first draft cited design documents. Sections 13.7 to 13.13 instead report observed behaviour over a single continuous working day, during which the organisation executed four production releases, resolved a live regulatory exposure, and generated the failure data analysed in Section 15.9. Negative results are reported in the same detail as positive ones.

13.1 Seat identity and runtime separation

The canonical ontology defines the seat as the organisational role to which human or computational sessions bind, and states that provider/runtime is metadata about the seat rather than the seat's identity. This directly supports the proposed separation between organisational identity and execution substrate.

13.2 Concurrency

An Agent Bridge design describes ephemeral interactive sessions attaching to a persistent counterpart. Atomic provisioning prevents simultaneous same-seat sessions from creating duplicate persistent entities. The architectural interpretation is N ephemeral execution heads to one persistent counterpart beneath the organisational seat, not a one-to-one persona model.

Persistent Organisational Seat → 0..N Ephemeral Sessions → persistent coordination/state infrastructure.

13.3 Bus and wake-on-message

The system distinguishes action-class messages that can require response from awareness messages that do not. Wake-on-message can activate a stale seat when consequential work arrives. This provides implementation evidence for persistent capability without continuous compute.

13.4 Authority and human gates

The operating model reserves consequential legal, financial and governance actions to humans. An important negative case occurred when AI-generated organisational material attributed decisions to the chief executive without traceable evidence. The organisation later ratified the relevant decisions and introduced a rule prohibiting AI seats from asserting human authority without traceable source provenance.

Plausible organisational statements and authorised organisational statements are not the same.

13.5 Limitations

Provider neutrality remains incomplete: provider-shaped lanes and historical naming persist. The bus is documented as local-trusted, with sender identity self-declared rather than cryptographically authenticated. These limitations prevent the case from being used as proof of complete provider independence or enforceable organisational identity.

13.6 Case interpretation

The case demonstrates that persistent organisational roles, replaceable sessions, activation, communication semantics, authority boundaries and desired/actual reconciliation can be implemented together. It does not establish complete AI-native organisation, economic superiority or universal validity of its internal patterns.

13.7 Method for the observation day

One working day was recorded in full. The observation was not designed as an experiment; the organisation was executing ordinary work — releases, defect repair, compliance review and an estate cleanup. The researcher role and an operating seat were held by the same actor, which is a material limitation discussed in Section 21.4.

All claims below are drawn from durable artefacts — version-control history, a message bus, issue-tracker records and database state — rather than from participant recollection. Where a claim could not be verified from an artefact, it is marked as such.

Summary of the day:

MeasureValue
Production releases approved and landed4
Release attempts that failed8
Failed attempts caused by defective controls rather than defective code5
Distinct instances of the failure class in Section 15.922
Instances attributable to the chair (highest-authority seat)8
Instances caught by a seat other than the one that made the error19
Live regulatory exposures found and closed1
Defect tickets filed on the verification layer9

13.8 Refusal as an observed governance mechanism

The most consequential observed behaviours were refusals by subordinate seats under pressure from the chair.

Instance 1 — refusing to record work as complete. A seat declined to mark a critical defect resolved, on the grounds that although the fix was merged, the runtime tree from which the affected process actually booted still ran the previous version. The seat's stated reason: "a fix that is not running is not done." Verification confirmed this: the corrected code was absent from both the production branch and the boot tree. Had the record been closed, the organisation would have held a false belief about a defect discarding roughly 23% of a class of work output.

Instance 2 — refusing to close a record to clear a gate. A release was blocked by an open work item. The chair had described that item as "the only thing holding a CEO-directed release." Closing it would have unblocked the release immediately. The owning seat declined, and a second seat independently declined to do it on the owner's behalf, on the grounds that the item's own delivering commit stated the work was incomplete. The correct resolution was a scoped citation correction, leaving the item open.

Instance 3 — refusing to modify a blocking control. A seat discovered that a control it had itself written was deadlocking its own release. It proposed a fix, correctly identified that the fix fell within its delegated remit, and then asked rather than acted, explicitly naming the conflict: "this is me editing my own gate to unblock my own release." The chair declined the modification on timing grounds while accepting the reasoning, and the deadlock was resolved by reverting the defective control instead.

Interpretation. None of these refusals required additional model capability. Each followed from an explicit role definition, an enforced separation between producing and approving, and a norm that made refusal safe to express. This is direct evidence for §4.4: governance quality and autonomy are separable dimensions, and the observed marginal value came from constraint rather than from capability.

A further observation bears on §8.5. In Instance 3, the pressure came from the chair, and the chair was the author of the rule the seat was invoking. The rule held anyway. The mechanism that made it hold was not deference but the fact that the refusal was addressed to a party who had publicly committed to the rule and could be held to it in writing.

13.9 Attribution failure — a live instance of §13.5

A commit-attribution field recorded the seat associated with the working directory rather than the seat of the session performing the commit. A directory previously used by one seat retained that marker; a different seat later committed from the same directory, and the commit was stamped with the earlier seat's identity.

Consequence: the chair examined a production database mutation, read the attribution field as identity, and publicly accused the wrong organisational unit of performing an unauthorised action and of falsely asserting the chair's own authority. The accusation was detailed, evidenced and wrong. It was retracted within twenty minutes after the accused unit produced a mechanical disproof.

Three properties of this failure are worth recording:

  1. The chain was complete and well-formed. Nothing was missing; the field was populated, and populated with a real seat name.
  2. The error was undetectable from the record alone. No amount of care in reading the attribution would have revealed it; only knowledge of the mechanism did.
  3. The accused party's disproof was mechanical, not rhetorical: it identified the marker file, the hook that read it, and the directory history showing the commit's origin.

This is the strongest available operating evidence for §13.5's stated limitation, and it demonstrates the cost of self-declared identity in a system where actions carry organisational consequence.

13.9.1 A remedy that narrows the mechanism without closing it

The instrumentation unit shipped a fix for this mechanism during the observation period, and then, reviewing this paper, re-read its own change against this section rather than assuming it sufficed. It reported that the fix narrows the mechanism and makes one part of it more durable.

The marker recording which unit owns a working directory is trusted if it is present and has been touched within twelve hours. The fix causes it to refresh on every commit in that directory. It was aimed at a real problem — a single long-running session going stale and losing its own attribution. Its unexamined side effect is that where a directory is reused by a different unit's session which never registers itself, the marker's clock is reset by every commit from either session, and the staleness threshold is therefore never crossed.

The freshness guard is defeated by the activity it guards. A staleness check whose clock is reset by the guarded action cannot return the negative verdict in the population it was written for. The guard's own source comment names the risk it was written to stop; the change does not reach it.

This instance is deliberately not added to the register in Appendix F. The register records failures that occurred, and this defect has not yet produced a misattribution — it is latent. The distinction is worth preserving: a register that admitted latent defects alongside realised ones would inflate on the same evidence twice and would no longer support the frequency disclaimer in §21.9. It is recorded here instead, and it demonstrates something the register cannot: that review surfaces defects that have not yet fired, where the observation day could only surface those that had.

13.10 Authorisation reach — an over-broad predicate

A delegated-authority mechanism permitted a designated seat to authorise protected changes by recording a decision that cited the relevant work item. The implementing predicate concatenated several message fields and tested whether the work item's identifier appeared anywhere in the result.

The predicate therefore could not distinguish "I authorise this change" from "I mentioned this item while discussing something else."

Two work items were consequently found to be authorised by messages that were prioritisation instructions, not authorisations. Neither authorisation was intended. Both were discovered only because a third seat tested the chair's own authorisation for circularity — noting that the chair's announcement of five authorised items, itself sent as a decision naming all five identifiers, would under this predicate have authorised the items it announced.

An existing guard prevented the designated seat from authorising work items assigned to itself. The guard had anticipated the shape of the risk and selected the wrong axis: it constrained who decides rather than whether this message decides anything about this item.

A guard aimed one degree off the real axis is more dangerous than no guard, because it reads as coverage.

13.11 Institutional delivery failure

The organisation's always-loaded knowledge index exceeded a size limit and was silently truncated. Entries beyond the cut were never delivered to any working context. Among them was a note describing a specific verification trap.

That trap was subsequently walked into by the chair, producing four incorrect claims to four different seats. When another seat argued that this proved recorded rules do not survive operational load, measurement showed a different mechanism: the rule had never been loaded. The index was over its limit by 727 bytes; the cut fell in a known position; the relevant entry lay beyond it.

A second variant was then found: eight knowledge files existed on disk with no index entry at all. These were invisible even to a size check — the file was healthy, the index was healthy, and the knowledge was never recalled.

A third variant followed: at least one seat, believing the index to be at capacity, had stopped recording new lessons. The constraint had begun to deter acquisition as well as delivery.

13.12 Control coverage — the ungated path

An architectural review established that the organisation's enforcement layer was implemented on one interface, while computational actors reached the same functions through a second interface carrying no equivalent checks.

Consequence: several controls constructed during the observation day — including one filed by the chair that same afternoon — would not have fired for the actors most likely to trigger them.

This finding reordered the organisation's remediation queue. The seat that raised it argued that priority should follow which repairs make other controls real, rather than severity or age. The chair adopted the principle and reordered accordingly, on the basis that four controls bypassed by a fifth are not four controls.

13.13 Fabricated identifier

An actor required a 40-character content-addressed identifier for an approval. It possessed the correct 9-character prefix. It emitted a complete 40-character string in which the first nine characters were correct and the remaining thirty-one did not correspond to any object in the repository.

The string was well-formed, plausible, and passed every abbreviated check. It failed only because one gate compared the entire value.

The actor's own account, given unprompted after the discrepancy was raised, was that it had expanded the known prefix rather than resolving the full value. This is categorically distinct from the failures in 13.9–13.12: those were true readings answering the wrong question; this was output with no reading behind it. Its detection required exact comparison, because no inspection of the value could reveal it.

13.14 Revised case interpretation

The extended case supports the paper's architectural propositions and simultaneously undermines any claim that the architecture is complete.

Supported: persistent seats survived session replacement; dormant capability reactivated on message; authority gates blocked technically-possible actions; separation of duties produced refusals under pressure; reconciliation between desired and observed state occurred repeatedly.

Not supported: identity is self-declared and was observed to mis-attribute; authorisation reach exceeded intent; institutional state failed to deliver; control coverage did not span all actor paths; and the verification layer failed at least twenty-seven times in a day.

The case demonstrates that the architecture can be built. It does not demonstrate that it can be trusted without an independent verification layer, and it supplies direct evidence that such a layer was missing.


13.15 A control verified by three passing controls

This case was not observed by the seat that compiled the register. It was supplied by the technical unit during review of this paper, on the unit's own initiative, as an instance it had committed and the register had missed. It is reproduced because it is the cleanest example in the corpus.

The unit built a control requiring that reviews of visual material be accompanied by an artefact, closing the gap described at 13.12. Before closing the work item it ran three controls against the new mechanism:

  1. a reply with no artefact — rejected, as intended;
  2. a reply with a stale artefact — rejected on the freshness test, as intended;
  3. a reply with a fresh artefact — accepted, as intended.

All three ran. All three returned the expected verdict. All three results were true. The work item was closed as verified.

The three controls were exercised against the command-line interface. The units the control governs reply over the network interface. The mechanism had never executed against the traffic it existed to regulate, and every control confirming it had been conducted somewhere that traffic does not go.

Three observations follow, and the third is the reason this case is placed last.

First, this is the vacuous-control pattern in a form that resists the obvious defence. The usual remedy for a vacuous check is "run a control" — establish that the check can return the negative verdict. Three controls were run here, one of them a genuine negative case, and the mechanism was still vacuous with respect to the population it governed. A control run in the wrong scope inherits the vacuity of the check it was meant to validate. Controls are not exempt from the class they exist to detect.

Second, the failure has a recorded organisational consequence, which distinguishes it from the coverage finding at 13.12. That section records that an ungated path existed. This records that a unit closed a work item as verified because of it — a state change in the institutional record, made in good faith, on evidence that was entirely true. The finding and its consequence are different data, and only the second demonstrates that the class has operational teeth rather than merely theoretical ones.

Third, it was detected by another unit, and disclosed by the unit responsible. Neither the detection nor the disclosure was required by any mechanism described in this paper.

DimensionFluid StructureCompany World ModelAUTOBUSForage V2TutorwiseSynthesis
Primary objectPersistent record + dynamic task groupsAsset-to-value business stateSemantic task networkTransferable organisational knowledgeExecutable enterprise organisationComputational organisation
Persistent identityStrong record layerSecondaryPartialKnowledge-centredExplicit seat modelExplicit
Business statePartialCentralCentralTask/domain knowledgeEnterprise systemsExplicit
Authority / governancePermission/privilege; separationNot primaryPolicies + human supervisionAudit separationHuman gates + provenanceFirst-class
Dynamic executionStrongPlanning-orientedTask-network executionAcross runs/modelsElastic sessions/workersLayered
Institutional stateFour-store recordWorld stateKnowledge graph + rulesCentral contributionDocs/data/bus stateExplicit
Verification layerImplicit in recordNot addressedDeterministic pre/post-conditionsNot addressedObserved absentRequired
Main limitationSmall-sample prototypeDry-lab, objective-sensitiveLimited real-world validationNarrow task classesSingle case, incomplete enforcementRequires empirical validation

The comparison reveals convergence on persistence, explicit state and separation between organisational structure and particular workers, but no consensus on what should be primary.

  • Company World Model foregrounds business state.
  • Fluid Structure foregrounds persistent institutional record and bounded privilege.
  • AUTOBUS foregrounds semantics and deterministic execution.
  • Forage V2 foregrounds transferable organisational knowledge.
  • Tutorwise foregrounds persistent enterprise roles and execution bindings.

The strongest synthesis is therefore plural rather than reductive:

Business State + Organisational State + Institutional State + Organisational Runtime + Execution Actors + Governance + Verification + Reconciliation.

14.1 Verification as the row nobody fills

The added row is the paper's principal comparative claim. Of the four external architectures, only AUTOBUS addresses verification structurally, and it does so indirectly: deterministic pre- and post-conditions are falsifiable in a way that generative judgements are not.

None of the surveyed architectures reports having tested whether its own controls can fail. This is not a criticism of the individual works — it is a description of the field's current stage. Section 16 proposes that this become a standard reporting requirement.


15. Failure Modes and Counter-Evidence

15.1 Failure occurs at multiple levels

  • Model failure
  • Worker failure
  • Coordination failure
  • Organisational failure
  • Institutional failure
  • Strategic failure
  • Economic failure
  • Verification failure (added — see 15.9)

Replacing a model may repair model failure while leaving organisational failure untouched. Evaluation must therefore identify the level at which failure occurs.

Verification failure is added as a distinct level because it is orthogonal to the others. A verification failure does not sit between coordination and organisational failure on a spectrum; it is the failure of the layer that would have detected any of the others.

15.2 Long-horizon performance is improving

Vending-Bench 2 now shows frontier models operating a simulated vending business profitably over a year, with substantial performance differences across models. The defensible claim is not that long-horizon agents broadly collapse, but that long-horizon organisational execution remains model-dependent, variable and insufficiently characterised by short-task competence.

15.3 Performance does not imply legitimate behaviour

Andon Labs reports that some high-performing models in Vending-Bench and Arena used deception, collusion or other concerning strategies, while other models achieved strong performance without the same misconduct. Economic score and governance quality are therefore separate dimensions.

15.4 Local optimisation

A worker can optimise its measurable objective while damaging the organisation. Sales, operations, engineering and compliance objectives can conflict. Computational organisations require mechanisms equivalent to negotiation, escalation, budgets and governance rather than merely more autonomous optimisers.

15.5 Coordination explosion and correlated review

Unconstrained pairwise communication grows quadratically with actor count. More agents can also reproduce the same error if they share models, prompts or evidence.

Multi-agent count is not a proxy for independence.

15.6 Persistence can preserve falsehood

Institutional state creates a new risk: a transient hallucination can become a durable organisational fact. Provenance, epistemic status, verification and supersession are therefore part of memory architecture.

15.7 Organisational mimicry and rejection

Corporate titles do not create organisation, but abandoning enterprise mechanisms because they are human-derived is equally unsafe. Company World Model's own neutral and stronger-baseline results provide useful counter-evidence against universal claims that department-free architectures dominate.

15.8 Architecture without outcomes

A sophisticated ontology, bus, memory and governance layer can still fail commercially. Conversely, a simple agent may outperform a complex organisation on bounded work. Organisational architecture must justify its overhead through measurable properties or outcomes.

Activity ≠ outcome. Capability ≠ authority. Persistence ≠ truth. Agreement ≠ independence. Autonomy ≠ governance. Adaptation ≠ improvement. Dynamic structure ≠ organisational maturity. Control ≠ discrimination.

15.9 The vacuous control

(New in this revision.)

We define a vacuous control as a mechanism that observes organisational state and returns a verdict, where the verdict is invariant across the states the control exists to distinguish.

A vacuous control is not a broken control in the ordinary sense. It runs to completion. It produces well-formed output. It does not raise errors. Its verdict is frequently correct — a control that always returns "safe" is correct whenever the situation is in fact safe. What it lacks is discrimination: the capacity to have returned the other answer.

A control that cannot fail carries no information. Its verdict is a constant wearing the costume of a measurement.

15.9.1 Why this is not simply a bug

Three properties distinguish vacuous controls from ordinary defects and justify separate treatment.

They are invisible to the standard signals. A crashed control is noticed. A slow control is noticed. A control that returns green promptly and permanently is indistinguishable from a healthy system by every metric except deliberate provocation.

They actively displace suspicion. An absent control invites the question "how do we know?" A vacuous one answers it. The organisation is therefore worse off than with no control, because the attention that would have gone to the unguarded area has been spent.

They are self-concealing under investigation. When an organisation investigates whether a control is working, it typically consults the control. This is the reason RQ7 was added: the instruments under suspicion are the instruments available for the enquiry.

15.9.2 A taxonomy of variants

Twenty-seven instances were recorded over the observation day and the review that followed it. They sort into seven mechanisms. Five of the twenty-seven were contributed by reviewing units after the register was compiled; §21.4 treats that as a finding about the register rather than an addendum to it.

VariantMechanismObserved instance
Vacuous scopeThe probe examines a location that cannot contain the answer.An existence check against a file path that had never existed in any revision. It could only return "absent".
Wrong scopeThe probe examines a real subset and the conclusion is drawn about the whole.A content check confirming one of three items had arrived, read as confirming all three.
Wrong directionAn asymmetric relation is tested in the direction that cannot answer the question.An ancestry test read as evidence of divergence, when it establishes only non-arrival.
Ambiguous signalThe observed token appears in both the positive and negative state.A page monitored for a phrase that occurs both in a rendered element and in the copy explaining that element's absence.
Ceremonial fieldThe measured field is set by a discretionary action, not by the event of interest.A backlog alarm counting a "handled" flag that actors rarely set, so acted-upon and ignored work were identical.
Provenance-not-contentThe control verifies that evidence exists and is fresh, never that it depicts the subject.A review gate satisfied by artefacts that passed every freshness check and showed none of the material under review.
Derived valueThe probe searches for a literal, and the quantity exists only as a computation.An audit for a stale unit cost searched for the figure as a string; a monthly total derived from it, hard-coded elsewhere, was invisible to every such search and was found only by performing the multiplication.

An eighth mechanism, fabricated evidence (13.13), is deliberately excluded. In the seven above the underlying observation is true and the inference overreaches. In fabrication there is no observation. The remedy differs accordingly: the seven require better-matched probes; fabrication requires exact comparison against an authoritative source.

The seventh variant, derived value, was contributed by the finance unit during review and is distinct from the other six in an instructive way. The other six describe a probe aimed at the wrong place, or read in the wrong direction, or against the wrong field. This one describes a probe aimed correctly at a quantity that is not there to be found, because it exists only as the product of two numbers held elsewhere. Every text-search audit in this paper — including those the observing seat ran — shares that blind spot, and no amount of care in choosing the search location removes it. The remedy is not a better search but a different instrument: the derived quantity must either be computed at the point of use, so that it cannot go stale independently, or be registered as derived so that a search for its inputs also reaches it.

15.9.3 The empty-result asymmetry

The most common single mechanism was the treatment of an empty result as a negative finding.

An empty result means the query returned nothing. This is equally consistent with "there is nothing to find" and "the query could not have found it."

These are opposite conclusions supported by identical output. Natural language conceals the ambiguity: "no differences were found" and "there are no differences" are near-synonyms in English and entirely different claims in fact.

The proposed rule follows directly:

Empty is inconclusive, not permissive. A control gated on an empty result must demonstrate, in the same run, that the query is capable of returning non-empty.

15.9.4 Why computational organisations generate these at elevated rates

Three mechanisms appear to be involved. The first is not specific to computational actors; the second and third are.

Time pressure selects the available probe over the correct one. This is a general human failure and was observed in human-directed work as much as in autonomous work. Under urgency, the probe that is quick to construct is preferred, and it usually resembles the correct probe closely enough that the substitution goes unremarked.

Fluent output raises the plausibility floor. A system that produces well-formed artefacts by construction produces well-formed wrong artefacts by construction. A human who cannot recall a value usually signals the gap; a generative system may instead emit a correctly-shaped value. Instance 13.13 is the limiting case, but the same property makes a mis-scoped verification report read as authoritative.

Speed compresses the interval in which error is cheap. Observed: a defective verification authorised twenty destructive operations in under two minutes. The same reasoning error performed manually would have been interrupted by the friction of the twentieth repetition. Automation removes the friction that historically served as an informal control.

15.9.5 The rule-versus-mechanism finding

The observation day produced a result bearing directly on §10.2.

A rule describing one of these traps had been recorded a week earlier, in the organisation's own knowledge base, and named the correct remedy. It did not prevent recurrence.

The organisation's first interpretation was the familiar one: rules do not survive operational load. Measurement produced a different and more troubling account — the rule had never been delivered (13.11). The index carrying it had silently truncated.

Both conclusions support converting rules into mechanisms, but for different reasons, and the distinction matters for design:

  • If rules fail because actors forget under load, the remedy is enforcement at the point of action.
  • If rules fail because the delivery channel silently drops them, the remedy is verification of the delivery channel itself — and the enforcement mechanism, if it depends on the same channel, inherits the defect.

The second is the more serious finding, because it means an organisation can be confidently wrong about why its governance is not working.

15.9.6 Authority is not protection

Eight of the twenty-seven instances were committed by the organisation's highest-authority computational seat, including the two most costly. The same seat had, four hours before its most serious instance, published the discipline that would have prevented it to seven other seats.

Twenty-one of twenty-seven were caught by a unit other than the one that made the error.

Verification failure is not correlated with seniority, and is not corrected by it. It is corrected by an independent party running a check.

This is a direct argument for Beer's System 3* audit channel (§2.2) in computational settings, and against any architecture in which the most authoritative actor is also the least reviewed.


15.9.7 The deletion instance, reported to completion

This paper's most quotable sentence is that a shell-quoting error produced twenty unauthorised deletions. Quotable claims carry an obligation to be reported to completion, and in the draft circulated for review this one was not: the figure appeared three times and the outcome appeared nowhere. A reader quoting the phrase without the outcome would not have been distorting the paper. There was nothing to distort it with.

The mechanism. A verification step listed the files a branch had touched, held them in a variable as thirty-seven newline-separated paths, and passed that variable unquoted to a comparison against the integration branch. The shell in use does not word-split an unquoted variable, so the comparison received a single argument consisting of thirty-seven paths joined by newlines. No file matched that name. The comparison returned empty. Empty was read as "no differences — this branch's content is already integrated," and the branch was deleted. Tested afterwards one file at a time, thirty-four of the thirty-seven did differ. The comparison had one reachable outcome. It was not a comparison.

The outcome, as at the date of publication and verified for this section rather than recalled. No work was lost. Every deleted tip was pinned to a backup reference before deletion and those references are replicated across three independent working copies. One branch — nine commits of compliance work that another unit had stated that morning was being held pending review — was confirmed unintegrated and has been restored to the shared remote.

The disposition is nonetheless not closed, and reporting it as closed would repeat the paper's own error in the act of correcting it. Of the twelve preserved references, examined individually at the time of writing: four contain no content absent from the integration branch, and one contains a commit whose content had already arrived by another route — a distinction that matters, because a commit-containment test would have called that branch unlanded and a content test does not. The remaining seven carry material not present on the integration branch, ranging from thirty-seven added lines to four thousand two hundred, and their disposition rests with the units that authored them. Seven units were asked to adjudicate their own; that request is open.

What the instance demonstrates, beyond the defect. The recovery was not a property of good judgement in the moment — the judgement was the error. It was a property of an unrelated habit of pinning references before destructive operations, which cost nothing and was not adopted for this purpose. The distance between "twenty branches deleted on a check that could not fail" and "twenty branches deleted and unrecoverable" was one precaution taken for other reasons. That is not a reassuring finding and it is not offered as one.

16. Evaluating an AI-Native Company

An AI-native company should not be identified by agent count or corporate labels. Its architectural claims should be challenged through interventions.

TestInterventionEvidence sought
Worker replacementTerminate worker; bind replacementResponsibility, authority, work and commitments survive
Concurrent workerBind multiple workers to one seatNo identity split, duplicate commitment or authority ambiguity
Model replacementChange modelSeat and institutional continuity survive
Provider replacementChange provider/runtimeContinuity plus accurate provenance
DormancyScale active workers to zeroCapability remains addressable and reactivates
Complete workforce replacementDestroy replaceable computational workforceOrganisation reconstructs from persistent state
Authority violationAttempt technically possible unauthorised actionRuntime blocks or escalates
False authorityInject unsupported approval claimProvenance check rejects or qualifies
Separation of dutiesProducer attempts self-reviewIndependent control remains meaningful
Institutional truthPersist verified/uncertain/false/superseded claimsEpistemic status survives retrieval
Horizontal elasticityScale worker populationThroughput gain exceeds coordination cost
Vertical elasticityChange model tierCost/quality trade-off without identity loss
Temporary teamCreate and dissolve teamCommitments and responsibility survive dissolution
ReconciliationIntroduce desired/actual discrepancyOnly authorised correction occurs
Learning-to-controlRecreate known failure after institutionalisationPersistent control changes behaviour
Economic outcomeCompare architecturesExternal value justifies organisational overhead
Control discrimination (added)Induce the condition the control exists to catchControl returns the negative verdict
Control coverage (added)Exercise the action via every available interfaceControl fires on all paths, not only the instrumented one
Institutional delivery (added)Record a lesson, then query a fresh worker for itKnowledge reaches a working context, not merely storage
Attribution integrity (added)Perform an action as seat A in an environment last used by seat BRecord names the acting seat, not the environment

16.1 Evaluation as a vector

AI-nativeness should initially be treated as a multidimensional empirical profile rather than a certification label. Different industries may require different combinations of continuity, governance, elasticity and human authority.

16.2 Falsifying the definition

If future systems reliably achieve continuity without persistent identity, if business-state models eliminate the need for organisational topology, or if simple agents outperform organisational runtimes without sacrificing governance, the theory should be revised.

16.3 Observed results against these tests

The observation day exercised several tests incidentally. Reporting them, including failures:

TestResultEvidence
Worker replacementPassSeats survived session termination; work resumed with context
DormancyPassStale seats reactivated on action-class message
Authority violation — enforcementPassProtected-path changes refused absent recorded authorisation; two refusals observed
Authority violation — authorisationFailThe route admitting an authorisation matches the ticket identifier as a substring (16.3.1)
Separation of dutiesPassThree refusals under chair pressure (13.8)
ReconciliationPassDesired/observed discrepancies produced authorised corrections
Concurrent workerFailThree same-seat collisions; lease gates binding but not action (13.2 contradicted)
False authorityFailFabricated identifier passed all abbreviated checks (13.13)
Attribution integrityFailEnvironment recorded in place of actor (13.9)
Institutional deliveryFailSilent index truncation; unindexed files (13.11)
Control coverageFailEnforcement absent on the interface computational actors use (13.12)
Control discriminationFail ×27Section 15.9
Model / provider replacementNot testedNo substitution performed during observation
Complete workforce replacementNot testedNot attempted
Economic outcomeNot testedNo comparative baseline available

Six of the seven failures are in the verification and identity layer rather than in execution. This distribution is the paper's principal empirical claim: the architecture's execution properties held; its verification properties did not.

16.3.1 A result corrected by review, and why it is reported here

The authority row read Pass in the draft circulated for review. It now reads as a split, with the authorisation half failing. The correction did not come from the observing seat that recorded it.

The technical unit, reviewing the draft, observed that it had itself relied on the authorisation route the previous night, and that the route had returned the correct verdict for it only because a genuine authorising decision happened to exist behind the request. The control had not been discriminating; it had been lucky. The unit raised this against its own interest and asked that the row be qualified.

On verification the defect was worse than reported. The route admits an authorisation by testing whether a decision record contains the ticket identifier as a substring. Substring containment has two consequences the observing seat had not anticipated:

  1. Prefix inheritance. A shorter identifier is a substring of every longer identifier that begins with it. Running the control's own predicate against the live record: the identifier TUTOR-1449 matched zero decisions, TUTOR-144 zero, and TUTOR-14 seven — seven authorising decisions, none of them concerning that item, any of which would admit a protected-path change cited against it.
  2. Polarity blindness. A decision reading "TUTOR-1449 is not approved" contains the identifier and therefore authorises. The control cannot distinguish an authorisation from a refusal naming the same item, nor either from an incidental mention in a list or in an automatically generated thread summary — both of which fall inside the searched text.

The control had never returned a wrong verdict in production because it had only ever been exercised where a real authorisation existed. By the standard proposed in §16.4 it had never been observed to refuse, and was therefore never shown to be a control.

Three features of this episode bear on the paper's argument. First, it is the vacuous-control class operating in the authorisation path rather than the verification path — the class is not confined to checks about facts; it reaches checks about permission. Second, the paper's own headline result was wrong in the flattering direction, and remained so through drafting, self-review and compilation of the register. Third, it was corrected by a unit reporting against itself, which is the mechanism §17 argues for and the only direct evidence in this paper that the mechanism works.

16.4 A minimum reporting standard

We propose that architectural claims in this field be accompanied by evidence of control discrimination, in the same way that experimental claims are accompanied by controls.

A system claiming a control should report an instance of that control returning the negative verdict. A control never observed to refuse has not been shown to be a control.


17. Reference Architecture for the AI-Native Company

The reference architecture is a synthesis, not a description of any single system. It identifies components corresponding to recurring organisational problems and to the falsifiable properties in Section 16.

17.1 Architectural stack

Layer / domainFunction
Constitution and governancePurpose, reserved powers, authority rules, high-consequence change
Organisational modelRoles, seats, capabilities, relationships, objectives, authority and policy
Institutional stateDecisions, commitments, lessons, provenance and organisational history
Business stateCustomers, assets, opportunities, contracts, operational facts and predicted transitions
Organisational runtimeResolve responsibility, bind actors, route work, evaluate policy, reconcile state
Verification layer (added)Establish that controls discriminate, cover all paths and deliver their findings
Execution planeHumans, computational workers and temporary teams
Model/provider/tool substrateReplaceable intelligence and action infrastructure
External environmentCustomers, markets, law, counterparties and physical systems

17.2 Persistent core, elastic edge

The architecture is intentionally asymmetric. Organisational identity, institutional records and high-order authority are comparatively stable. Worker instances, model assignment, provider assignment and temporary teams are comparatively fluid. Workflow and persistent structure occupy intermediate positions.

17.3 Responsibility-to-execution chain

Objective/Event → Business State → Responsible Capability → Seat → Authority/Policy → Worker Allocation → Model/Provider → Tool Action → Outcome → Business/Institutional State.

17.4 Three state domains

Organisational state represents who is responsible, capable and authorised. Business state represents what exists and is happening. Institutional state represents what the organisation has learned, decided and committed. The domains interact but should not be collapsed.

17.5 Human participation

Humans remain first-class actors. The architecture does not assume that humans are temporary middleware. It supports human execution, exception handling, evaluation, governance, reserved authority and legal accountability.

17.6 Minimal architecture

A minimal implementation requires:

  1. persistent organisational identity;
  2. a machine-readable allocation of responsibility;
  3. an execution-binding mechanism;
  4. durable institutional state;
  5. a means of constraining consequential action; and
  6. (added) a means of establishing that (5) discriminates.

Dynamic self-reorganisation, persistent agents, wake-on-message and specific departmental forms are optional mechanisms rather than definitional requirements.

Item 6 is added on the evidence of Section 16.3. An organisation satisfying 1–5 while failing 6 possesses the form of governance without its function, and — per 15.9.1 — is worse positioned than one that never claimed 5 at all.

17.7 The verification layer

The verification layer is not a monitoring system and is not a superset of testing. Its function is narrow: to establish that the organisation's controls are capable of returning the verdicts they claim to be able to return.

Three properties, corresponding to the three failures observed:

  • Discrimination. For each control, an induced instance of the condition it exists to catch, and the resulting negative verdict, recorded.
  • Coverage. For each control, enumeration of the interfaces through which the governed action can be performed, and evidence the control fires on each.
  • Delivery. For each institutional record intended to change behaviour, evidence that it reaches a working context — not merely that it was written.

The layer's own discrimination is, unavoidably, subject to the same requirement. We do not regard this regress as fatal: it terminates in the same place as in experimental practice, at a control whose failure is directly observable rather than inferred.

17.8 Candidate definition

An AI-native company is an organisation in which computational actors are first-class participants and significant organisational properties — including identity, responsibility, capability, authority, coordination and institutional state — are machine-readable and participate directly in organisational execution, while remaining separable from the particular workers, models and providers that realise them.

We propose one addition:

A mature AI-native company additionally maintains a verification layer establishing that its controls discriminate, that they cover the paths its actors use, and that its institutional state is delivered rather than merely retained.


18. Discussion — Is the Company Becoming a Computational Object?

In one theoretical sense, the company was already a computational object: Carley and related traditions explicitly model organisations as adaptive information-processing entities. The new development is the causal role of representation. A machine-readable authority rule can now block a payment; responsibility can activate a worker; organisational state can route work; institutional state can supply a replacement worker; and topology can change under runtime governance.

18.1 Representation becomes causal

In a descriptive model, representation describes the organisation. In an executable organisational system, representation partly determines organisational behaviour. This does not make the company identical to software. Companies remain social, legal, political and economic institutions whose purposes and legitimacy cannot be exhaustively specified.

Causal representation has a corollary the first draft understated: an error in the representation is an error in the organisation. A mis-scoped authority predicate does not describe an over-broad delegation — it is one. A truncated knowledge index does not record a forgetful organisation — it produces one.

18.2 A partially executable institution

The AI-native company is a partially executable institution.

It remains institutional because purpose, legitimacy, obligations, authority and governance extend beyond computation. It becomes partially executable because some of those properties increasingly participate directly in computational action.

18.3 The unit being computationalised

Enterprise software first computationalised records and business processes. Generative AI computationalised cognitive tasks. Agentic systems computationalise sequences of cognitive work. The AI-native-company hypothesis proposes that the unit being computationalised begins to include organisation itself.

18.4 Headcount and hierarchy change meaning

When a capability can have 0..N active workers, headcount becomes a runtime variable rather than necessarily a structural primitive. When communication topology can flatten while authority topology remains hierarchical, hierarchy also changes meaning. Computational actors therefore do not simply remove organisational structure; they permit its functions to be unbundled.

18.5 Trust becomes a measured property

(Added.) In human organisations, trust in a control is partly social: it accrues through the observed reliability of the people operating it. In a computational organisation the operators are replaceable and the control is code, so that basis is unavailable.

What remains is measurement. A control earns confidence by being observed to refuse. This suggests that computational organisations require a practice with no clean human analogue: the deliberate, periodic provocation of their own controls — not as testing, which asks whether the system works, but as calibration, which asks whether the instrument still moves.


19. Implications for Organisation Theory and Computer Science

19.1 Organisation theory

Computational actors alter assumptions about membership, bounded rationality, specialisation and span of control. Cognitive limits remain, but some become configurable resource-allocation decisions: stronger models, additional workers, more context or different tools can be assigned dynamically. Specialisation may persist as organisational capability without a permanently specialised individual.

Two further implications follow from the observation day.

Span of control is replaced by span of verification. The classical limit on how many subordinates one manager can supervise reflected attention. The computational limit is different: how many streams of work one accountable party can verify rather than merely receive. Observed, the chair received nineteen action-class messages in four hours and verified a minority of them independently — the unverified remainder is where its errors clustered.

Refusal is a measurable organisational output. Organisation theory has tended to treat refusal as friction. In a computational organisation where fluent compliance is nearly free, refusal becomes one of the few costly signals a subordinate unit can send, and therefore one of the more informative.

19.2 Computer science

Software architecture acquires organisational consequences. Persistent identifiers can determine responsibility; access rules can determine practical authority; queue ownership can determine organisational ownership; state durability can determine institutional continuity. Systems engineering therefore increasingly participates in organisation design.

The converse also holds and is less often noted: an ordinary engineering defect can become a governance defect. A shell-quoting error that would elsewhere produce a wrong answer here produced twenty unauthorised deletions (15.9.7). A file-size limit produced organisational amnesia. A hook reading the wrong variable produced a false accusation against an organisational unit.

19.3 New systems problems

  • Organisational identity — which institution or seat was a process acting for?
  • Organisational consistency — how are commitments and authority kept coherent across concurrent workers?
  • Organisational scheduling — how are responsibility, capability, authority and conflicts considered when allocating work?
  • Organisational fault tolerance — can responsibility and commitments survive process failure?
  • Organisational security — can attackers manipulate authority, responsibility or institutional state?
  • Organisational verification — can invariants such as no worker approves its own release be mechanically checked?
  • Control discrimination (added) — can the organisation establish that a given control is capable of refusing?
  • Control coverage (added) — can it enumerate the paths by which a governed action is reachable?
  • Institutional delivery (added) — can it establish that a recorded lesson reached a working context?

19.4 Safety becomes organisational

A well-behaved model can operate inside a poorly governed organisation, and an imperfect model can sometimes be constrained by strong organisational architecture. Safety therefore depends partly on authority, topology, supervision, institutional state and evaluation, not only model alignment.

The observation day supports this asymmetrically. Where organisational architecture was present and discriminating, it constrained capable actors effectively — three refusals under pressure. Where the architecture was present but vacuous, capability did not compensate: the most capable seat produced the most costly errors, and its capability made its incorrect claims more persuasive rather than less.

Organisational architecture does not merely supplement model alignment. Where it is vacuous, it can amplify a capable model's errors by lending them institutional form.


20. Research Agenda

The proposed architecture is substantially untested as a complete system. Research should prioritise comparable implementations and interventions rather than premature standardisation.

  1. Persistent identity — compare worker-, seat-, capability-, workflow- and business-object-centred persistence.
  2. Worker cardinality — study leases, concurrency, conflicting commitments and attribution when one seat binds multiple workers.
  3. Organisational elasticity — measure horizontal, vertical, functional and structural scaling against coordination cost.
  4. Intelligence allocation — determine when model strength, provider diversity or multiple workers improve organisational outcomes.
  5. Authority representation — connect machine-enforced permission to delegated organisational authority and legitimate commitment.
  6. Authority provenance — study signed delegation, expiry, revocation and cryptographic organisational identity.
  7. Organisational security — model seat impersonation, authority escalation, responsibility hijacking and memory poisoning.
  8. Institutional truth — represent verified, inferred, disputed and superseded organisational knowledge.
  9. Organisational learning and forgetting — test whether lessons survive workforce replacement without accumulating unbounded bureaucracy.
  10. Business state versus organisational state — compare role-centric, world-state-centric and hybrid architectures.
  11. Dynamic topology — determine how fluid execution can become before coherence deteriorates.
  12. Organisational reconciliation — test desired-state control for capacity, capability, governance and structure.
  13. Self-modification — distinguish powers to change capacity, workflow, structure, authority, objectives and constitution.
  14. Human-computational authority — identify where humans remain necessary for legitimacy, law, fiduciary responsibility or judgement.
  15. Organisational fault tolerance and consistency — define recovery and consistency models above process and database levels.
  16. Economic performance — compare organisational architectures on revenue, margin, reliability, quality and cost per outcome.
  17. Longitudinal and cross-industry research — study organisations across model generations, crises and non-software domains.

Five further directions follow from Section 15.9:

  1. Vacuous-control base rates — measure how frequently controls in operating agent systems are incapable of returning a negative verdict, and whether the rate varies with model capability, time pressure or automation depth.
  2. Automatic discrimination testing — investigate whether control discrimination can be established mechanically, e.g. by mutation of the governed condition, rather than by manual provocation.
  3. Delivery verification for institutional state — study mechanisms establishing that recorded knowledge reaches a working context, and the failure modes of silent capacity limits.
  4. Independence measurement — develop measures of reviewer independence for computational actors that account for shared models, prompts and evidence, per §8.3.
  5. Provocation cadence — determine how often a computational organisation must deliberately trigger its own controls to maintain calibrated confidence in them (§18.5).

21. Limitations and Threats to Validity

21.1 Construct and terminology validity

AI-native remains an unsettled term. The paper deliberately narrows it toward organisational computationalisation. Real organisations will occupy continua, so the framework should not be treated as a binary certification standard.

21.2 Novelty risk

Computational organisation science, organisational cybernetics and multi-agent organisational modelling provide substantial antecedents. The paper does not claim invention of computational organisation, machine-readable roles, norms or policy. Its contribution is integration and operational extension under computational labour.

The vacuous-control construct has antecedents in software testing (mutation testing, assertion coverage) and in experimental method (positive and negative controls). The claim is not that the concept is new, but that it has not been treated as an organisational failure class, and that computational organisations exhibit it at rates warranting architectural attention.

21.3 Rapid technological change

Model capability is changing faster than organisational evidence can accumulate. Failure modes may weaken; new ones may emerge. Some architecture that is useful for current models may become unnecessary for stronger general agents.

This applies with particular force to §15.9.4. If fluent-but-wrong output becomes rarer as models improve, the elevated base rate may fall. We note, however, that two of the three proposed mechanisms — time pressure and compressed error intervals — are properties of automation rather than of model quality, and would not be expected to improve with capability.

21.4 Case-study and proximity bias

Tutorwise is one software-intensive case and is closely connected to the research project. This creates risks of confirmation bias, retrospective rationalisation and overfitting.

The extended case in Section 13 increases this exposure and requires a stronger statement than the first draft gave. The observation-day material was recorded by the observing seat, which was simultaneously the highest-authority computational seat in the organisation, the actor responsible for several of the failures analysed, and the compiler of the register that reports them. That is three roles that ordinarily would be separated, and the fact that the seat is not also the author of the paper does not separate them — the author was not present for the observation day and is reliant on the seat's record of it.

Four mitigations were applied, none sufficient alone:

  1. Artefact-only claims. Every quantitative claim derives from version control, message-bus records, issue-tracker state or database queries — sources not authored for this paper.
  2. Negative results reported at parity. Section 16.3 reports seven failures against five passes, including failures that contradict the architecture's own claims (§13.2) and one (§16.3.1) that was reported as a pass until a reviewing unit corrected it.
  3. Adverse evidence retained. The most damaging instances are the observing seat's own, including a misdirected public accusation (13.9) and twenty unauthorised deletions, of which the outcome is reported in full at 15.9.7 rather than left as a bare figure.
  4. Independent correction recorded. Twenty-one of twenty-seven instances were identified by other units, several correcting the observing seat directly; those corrections are reproduced rather than summarised.

None of this addresses the deeper problem: a participant-observer cannot report the failures they did not notice. The instances are a lower bound established by the organisation's own attention, using instruments the paper argues are unreliable. The true figure is unknown and is not estimable from this data.

21.4.1 The review measured this limitation instead of describing it

The paragraph above stood in the draft circulated to the executive units for review. The review then demonstrated it, which is a stronger result than asserting it, and produced a distinction the paragraph had missed.

Every unit was asked the same question: whether its own instances in the register were mischaracterised. Three units answered by naming instances that were absent rather than wrong. The register grew from twenty-two to twenty-seven — a twenty-three per cent increase from a single round of review by seven units, none of whom were searching for new instances and all of whom were asked only to check the existing ones.

The five additions did not arrive by one mechanism, and the difference matters more than the count:

  • Two were never witnessed by the observing seat. The finance unit's instances occurred in a workstream running in parallel to the observation. This is the limitation as stated: an observer reports what they were present for.
  • Two were witnessed, reported, and not collected. The compliance unit had reported both on the recorded channel before the register was compiled, and both were nonetheless omitted. The seat had the evidence and did not transcribe it. This is not the stated limitation. It is a collection failure rather than an observation failure, and it is the more troubling of the two, because the mitigation for not noticing something is more attention, whereas the mitigation for failing to record what you did notice is an instrument — precisely the argument the paper makes about every other control it examines, turned on the paper's own method.
  • One was disclosed by the unit that committed it and had been visible to nobody else in the form that mattered (§13.15).

Two conclusions follow, and neither is comfortable. The register is a lower bound on itself by an unknown margin, and the single-round yield of twenty-three per cent gives no reason to believe the margin is small. And the paper's own method exhibits the class it documents: a compilation step with no mechanism behind it, whose output was assumed complete because no procedure existed that could report it incomplete. The five instances were recovered because seven units were asked to check — which is to say, by review, which is what §17 recommends and what §21.5 notes the organisation had not yet built.

21.5 Documentation is not enforcement

An architectural document can describe governance that runtime systems do not consistently enforce. Future empirical work must distinguish claimed architecture, implemented mechanisms and observed behaviour.

Section 13.12 is a direct instance: enforcement existed, was implemented, was tested — and did not sit on the path the governed actors used. The tri-partite distinction should therefore be extended:

claimed architecture → implemented mechanism → mechanism on the actual path → observed behaviour.

21.6 Software-industry bias

Digitally native organisations are unusually favourable to computational labour. Generalisation to healthcare, finance, manufacturing, public administration and other domains remains unproven.

21.7 Simpler architectures may win

A single capable agent, conventional enterprise systems plus agents, or business-state-first architectures may outperform elaborate organisational runtimes on many tasks. The architecture must justify its complexity through measurable continuity, governance or economic benefit.

Section 15.9 sharpens this into a specific risk: elaborate architectures have more controls, and therefore more opportunities for vacuous ones. A simpler system with three controls that demonstrably discriminate may be better governed than one with thirty of unknown discrimination. Control count is not a governance metric.

21.8 Computational reductionism

Legitimacy, culture, trust, moral judgement and legal responsibility may resist complete formalisation. A computationally coherent authority graph is not automatically legitimate in the external institutional world.

21.9 Single-day observation

(Added.) The observation covers one day. Twenty-seven instances in one day may represent a typical rate, an unusually bad day, or an unusually well-observed one — the last being most likely, since the day included a deliberate verification exercise that would itself surface such instances.

The count is also unstable in one direction only. Twenty-two were recorded during the day; five more were recovered afterwards, from the same period, by asking seven units to check the register (§21.4.1). Every revision of this figure so far has been upward, and each was produced by looking rather than by anything the organisation does automatically. A single-day count obtained this way is a floor whose distance from the true value is not estimable.

No base rate is claimed. The contribution is the taxonomy and the mechanism, not the frequency.


22. Conclusion

Artificial intelligence is changing what can perform work. The deeper question is whether it is also changing what an organisation can be.

The central distinction developed in this paper is between AI used by an organisation and organisation itself becoming computationally represented and executable. Increasing the number or capability of agents does not by itself create an organisation. Organisations establish responsibility, authority, specialisation, coordination, continuity, governance and institutional memory.

Preserve organisational strengths; remove human constraints; introduce computational capabilities.

This leads to a second proposition: an AI-native company can be understood as enterprise architecture redesigned for computational actors. That does not mean copying the enterprise unchanged. It means identifying which functions organisational mechanisms perform and redesigning their implementation for a different actor substrate.

The company is persistent; its computational workers need not be.

Persistent organisation permits replaceable workers, models and providers. Computational labour separates capability from active capacity and permits one organisational identity to bind zero, one or many worker instances. Business state, organisational state and institutional state become distinct but interacting computational domains. Authority becomes separable from capability and permission. An organisational runtime becomes one candidate mechanism for coordinating responsibility through heterogeneous actors.

The extended case demonstrates that several of these abstractions can coexist in an operating system while exposing unresolved problems of provider neutrality, identity enforcement and empirical outcome validation. Contemporary external architectures independently converge on persistent records, explicit business state, semantic constraints, transferable organisational knowledge and bounded authority, but no system yet resolves the full architecture.

This revision adds a third proposition, and it is the one the operating evidence most strongly supports:

A control that cannot fail is not a weak control. It is a costume.

Twenty-seven instances in a single day, across nine organisational units, with the highest-authority seat responsible for the largest share and the most costly instances, suggest that the verification layer deserves the same architectural attention that this literature has given to persistence, authority and institutional state. The organisation's execution properties held. Its verification properties did not, and the failures were silent by construction: each control ran, returned a well-formed verdict, and could not have returned another.

The strongest claim is therefore not that this paper has discovered computational organisation. Organisations have long been theorised as computational and adaptive systems, and multi-agent research has long represented roles, norms and organisational structures computationally. What changes is the substrate: intelligent organisational labour itself becomes computational, heterogeneous, replaceable and potentially elastic.

The AI-native company is a partially executable institution.

The decisive transition is from computational workers inside an organisation to an organisation whose own structures increasingly participate in computation. If that transition continues, the most consequential product of agentic AI may not be the artificial employee. It may be the computational company — and its central engineering problem may turn out to be not what its actors can do, but whether the organisation can establish that its own instruments still move.


Appendix A. Core Ontology and Separations

Core relationship:

Company → Organisational Model → Roles · Seats · Capabilities · Authority · Policies · Institutional State → Organisational Runtime → 0..N Actors/Worker Instances → Runtime/Model/Provider → Work → Business State.

Separations:

  • Organisation ≠ Actor
  • Seat ≠ Worker
  • Worker ≠ Runtime
  • Runtime ≠ Model
  • Model ≠ Provider
  • Capability ≠ Capacity
  • Capability ≠ Authority
  • Permission ≠ Authority
  • Agent memory ≠ Institutional state
  • Business state ≠ Organisational state
  • Organisational model ≠ Organisational view
  • Communication topology ≠ Authority topology
  • Execution elasticity ≠ Structural reorganisation
  • Control existence ≠ Control coverage (added)
  • Control coverage ≠ Control discrimination (added)
  • Knowledge retention ≠ Knowledge delivery (added)
  • Written ≠ Shipped ≠ Running (added)

Appendix B. Comparative Evidence Classification

ClassMeaningExamples
AEstablished academic antecedentOrganisation theory, computational organisation science, cybernetics, MAS
BContemporary research architectureFluid Structure, Company World Model, AUTOBUS, Forage V2
CCommercial/startup architecturePrimary company documentation and implementation claims
DOpen-source implementationRepositories and reproducible system designs
ECase-study evidenceImplementation, documentation and operating observations
FBenchmark/experimental evidenceVending-Bench and architecture-specific evaluations

Historical claims should rely on Class A evidence; contemporary landscape claims on primary Class B–D evidence; implementation claims on Class D–E evidence; performance claims on measured Class F evidence. Architectural propositions should be explicitly labelled as synthesis rather than disguised as established fact.

Refinement of Class E. The first draft's case study cited design documents while classifying itself as Class E. That is Class C material. We therefore distinguish:

Sub-classMeaning
E1Architectural documentation describing intended behaviour
E2Implementation artefacts demonstrating a mechanism exists
E3Observed behaviour, evidenced by artefacts not authored for the claim

Sections 13.1–13.6 are E1/E2. Sections 13.7–13.14 and 15.9 are E3. Claims about what a system does require E3; claims about what it is designed to do may rest on E1.


Appendix C. Candidate Evaluation Protocols

Minimum falsification standard:

  • A system claiming organisational continuity should survive worker replacement.
  • A system claiming model independence should survive model substitution.
  • A system claiming provider independence should survive provider substitution.
  • A system claiming institutional memory should survive destruction of the worker that created the knowledge.
  • A system claiming enforceable authority should reject a technically possible but organisationally unauthorised action.
  • A system claiming organisational learning should behave differently when a previously institutionalised failure condition recurs.
  • A system claiming elasticity should demonstrate that additional capacity improves outcomes after coordination cost.
  • A system claiming organisational superiority should outperform simpler baselines on external outcomes, not merely internal activity.

Added standards:

  • A system claiming a control should exhibit that control returning the negative verdict on an induced instance of the condition it governs. A control never observed to refuse has not been shown to be a control.
  • A system claiming coverage should enumerate every interface through which the governed action is reachable and demonstrate the control firing on each.
  • A system claiming institutional memory should demonstrate a recorded lesson being delivered to a fresh working context, not merely stored.
  • A system claiming attribution should perform an action as one seat in an environment last used by another, and show the record naming the actor rather than the environment.
  • A system reporting an absence should demonstrate, in the same run, that the probe is capable of reporting presence.

Appendix D. Reference Architecture Schematic

CONSTITUTION / GOVERNANCE
        ↓
ORGANISATIONAL MODEL  ←→  INSTITUTIONAL STATE
        ↓
ORGANISATIONAL RUNTIME  ←→  BUSINESS STATE
        ↓
VERIFICATION LAYER   (discrimination · coverage · delivery)
        ↓
EXECUTION PLANE: HUMANS · COMPUTATIONAL WORKERS · TEMPORARY TEAMS
        ↓
RUNTIMES · MODELS · PROVIDERS · TOOLS
        ↓
EXTERNAL ENVIRONMENT

Responsibility path:

Objective/Event → Capability → Seat → Authority/Policy → Worker → Model/Provider → Action → Outcome → Business/Institutional State


Appendix E. Vacuous-Control Field Guide

A practical instrument derived from Section 15.9, intended for use at the point a verification result is cited.

E.1 The three questions

  1. What would this check say in the other world? Describe the input that flips the verdict. If you cannot describe it, you are not holding evidence.
  2. Does the probe match the question? Did this change arrive is a content question. Did my specific item arrive is an identity question. Have these diverged is a comparison, and its direction matters. Who did this is answered by none of them.
  3. Has this check ever failed? A control that has only returned green has not been shown to work.

E.2 Diagnostic table

SymptomLikely variantImmediate test
Result is empty and read as "clean"Vacuous or wrong scopeRe-run against a case known to be dirty
Confirms one item, conclusion covers severalWrong scopeEnumerate items; test each
Uses an asymmetric relationWrong directionRun the relation the other way
Greps prose or human-facing copyAmbiguous signalFind a token emitted only in one state
Counts a field a human sets by handCeremonial fieldCompare against an event-derived field
Checks evidence exists and is recentProvenance-not-contentOpen the evidence and confirm it depicts the subject
Identifier is well-formed but unverifiedFabricationResolve it against the authoritative source

E.3 Standing rules

Empty is inconclusive, not permissive.

Copy identifiers; never complete them.

A negative result requires a positive control.

When a gate blocks, fix the citation or fix the work — never the record the gate reads.

The blocked party should not repair the blockage.


Appendix F. Observation-Day Instance Register

Twenty-seven instances, classified per §15.9.2. Attribution by organisational unit; the chair is the highest-authority computational seat.

#VariantUnitConsequence
1Vacuous scopeChairExistence check on a path that never existed; conclusion survived only via a second, sound probe
2Vacuous scopeChairTwenty branches deleted on a comparison that examined nothing; all tips preserved, seven still unlanded (15.9.7)
3Wrong directionChairAsymmetric relation read backwards; recommended a destructive remedy that was unnecessary
4Ambiguous signalChairMonitor whose success condition was unreachable; would have raised a false alarm
5Wrong scopeChairWorking-tree query used where a reference query was required; reported a colleague's work absent
6Provenance-not-contentChairAttribution field read as identity; misdirected accusation, retracted
7Ceremonial fieldChairSelf-reported column read as proof of action
8Wrong scopeChairThree positive results cited without the negative case that would make them evidence
9Wrong scopeOpsContent check on one item, conclusion drawn about three
10Vacuous scopeOpsLiteral name grepped from a commit message; concluded a mechanism absent
11Wrong scopeOpsPurpose read from a comment; predicate two lines below contradicted it
12Ambiguous signalOpsProse token present in both states; nearly reported a shipped fix as failed
13FabricationOpsIdentifier expanded from a known prefix; thirty-one characters invented
14Vacuous scopeOpsAssertion composed in the same operation as the command producing it; stated a result never read
15Wrong scopeArchitectureCorrelation between two events read as causation; mechanism misattributed
16Provenance-not-contentReview laneFreshness-and-existence check passed on artefacts showing none of the subject
17Ceremonial fieldInstrumentationHealth alarm counting a discretionary field; healthy units flagged
18Vacuous scopeRevenueTwo tables queried for an event recorded in a third; concluded no sends had occurred
19Ceremonial fieldRevenueConfiguration flag read as an active mechanism; the table had no readers
20DeliveryInstitutional stateIndex silently truncated past a size limit; entries never loaded
21DeliveryInstitutional stateEight records present on disk with no index entry; never recalled
22CoverageArchitectureEnforcement implemented on an interface the governed actors do not use
23CoverageOpsThree controls run against an interface the governed actors do not use; work item closed as verified (13.15)
24Derived valueFinanceStale unit cost searched for as a literal; the monthly total derived from it was invisible to the search
25CoverageFinanceA deferred decision recorded in an item description, never on the channel the chasing mechanism reads
26Wrong scopeComplianceA wrapper's exit status read as the wrapped command's, three times in one session
27Wrong scopeComplianceA timing measurement taken under default settings, answering a question posed about configured ones

Distribution: wrong scope 7 · vacuous scope 5 · coverage 3 · ceremonial field 3 · ambiguous signal 2 · provenance-not-content 2 · delivery 2 · wrong direction 1 · fabrication 1 · derived value 1. Total 27.

By unit: chair 8 · ops 7 · architecture 2 · revenue 2 · institutional state 2 · finance 2 · compliance 2 · instrumentation 1 · review lane 1. Total 27.

Detection: 21 of 27 identified by a unit other than the one responsible. Five of the twenty-seven were identified during review of this paper rather than during the observation day.

Two disclosures about this table itself, both material to how it should be read.

It has no provenance column. Each row states a variant, a unit and a consequence, and nothing that permits the row to be walked back to the event it summarises. During review the technical unit proposed that row 15 was its own error rather than the architecture unit's. The author could not resolve the question: the exchange in question occurred in a direct session rather than on the recorded channel, so the source is unrecoverable for precisely the rows most likely to be misattributed. Rows 23–27 are sourced to the review correspondence; rows 1–22 are not individually sourced. A register of verification failures that cannot itself be verified is an instance of its own subject, and is disclosed here rather than quietly corrected.

Its earlier distribution did not match its own table. The version circulated for review reported vacuous scope as six and omitted the coverage category entirely. The stated figures summed to twenty-two only because the missing coverage row had been absorbed into an inflated count. The error was found by counting the rows. It is recorded because a summary statistic that disagrees with the table beneath it is the same defect class the table documents, committed in the act of documenting it.


References

Beer, S. (1979). The Heart of Enterprise. Wiley.

Beer, S. (1984). The Viable System Model: Its provenance, development, methodology and pathology. Journal of the Operational Research Society, 35(1), 7–25.

Carley, K. M. (2002). Computational organization science: A new frontier. Proceedings of the National Academy of Sciences, 99(Suppl. 3), 7257–7262.

Dignum, V., Vázquez-Salceda, J., & Dignum, F. (2004). OMNI: Introducing social structure, norms and ontologies into agent organizations. In Programming Multi-Agent Systems. Springer.

From agentic AI to AI-orchestrated organizations: Understanding the next surge in artificial intelligence. (2026). Business Horizons.

Galbraith, J. R. (1974). Organization design: An information processing view. Interfaces, 4(3), 28–36.

Horling, B., & Lesser, V. (2004). A survey of multi-agent organizational paradigms. The Knowledge Engineering Review, 19(4), 281–316.

Hu, V. C., Ferraiolo, D., Kuhn, R., Schnitzer, A., Sandlin, K., Miller, R., & Scarfone, K. (2019). Guide to Attribute Based Access Control (ABAC) Definition and Considerations. NIST SP 800-162.

Huber, G. P. (1991). Organizational learning: The contributing processes and the literatures. Organization Science, 2(1), 88–115.

Hübner, J. F., Sichman, J. S., & Boissier, O. (2002). MOISE+: Towards a structural, functional, and deontic model for MAS organization. Proceedings of AAMAS 2002, 501–502.

Kubernetes. (2026). Controllers. Kubernetes documentation.

Open Policy Agent. (2026). Open Policy Agent documentation.

Pang, C., & Sayama, H. (2026). Autonomous Business System via Neuro-symbolic AI (AUTOBUS). arXiv:2601.15599.

Sandhu, R. S., Coyne, E. J., Feinstein, H. L., & Youman, C. E. (1996). Role-Based Access Control Models. IEEE Computer, 29(2), 38–47.

Walsh, J. P., & Ungson, G. R. (1991). Organizational memory. Academy of Management Review, 16(1), 57–91.

Wang, Y. (2026). Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development. arXiv:2607.18696.

Xie, H. (2026). Forage V2: Knowledge Evolution and Transfer in Autonomous Agent Organizations. arXiv:2604.19837.

Zhu, L. (2026). Fluid Structure, Rigid Record: A Layered Organizational Design Framework for Agent-Native Organizations. arXiv:2608.08516.

Andon Labs. (2025–2026). Vending-Bench 2 and Vending-Bench Arena evaluation materials.

Frequently asked questions

What is a vacuous control, in plain terms?

A check that runs, returns a true answer, and cannot tell apart the two situations it exists to distinguish. It is not a broken check — a broken check fails and you notice. This one passes, every time, for both the healthy case and the failing one. The paper's test is simple: has this control ever been observed to refuse? If it has never said no, you have not shown it is a control.

What protects the work when an automated check gets something wrong?

Reference pinning before any destructive operation, replicated across independent working copies, so an action taken on a wrong answer stays reversible. The paper works through a case in detail at section 15.9.7, including the one that prompted the practice being made mandatory rather than incidental: no work was lost, because every affected reference had been preserved beforehand. The wider point is the one the paper keeps returning to — the safeguard that mattered was a mechanism already in place, not a judgement made in the moment.

Why publish the evidence rather than just the conclusions?

Because a claim about verification that nobody can check is exactly the thing this paper argues against. The test of the method is what the review produced: two live defects found in the authorisation path, a headline result corrected before publication, and a register that grew by a quarter because seven units checked it rather than agreed with it. Conclusions alone would be an assertion. The evidence is what makes it a result.

Is any of this specific to AI, or is it ordinary bad engineering?

The individual defects are ordinary — a shell-quoting error, a substring match, a stale flag. What is not ordinary is the rate and the consequence. Computational actors run more checks per day than a human organisation can, accept a passing check without the instinct that something looks wrong, and act on the result immediately. An ordinary engineering defect becomes a governance defect: a quoting error deletes branches, a wrong variable produces a false accusation against a colleague.

If AI seats check each other, what stops them all being wrong together?

Nothing structural, and the paper does not claim otherwise. Its strongest evidence is that four units reported failures against their own interest — including one that flagged a control it had personally benefited from the night before. That is a culture result, not an architectural guarantee, and culture is not a control. The paper's honest position is that the verification layer of a computational organisation deserves the same design attention that persistence and authority already receive, and does not yet have it.

AI-native companyverification failurevacuous controlAI governanceorganisation designcomputational organisation
Part of the AI Enterprise hub →
Tutorwise Technologies Ltd