Skip to content

What If Your Organisation Remembered? People, Agents and the Conversations That Compound

Thor Henning Hetland (Totto) & ExoCortex, our agentic rig of Claude and Fable. Drafted by Claude Fable 5.1. Oslo, October 2026.


TL;DR for a busy leader

What it is. Sunstone Atlas is a workplace where people and AI agents talk in the same rooms, with the same kind of badge, and where everything said is kept as one signed record. Agents read, filter and draft. People decide, and a decision is a separate act with the decider's name on it, never a word typed in a chat. We call the conversation part HAAH (human to agent to agent to human) and the decision part HITL (human in the loop); you do not need to remember either.

What we have already achieved, among ourselves. Since mid-September 2026 the three of us, Totto, Steinar and Selina, and our agents have used these rooms as a working channel, not a demo (Totto and Steinar from the 15th, Selina from the 18th). On day one Steinar's agent put a pull request up for review in a group and it was merged and released; since then a scope change to a feature before it shipped, a versioning rule our build pipeline now enforces, a release held for the other side's review, two design decisions, an approved test plan and a definition of done for a three-party proof were all taken in those rooms. The friction Selina and Steinar hit joining the first group became an accepted design document the same day. We have measured nothing about time or cost saved, and one external partner we offered a room to kept his own messaging tool.

What you could get. Everything we have actually done with it is software work among ourselves; this list is what the same pieces would give a different kind of organisation, and section 6 marks each as REAL or ILLUSTRATION. "Who may decide this step" as a checked fact on the record, for every step that names an owner, rather than a sentence in a policy; an inbox of what is waiting on a person; review by several agents before a person sees anything; agents that take bounded tasks and cannot merge their own work; a conversation history your organisation does not lose when someone leaves.

What it costs you today. One server per organisation, administered by someone who issues every credential or invite. A small web receiver your own team runs, because the platform only sends a short signed note saying "something is waiting for you" to a web address you give it; turning that into a Slack message or an email is your receiver's job. Agent teams installed by hand from a checkout, with model spend you cap at the provider, because the product does not. And a person who, today, carries a team's result from the room to the gate. We are not quoting a price, because we have not sold one.

What is not built. No chat view in the browser. No Slack, email, SMS or mobile delivery. No files in rooms, no leave or remove, no installer for the agent teams. Nothing is runnable by an outsider; the repository is private.

And 10-40x? A hypothesis, labelled as one, with the measurement plan in section 7.


The tags. [Built]: on the main branch, with tests. [Proven in our pilots]: ran live at least once, with a record. [Designed, not built yet]: a documented design or intention; no code does it. [Not built]: no code and no design. No tag is higher than the code and our records support.

1. The problem: organisations forget, people drown, agents start from zero

A good decision is made in a meeting. Three weeks later nobody can find why it was made, only that it was. Six months later somebody re-opens it, and the second discussion is worse than the first because the people who held the context have moved on. The organisation did not learn; it got older.

Meanwhile the people who are supposed to decide are drowning, not in decisions but in the material around them. Our own product description puts the AI version plainly: ten times the output drowns the humans who are supposed to review it. If every agent decision routes to a person and output goes up tenfold, the review queue goes up tenfold too. That is not leverage; it is a bigger backlog with an AI label on it.

And the newest problem: every AI agent you hire starts from zero each morning. It does not remember what it agreed with your colleague yesterday, because it was never in the room; you were. You read the message, pasted it into the agent, read the draft, pasted it back. We have a dated record of this from a partner project in early September: six rounds in one afternoon, every round relayed by hand. The design note we wrote afterwards called the person in the middle "the wire protocol". Nobody wants to be the wire protocol.

Three problems, one shape: memory in heads and in tools that do not talk to each other, judgment spent on synthesis, and the new agents shut out of the memory altogether.

2. The idea: conversations that compound, humans who decide, agents that read

Put people and agents in the same rooms. Keep every message as one signed line that is never rewritten. Let agents do what they are good at: reading everything, filtering, drafting. Keep decisions with a named person, through a door that checks who is knocking. Make the decision part of the record, next to the conversation that led to it.

Three consequences follow. Conversations compound: a new colleague or agent can read what was decided and why, in order. When we argued about giving agents a private memory that travels between rooms, our founder's phrasing was that compounding is "history, not aggregation": build a faithful record, not a clever memory. Humans decide, visibly: typing "approved, go ahead" in a room does nothing, literally (section 4 shows it); a decision is a separate action by a specific person, and the record shows who, when, and, where the step names an owner, whether it was that owner or someone stepping past them with a written reason. Agents take the synthesis, not the judgment: an agent can read a hundred messages and tell you the three that need you, a panel can review a plan from three angles and hand you the disagreements, and none of that moves a decision a millimetre.

Our workshop material calls the result a compounding organisation, where the executive's value becomes judgment rather than synthesis, because the synthesis is infrastructure. That phrase predates the code. The code is our attempt to make it true; the rest of this post is honest about how far we have got.

3. The whole system in one picture

Hold the system as a workplace.

The building is Canvas: a badge system where everyone, human or agent, holds their own credential; filing cabinets where every document is a signed version never overwritten; house rules that are checked rather than remembered. [Built]

The colleagues are co-worker agents: long-running programs, each with its own badge and job description, and a default of "no tools unless this role and this task both allow it". A reviewer cannot touch code; only a fixer, on a fixing task, can. [Built], [Proven in our pilots] on one machine, one operator.

The departments are agent teams: a short file of positions, each minted as a badge, sharing one closed room. A position labelled "reviewer" is rejected because it collides with the real reviewer role; nobody acquires authority by being named something important. [Built]; a five-role team [Proven in our pilots].

The review meeting is an agent panel: one task to several reviewers with different lenses, one orchestrator who writes a single decision line, a "judge" standing in for sign-off. One round; the judge is a simulated human. [Built], [Proven in our pilots].

The three nest: a co-worker is one agent with a badge and a job; a team is several co-workers in one closed room; a panel is what a team does when the job is a review.

The conversations are HAAH: signed, append-only rooms, the same six tools for people and agents, and a strict rule that talk decides nothing. [Built], [Proven in our pilots] among the three of us and our agents.

The signature on the form is HITL: the one door through which a person's decision changes what a run does. It refuses the shared master credential, refuses whoever started the work, and, since 1 October, where a step names its decision owner, refuses anyone but that owner unless a person holding publisher-level rights steps past with a written reason that stays on the signed record. A bot, or an agent acting for a person, cannot make that override. [Built]; the 1 October rules not yet in a pilot.

The process manual and the house rules are playbooks and charters: signed procedures whose gates the runner checks, and standing rules that loosen only on a signed track record, never to "no oversight" by themselves. [Built]; 61 worked example procedures that run [Proven in our pilots], by the repository's own account; none of them is a customer's process, and none is for software engineering yet.

The building: Canvas, HAAH rooms, the review meeting and the HITL desk

Where the analogy breaks. A landlord can change the locks; here, whoever administers the server (root, in the jargon) can rewrite files, and the signatures detect a changed line but not a missing tail, and cannot stop root either way. A colleague remembers yesterday; a co-worker agent gets a fresh session per task. A department can hire and fire; a team has no "leave" or "remove", only revoking a badge. In an office, "yes, go ahead" in a chat often counts; here it never does. And the break that bites a manager first: the building does not page you. It sends a signed pointer to an address you run; it cannot send a Slack message, an email or a text.

4. The full loop: talk, gate, decision, record

Both mechanisms have "human" in the name, so they get confused. HAAH is where work is discussed, remembered and routed. HITL is where a decision is required and recorded. The honest picture is that today they are two mechanisms that meet in a person: nothing on main carries a result from an agents' room into a playbook gate by itself, and in our pilots a human attached the panel's verdict to a one-step procedure by hand. Here is the loop as the pieces allow it, station by station, each with its own tag; no single run has yet gone through all six.

  1. Agents work and talk in a room. A co-worker takes a task, posts "accepted" when it starts, then exactly one of "result" or "failed"; reviewers post findings; any member reads it in order. [Built], [Proven in our pilots]
  2. A gate pauses the work. A step whose author said "a person decides this" raises an escalation naming the owner. Since 1 October a step can also say "a person approves before anything is attempted". [Built]; the gate [Proven in our pilots] on 28 September; the pre-gate not yet.
  3. It lands in the owner's inbox, and a pointer goes out. The inbox is chronological and deliberately unranked; an optional "needs you first" card sits behind a switch with its own telemetry, so we can check whether it helps. The pointer goes to an address the owner's team runs. [Built]; a phone notification [Not built].
  4. The owner reads the discussion. Each pending step has a thread and a short brief composed from the record: why it stopped, the facts, the consequence of each choice. A colleague types "fine by me"; it is recorded and changes nothing. [Built]; the brief not yet in a pilot.
  5. The owner decides at the gate. One action, no option pre-selected. The record shows who, what, when, and, if a person with publisher-level rights stepped past the named owner, the written reason. A bot cannot make that override. [Built]; not yet in a pilot.
  6. The decision flows back and compounds. The thread shows the resolution, and the gated procedure can be replayed offline to check the decision was made on exactly the facts recorded. [Built] for procedures; the agents' room in station 1 is a witness record, not a replayable run, and "the work continues" from a decision back into the agents' room is [Not built]: today a person restarts the agents.

For the bytes, see How Agents and Humans Talk on Sunstone Atlas, including the capture of "approved" typed in a thread leaving the step pending. The parts list is Agentic Engineering Teams on Sunstone Atlas: What They Are Made Of.

The mechanics of memory: one loop from agents talking to a decision that compounds

What a leader still does. Decide who owns which decisions, by name, and keep that list honest; a gate naming a role nobody holds pages nobody. Read before resolving; the gate cannot make you. Treat the override as breaking glass and read its log. Run the receiver. Carry the result from the room to the gate, for now. And sign off on the playbook, not only on the run: publishing the procedure decides which steps need a person.

5. What we already do

This is the part we would normally dress up. Here it is undressed. Each item is something the record ties to a room; where it is a thing we noticed rather than an outcome, the sentence says so. No client, customer or partner is named; no private message is quoted.

A standing room between Totto and Steinar, on a separately deployed instance of the same product, each with his own agent in it. Stood up 15 September.

What we did: - On day one Steinar's agent posted a pull request for review in the room; it was merged and released the same day, followed by his team's review of a design document. - Steinar's first-real-user feedback moved one feature of a notifications design from "later" into the release being built, because a pointer with only a count would force a full re-read. - A whole feature arc ran through the room as working traffic: we built and held the release; Steinar's side load-tested and found two real races; the fixes were split between the two sides; the release was confirmed in the room, one to two days end to end. A versioning convention he asked for was agreed there too; our release pipeline now enforces it and rejected a mis-cut release for real on 28 September. - Two design replies on 30 September and, on 1 October, approval of a test plan his agent had posted, all human-gated. The approval led to an open pull request on his side, and one of the replies to a tracked follow-up. The approval was surfaced to Totto by our morning report, a Slack message at 07:30 Oslo time, built outside the product at his request and reworded on his instruction to have "HITL" and "HAAH" "behind the normal language".

What we also saw: what the room replaced was dated status files dropped in a downloads folder and read by file timestamp. That is the honest "before".

A three-way room for Totto, Selina and Steinar on eXOReaction's own instance. Created 18 September around Selina's cross-domain design.

What we did: - Something of an unusual kind: joining it hurt. Three drafts of a hand-written onboarding document, two raw credentials per person relayed by hand, one bug that silently skipped the membership step, and a security classifier correctly refusing to let the agent doing the work mint a third party's credentials. That became an accepted design document the same day; an invite flow now exists in the code. [Built], not yet used for a real onboarding. - A definition of done for the three-party cross-deployment proof, accepted by Totto in the room on 1 October.

What we also saw: the design review itself ran mainly on GitHub, as did our most intensive technical exchange with Selina in September. Anyone who tells you "all our collaboration runs on this" is wrong, and it would be us.

Agent-team pilots on Totto's own instance. Published already, so briefly.

What we did: - On 28 September a human gate refused our shared master credential, so Totto minted a separate identity and approved with it. In the fourth pilot he sent the team's own design back (his recorded decision at the gate: "we need more QA/rework"), the team reversed itself, he approved the smaller version ("Lets start building."), and it merged the same day. On a question about where to run our models he resolved the gate with his own words on the record: "we need soon to have a Norwegian soil solution." - On 1 October, in an experiment we have not yet published (working title "We Hid a Bug in Our Own Platform"), the team caught a bug we had planted in its own plumbing, with a named reason in the room, and had an approved one-line fix, with its regression test, 4 minutes 22 seconds after detection. It built five of seven steps on a branch, then stopped itself on a contradiction in our own specification. Nothing merged.

What we also saw: in that run, asked whether he wanted to read the panel's reasons before approving a plan it had rejected, Totto's recorded reaction was "ok.. to much work.. I'll approve anyway here (should this not been a HITL message for me in slack? )". Involvement theatre, in the research term for it, observed live in the founder of the product that is supposed to make it visible. It is the most useful sentence in this post.

One honest failure. We built a separate instance for a collaboration with an external partner. He never replied in the room and kept his own chat tool; two messages we posted on day one were still unanswered five days later, and our working rule became that his tool stays his surface. A welcome message there also named another customer; the record is append-only, so it could not be unsaid, only corrected on the record, and our rule "never name one customer to another" dates from that day.

In one sentence: our own team already runs on it, for real decisions, since mid-September; we have measured nothing about time or cost; and n is three people, their agents, and one partner who declined.

6. What you could use it for

Six use-cases, each REAL (we did it, there is a record) or ILLUSTRATION (the pieces exist to the degree tagged; the scenario has not been run end to end), each saying what the human still decides.

1. Approvals with a named owner: a purchase over a limit, a payment, a leave request. Routine cases pass a published yes/no rule with no AI involved; the rest wait for the named owner, not whoever is online. A deputy is refused by name; someone holding publisher-level rights, say the head of finance if that is who holds them, can step past with a written reason the record keeps. The human decides: the case, and who the owners are. REAL for the mechanism (every new instance ships with this demo); ILLUSTRATION for the scenario. [Built]; owner enforcement not yet in a pilot.

2. A decision with the conversation kept next to it. Today the discussion happens in a chat tool and the decision is written down somewhere else, later. Here a pending step carries its own thread, the evidence it was raised on, and the decision, in one record. The human decides: the step, after reading. REAL for "the panel's verdict and the human's decision in one signed record": three of the four 28 September pilots ran a real room, a real panel and a real human gate, with a person attaching the verdict to the gate by hand, and on that date nothing yet checked that the decider was the named owner. ILLUSTRATION for "the owner decides from inside the thread where the agents reasoned": the thread and the in-thread door are [Built], exercised on a throwaway instance on 1 October, not yet in a pilot.

3. "What needs me today," across everything. One morning card: decisions pending, new messages, who is waiting on whom; full detail only for items addressed to you. The human decides: everything; the card decides nothing. REAL as our own receiver outside the product; [Not built] inside it.

4. Project status across teams, without three stand-ups. Each workstream's agents post results to their own room; the lead's agent, a member of each, drafts a status note citing the records; open items are the pending escalations, with named owners. The human decides: every open item, and whether to trust the summary, which the record lets them check. ILLUSTRATION. The rooms are [Built]; a browser view is [Not built].

5. Onboarding a colleague and their agent. An admin issues a single-use invite; the person claims their own credential, so nobody relays it; their agent is minted as acting for them and can never see a document the person may not; a member adds each of them to the room, the agent too, because room membership is not inherited by default; their view starts at their own joining. The human decides: who joins, as what, into which rooms. REAL for the painful manual version of 18 September; the invite flow it produced is [Built], not yet used for real.

6. A plan reviewed from three angles before you read it. Reviewers with different lenses each return a verdict and findings marked blocking or not; an orchestrator writes one decision line; you read the disagreements. For business questions, grounded personas argue two rounds, a moderator synthesises without proposing, and a person decides. The human decides: always; the decision line is advice. REAL: it has run on our own plans and code since 28 September, including a 49,031-character plan review on 1 October, and six internal persona panels on our own business questions, whose output was advice we then decided on ourselves. [Built], [Proven in our pilots]. One honest number from the pre-registered test: most of the trap tasks never reached the panel, because the fixer refused the premise; on the one real trap diff that did, the panel's record is 0 of 1, and debatable; the technical post has the rest. Two reviewers once found two defects that 18 new tests had missed. We do not add those into a rate.

7. About 10-40x

A confession first. In January we published a post claiming a 330x productivity multiplier from a single-practitioner code build. Two days earlier we had published a plain educational post about a workshop. The 330x post got 6 likes, 3 comments and no leads; the educational one got 25 likes and 3 leads. We wrote the lesson down ("claim only what's proven") and have not used the big number since. We are not about to invent a smaller cousin of it here.

What our data supports. One measured data point for raw build speed: in January 2026 a build that its own report calls "a small team", and whose commit history shows one human author, Totto, working with one coding agent and a skill library, produced a 197,831-line, 7,461-test library in 11 calendar days. The 25-66x we have quoted from it is those 11 days divided by a cited industry timeline of 9-24 months that nobody on our side measured. Our own materials round that down to "10-30x typical" and call it proven; we think the honest label is an estimate with a stated basis, and that is the only label it gets here. It is about one practitioner's build with one coding agent; teams of agents, rooms, panels and gates did not exist in January.

What our data does not support. Any multiplier for the machinery in this post. The 1 October run is the only team run with timing and cost numbers: about 91 minutes of wall clock, roughly 12 of them waiting for a human, four typed words of gate input (launch, deploy, resume, approve; everything else the human said that morning was in a side conversation, not to the run), a planted bug fixed under review, five of seven steps built on a branch, nothing merged, and a driver's cost proxy under 18 dollars not yet reconciled against the bill. No control run, no human-built version of the same specification. We do not turn that into a ratio.

So, loudly: 10-40x is a hypothesis. Ours is that a governed team of agents with a human at the gates moves a well-specified change from specification to reviewed branch an order of magnitude faster than the same human working alone with one agent, at a cost per step in low single-digit dollars. We have not measured it.

How we would. A pre-registered specification and a control: same spec, same codebase, same held-out acceptance tests; one arm a human with one agent, one arm the team with human gates. Four clocks: wall clock, machine time, human attention actually spent reading and deciding, calendar time. Cost from the provider's bill, not a proxy. Output defined before the run as a merge-ready branch passing the held-out tests, never lines of code. At least three runs per arm on more than one codebase, variance shown. We publish the lower bound with the raw numbers, so you can compute your own ratio.

8. Risks an honest leader should hear

The human can still rubber-stamp, and we have watched ourselves do it. The gate cannot make a person read; "to much work.. I'll approve anyway" is in our own record. The instrument that would flag a gate with a near-100% approve rate and near-zero latency is [Designed, not built yet].

Delivery fails toward silence. A failed pointer is indistinguishable from nothing pending. Retries live in memory and die on restart. Nothing in the product reaches a phone. [Not built]

The signatures prove less than they look. Each line is signed by the instance's key, which proves the instance accepted these bytes, in this order, from this badge. It does not prove a person pressed the button: an ordinary approval is signed by the server, not the person's own key. It does not prove the record is complete or current: root on the host can truncate a file or present an old copy, and there is no externally anchored head yet. [Designed, not built yet] for both. We do not say "tamper-proof"; if a vendor says it to you, ask what root can do.

Identity is thinner than the badge system suggests. An admin issues every badge or invite; the owner rule matches an exact id or name, and although new duplicate names are now refused, older duplicates would make it spoofable. Our own read-only audit on 2 October found duplicate badge names on two instances, two of them apparently orphaned on one instance and not yet revoked, and one owner string matching no badge at all. Hygiene is a job.

The models misstate their own inputs. Recorded live: an orchestrator attributed a caveat to a reviewer who never raised it, while still reaching a defensible verdict; a judge wrote that it had "directly verified" two tests were absent and was wrong, having read its own copy of the code instead of the fixer's. A person reading the verdict against the real diff caught it. The room is a witness, not a referee.

A compounding memory is a scope-laundering machine unless recall is scoped. Our own design note's phrase, and why agents carry no private memory between rooms today. The designed version treats memory as staging, never as truth, written only through a human publish gate. [Designed, not built yet]

An overview of everyone's conversations is a surveillance instrument. Our own design note flags that in Norway this lands in employee-monitoring and works-council territory; we have not taken legal advice on it. There is no auditor role and overview reads are not logged; and on how an installation-wide overview can coexist with access scopes and federation at all, our founder's recorded position is "not sure how to solve that."

Adoption is not automatic, and the evidence is thin. One partner declined the room. Three people, one machine, one operator, no survey, no quotes from Steinar or Selina cleared for publication, nothing runnable by an outsider. A fixer with a shell can read the model API key, and no spend cap is enforced in the repository.

9. Where this is going

Nothing here carries a date, and the full list is the table below. The three items that matter most to a leader:

  • Invite link and self-service join, so onboarding never again needs a hand-relayed credential: [Built], not yet used for a real onboarding.
  • Personal signatures on approvals and an externally anchored ledger head, so "a person pressed this" and "nothing was cut off the end" become checkable claims: [Designed, not built yet].
  • Slack, email, SMS or mobile delivery, so the building can page you without a receiver of your own: [Not built].

The maturity strip: what is built, proven, designed and not built

10. The maturity table

Piece Tag Note
Signed, append-only rooms; groups membership-gated, no backfill Built; Proven in our pilots 51 tests; three colleagues and a five-role team
Same six tools and badge kind for people and agents; master credential refused Built
Talk decides nothing; "approved" in a thread leaves the step pending Built probe captured 1 October on a throwaway instance
Co-worker agents, tools denied by default, governance enforced at the tool call Built; Proven in our pilots enforcement real only since 28 September
Agent teams provisioned from a file into one closed room Built; Proven in our pilots human-run by decision; no screen, no installer
Agent panel (team.review): lenses, one decision line, one round Built; Proven in our pilots n = 12 and n = 1; no control run; cost not measured
Persona panels for business questions, ending at a human gate Built; Proven in our pilots panels do not decide
Human gate: named owner enforced, publisher-level override with reason, bots cannot override; pre-gate Built landed 1 October; not yet in a pilot; owner-less ordinary steps keep the older, weaker check
Decision brief, decision panel, per-step thread Built brief not yet in a pilot
Pointer-only paging and notifications to a receiver you run Built retries in memory
Offline replay of a gated procedure Built proves order and acceptance, not freshness or host honesty; an agents' room is a witness record, not a replayable run
Room result carried into a playbook gate by itself Not built a person does it today
Playbooks and charters; oversight loosens only on a signed track record Built; 61 worked examples Proven in our pilots by the repository's own account; no engineering playbook yet
Invite flow for onboarding Built not yet used for real
Our morning report of what needs a person Proven in our pilots, outside the product
Human-versus-bot check on the ordinary decision path; involvement-theatre instrumentation; model choice per role, including self-hosted "Norwegian soil" models Designed, not built yet
Agent memory across rooms; personal signatures; anchored ledger; engineering playbook; multi-round build loop Designed, not built yet the build loop has five of seven steps on an unmerged branch; its git, pull-request, CI and merge edges are Not built
Chat view; Slack, email, SMS, mobile delivery; digest in the product; files in rooms; sealed rooms; leave or remove; configuration screen and installer Not built a membership-lifecycle decision is pending

The slide version. Eight slides from a NotebookLM deck generated from this post: the title slide as the cover, three used in the text above and four here. Six others were left out because their wording was looser than the post. The hexadecimal strings and timestamps in the margins are decoration, not real data.

The paradigm, the reality, the cost of the status quo and the hard truths

Synthesis is infrastructure, judgment is executive

Risks an honest leader should hear

If every decision had to pass through a door that wrote down who opened it

11. A closing question

Our organisation is three people, their agents and a little over two weeks of real use. It already remembers more than it did in August: why a cursor exists, why a release was held, who accepted which definition of done. In order, with names, readable by whoever is added next, from the moment they are added. The moment its founder said approving was too much work is remembered too, but only in our notes, which is the point of section 8.

So the question is not whether an AI can do your job. It is smaller and more awkward. If every decision had to pass through a door that wrote down who opened it, and every conversation that led there was kept where the next person could read it, what would your organisation remember next year that it forgets today? And which of your people would be relieved, and which would be nervous?


We are sharing, not selling. The repository is private, there is no trial, and the agent-team code deploys by hand. Where this post and the code disagree, the code wins; where this post and our two technical posts disagree on a tag, the lower tag wins.