ARA — ANALYTICS · RESEARCH · ARCHITECTURE

Building in
the age of agents.

Getting AI to do impressive things isn’t hard anymore. Getting it to do the right thing every time, in front of your customers, is. That’s where we come in: data and AI architecture for organizations that want to move fast without cutting corners on trust.

  • A / ANALYTICS: so you can see what’s actually going on.
  • R / RESEARCH: so you know what’s real and what’s hype.
  • A / ARCHITECTURE: so what you build holds up.

BIG-COMPANY AI EXPERIENCE, FOR SMALL BUSINESSES AND NONPROFITS.

FIG. 01 — THE THESIS

YOU CAN’T BOLT TRUST ON AT THE END.

In the age of agents, trust is the architecture.

Anyone can wire up an AI agent in an afternoon. Building one you’d trust with your customers, your money or your data takes real design work up front. That’s the part we do.

T-01 · GROUNDED

Is it working from data you can trust? When AI gets something wrong, the real cause is usually bad data underneath it.

T-02 · BOUNDED

Does it know what it’s allowed to do, and does it stop there? Clear limits, the right permissions, and a way to hand things off to a person are what keep a helpful assistant from turning into a liability.

T-03 · ACCOUNTABLE

Can a person see why it did what it did, and stand behind it? That record should be there from day one, not pieced together after a lawyer asks for it.

Grounded. Bounded. Accountable.

If it doesn’t pass all three, we don’t ship it.

FIG. 02 — TRACK RECORD

30 years

Building enterprise data, analytics and AI systems, first hands-on and later as an advisor to executive teams. ARA was built on that experience.

We’ve been at this through every wave, at companies where guessing wasn’t an option.

We go back to when data warehouses were built by hand, through the analytics and machine-learning years, and into today’s AI agents. Our founder spent three decades advising and building for some of the world’s largest media, telecom and tech companies, where a bad call can cost millions and you don’t get many second chances.

  1. 1997 — DATA WAREHOUSING & BI
  2. ≈2012 — ANALYTICS & ML
  3. TODAY — GENERATIVE & AGENTIC AI
  1. 01

    C-suite advisory at the world’s largest media & entertainment enterprises.

  2. 02

    Analytics and AI for major telecommunications carriers.

  3. 03

    Data architecture for global high-tech platforms with audiences in the hundreds of millions.

That’s where we learned the craft. Now we bring the same discipline to small and mid-sized businesses and nonprofits, who usually can’t get this kind of help without a big-company budget.

FIG. 03 — WHAT ARA DOES

Three ways we can help.

We don’t resell anyone’s platform, so we’re free to recommend whatever fits: Databricks, Snowflake, BigQuery, any of the major clouds, open source, or the leading AI platforms.

01 · CROSS-PLATFORM ANALYTICS

Dashboards, reporting and analysis built on the tools you already have, or on whatever fits best if you’re starting fresh.

02 · DATA & AI TECHNICAL ARCHITECTURE

We design the data and AI systems underneath your agents so they’re reliable and safe from the start instead of patched later.

03 · DATA & AI RESEARCH & ADVISORY

Straight advice on what’s out there: what actually works, what’s still hype, and what’s worth trying now versus waiting on.

FIG. 04 — SUCCESS STORIES

Things we’ve actually built.

Abstract illustration: a phone outline holding a two-by-three grid of six person tiles above one full-width help bar

6 photo tiles — the hard-coded contact limit

CareDial: the phone app with six buttons

An iPhone app for an aging parent: six photo tiles, one help button, and a caregiver who manages it all from their own phone. Most of the design work went into deciding what to leave out.

See how it’s built

FIG. 04.1 — CAREDIAL

ENGAGEMENT

ARA (internal product)

STACK

iOS / SwiftUI · CloudKit

TIMELINE

2026-07

01 — CHALLENGE

For a lot of older people, a smartphone causes more confusion than it solves. The contacts app scrolls forever, a mis-tap opens something unrecoverable, and the person who could fix it lives in another city. So the phone stops being used for the one thing that matters most: calling family. CareDial started as a fix for our own family. Elderly parents needed a simple way to call, text and FaceTime a few key people, and we needed a way to manage those contacts remotely. The fix turned out to be less software. We built it for us, then polished it so other families could use it too.

02 — APPROACH

CareDial is an iOS app with two faces. On the parent’s phone it is one screen: up to six large photo tiles under the question “Who do you want to talk to?” Tap a face, the call starts. A caregiver can enable FaceTime or texting per household, and a red Help button calls a designated person, either with a confirm step or by dialing straight through. Settings exist but are deliberately hard to find: a low-contrast gear icon that opens only after a triple tap, so the parent cannot wander into configuration by accident.

The other face is the caregiver’s. On their own phone they add, edit and reorder contacts, set the help contact, and switch features on. Changes sync to the parent’s device through Apple’s CloudKit, over a shared zone the two phones pair with a short code. The parent’s side is read-only, and the app keeps a local copy of the contacts so it still works when sync doesn’t. If the parent does reach the settings sheet, they see one calm message: the contacts on this phone are managed by their caregiver.

03 — WHAT WE DELIVERED

A complete, working SwiftUI app: role selection and pairing, the parent’s call screen, the caregiver’s management dashboard, CloudKit sync with local caching, and the help-button safety path. It’s native Swift throughout with no third-party dependencies, so there’s nothing to maintain but the app itself.

04 — OUTCOME

CareDial shipped. It is live on the App Store and it is free. It launched recently, so the only users so far are the family it was built for, which is the right place to start. It’s also a good example of how we approach client work. Most apps keep piling on features. This one works because it doesn’t: the contact limit is hard-coded at six, on purpose.

“Intentionally low-contrast so elderly users overlook it. Caregivers triple-tap to open Settings.”

COMMENT IN THE CAREDIAL SOURCE CODE

CAPABILITIES

iOS / SwiftUI development · CloudKit sync & device pairing · accessibility-first product design · role-based UX · safety-path design

Abstract illustration: a leveled scorecard of seven bidder bars against a column grid, three struck through, the top-ranked row marked

7 bidders scored — three set aside with documented reasons

Bid leveling a board can defend

An owner’s rep firm was comparing contractor bids by hand. We built a scorecard they can rerun on every project. It does the math, flags anything that needs a human call, and explains its results in plain English for a volunteer condo board.

Read the leveling story

FIG. 04.2 — BID LEVELING

ENGAGEMENT

An owner’s representative firm in Florida — cross-platform analytics solutions

STACK

Python · CSI MasterFormat taxonomy · Cowork plugin

TIMELINE

2026-06

01 — CHALLENGE

The client is an owner’s representative firm that manages restoration and renovation projects for condominium and HOA boards in Florida. When a project goes out to bid, contractor proposals come back long, inconsistent, and organized however each contractor pleases. The firm leveled them by hand: copying line items into Excel, reformatting, rebuilding formulas. Then it had to explain the result to volunteer board members with no construction background, who carry fiduciary responsibility for multi-million-dollar awards under Florida’s post-Surfside condo statutes. It was slow, it was easy to make mistakes, and it was hard to defend if anyone asked where a number came from.

02 — APPROACH

We didn’t build an AI that reads the bids and picks a winner. We split the work the way a careful estimator would. Everything mechanically present in the bid matrix is computed: totals, price per square foot, price tiers, variance, ranking. Everything that is judgment stays labeled as judgment. The cost baseline and the square-footage basis come from the firm’s estimator each run, and the qualitative ratings are drafted by AI, then reviewed and overridden by a human before any board sees them. Line items map to a CSI MasterFormat taxonomy extended for restoration trades, with leveling rules that keep base bids, alternates, allowances and exclusions in separate buckets. If a required input is missing, the tool stops and asks. It never guesses.

03 — WHAT WE DELIVERED

A repeatable bid scorecard the firm runs on each new project, installed on their own machines as a Cowork plugin. They upload the bid files and press play; leveling that took hours of copying and reformatting now takes minutes. The scorecard ranks the bidders and drafts an award recommendation, a capability the firm didn’t have before this project. The handoff package includes an operating guide, a sample output, and a plain-English assumptions and methodology memo written for board members, which separates hard arithmetic from informed judgment section by section. We also delivered a morning business-pulse digest for the firm’s leadership.

04 — OUTCOME

It proved itself on the very first project we tested it on. The tool’s independence check found that the supposedly independent cost baseline matched one bidder’s subtotal almost exactly, a strong sign the yardstick had been anchored to that bidder’s own number. It surfaced a duplicate bid from the same contractor with two different totals. Seven bidders were scored and three were set aside with documented reasons. The board received a memo that states plainly which numbers are calculated, which are estimates, and which still need confirmation before an award. The time savings got their attention, but the error-checking is what sold them: the leveling rules catch miscategorized line items, flag entries for review, and find errors with suggested corrections. They are rolling the tool into their standard bid process now.

“The computer does the arithmetic; the human supplies the professional estimates.”

ARA METHODOLOGY MEMO, PREPARED FOR THE BOARD

CAPABILITIES

document intelligence · bid leveling & normalization · CSI MasterFormat taxonomy · Python data pipeline · AI-drafted, human-reviewed scoring · board-ready reporting · Cowork plugin packaging

Abstract illustration: three numbered agent nodes in a vertical build-order column, wired by right-angle lines to one hub 01 02 03

AI agents in Copilot and Claude that take busywork off a small team

More hands for a four-person team

A four-person team at a global nonprofit knew exactly what they wanted to automate but had no time to do it. We found where they were stuck, trained them hands-on, and helped them build AI agents to take on the work they couldn’t get to.

Read the enablement story

FIG. 04.3 — AGENTS AND AUTOMATION

ENGAGEMENT

A nonprofit digital-outreach organization (Asia-Pacific) — data & AI research and advisory (pro bono)

STACK

M365 Copilot · Agent Builder · Power Automate · Planner · Power BI · Claude Code

TIMELINE

2026-07

01 — CHALLENGE

The client is a nonprofit digital-outreach organization whose Asia-Pacific team is four people running Meta and Google campaigns to audiences across the region, in multiple languages. The front half of the loop took all their time: content, ad processing, campaign management. The back half kept going cold. Following up with people who respond, reconciling campaign data against donor records and reporting all sat at the end of a queue that never emptied. The team had already drawn its target workflow on a whiteboard.

02 — APPROACH

Before talking tools, we figured out the real problem. They didn’t need anyone to explain what AI could do. They already knew. They just didn’t have the hours. We produced an AI opportunity map that named the highest-value builds in priority order, audited the team’s actual Microsoft licenses so every recommendation matched what they could use, and designed a three-day hands-on workshop around a simple stack: Planner as the task backbone, Power Automate as the plumbing, a Copilot agent grounded on the team’s own SharePoint content, and a Power BI builder day. A license contingency kept the workshop standing even if the premium Copilot licenses fell through. From the first email exchanges to onsite delivery took about a month.

03 — WHAT WE DELIVERED

The opportunity map, with three named agents in build order. The full workshop design and facilitator materials. The license and capability audit.

04 — OUTCOME

The workshop ran onsite over three days, and each night we reworked the material based on what came up that day. The team came away building. They stood up a communications agent in Copilot, grounded on a SharePoint knowledge library of their own communication style, tone and SOPs. That agent was on the opportunity map, and they built it. Then they went past the roadmap entirely and built a subagent team in Claude Code that handles translation, updating large numbers of files across many languages. We can’t share much about their work, but we can share the time: an operation that took days by hand now takes minutes. Their biggest challenge is still finding people to support the work, but the agents are already saving them real time. The engagement was pro bono.

“The map, but not the hands to walk it.”

THE TEAM’S LEAD, DESCRIBING THE GAP

CAPABILITIES

AI opportunity mapping · M365 Copilot & Agent Builder · Power Automate & Planner · Power BI · Claude Code subagent teams · workshop & curriculum design · license/capability audit

FIG. 05 — INSIGHTS

Field notes from work we shipped.

FIELD NOTE 001
6 MIN READ · 2026-07-06

Apples to apples has to be built

We taught AI to compare 100-page contractor bids. Reading them was the easy part. Writing down the rules was the real work.

Read the note

A model can read a hundred-page contractor bid in seconds and give you a clean summary of it. That stopped being interesting a while ago. What it cannot do on its own is tell you whether two of those bids are describing the same building.

That’s the part that takes real work.

We built a bid-leveling tool for an owner’s representative firm in Florida that manages restoration work for condominium and HOA boards. Seven proposals arrive for a project. Each contractor organizes its numbers the way its own estimating shop prefers. The person who eventually has to choose is a volunteer board member with no construction background and a fiduciary duty for a multi-million-dollar award under Florida’s post-Surfside statutes. The job was to give that member a ranking they could defend.

WHERE COMPARABILITY ACTUALLY COMES FROM

The bids don’t arrive comparable. You have to make them that way.

Take one scope of work. One bidder prices it as a line in the base bid. Another moves it to an alternate, so it only appears if the board elects it. A third carries it as an allowance, which is a placeholder with a number attached rather than a price. A fourth writes it into the exclusions and never mentions it again. Four honest bids, four different totals, and none of the differences are about who is cheaper.

So the first real decision is the taxonomy. We mapped line items to CSI MasterFormat, extended for restoration trades. The standard covers the industry; the extension is where the actual work went, because the mapping is a judgment about what a given line item is, made once and then applied consistently.

The second decision is to keep base bids, alternates, allowances and exclusions in separate buckets and never let them merge. Collapse them and the word “lowest” stops describing price. It starts describing which bidder chose to include the most.

Neither of those is an AI problem. They’re construction know-how, and someone has to write it down before you can automate anything.

WHAT THE COMPUTER DOES AND WHAT THE HUMAN DOES

We drew that line on purpose and then made the tool enforce it.

Anything mechanically present in the bid matrix is computed: totals, price per square foot, price tiers, variance, ranking. Arithmetic does not need an opinion, and asking a model for one is how you end up with a number nobody can trace.

Anything that is judgment stays labeled as judgment. The cost baseline and the square-footage basis are supplied by the firm’s estimator on every run. They are not inferred, and they are not remembered from last time. The qualitative ratings are drafted by AI and then reviewed, and overridden where needed, by a person before a board sees any of it.

And when a required input is missing, the tool stops and asks. It never guesses.

That sounds like a small detail, but it’s what makes a professional willing to sign off on the output instead of double-checking all of it.

BUILD THE CHECK THAT CAN EMBARRASS YOU

On the validation project, the independence check found that the supposedly independent cost baseline matched one bidder’s subtotal almost exactly. That is a strong signal the yardstick had been anchored to the number it was supposed to be measuring. The tool also surfaced a duplicate bid from the same contractor carrying two different totals. Of the seven bidders scored, three were set aside with documented reasons rather than quietly vanishing from the comparison.

So add validation, but point some of it at your own inputs as well as the bidders’. Those are the mistakes nobody else in the process is looking for.

THE OUTPUT HAS AN AUDIENCE WITH LIABILITY

The reader of this work is a volunteer, not an analyst, and they carry real exposure for the decision. So the handoff package includes a plain-English assumptions and methodology memo that separates hard arithmetic from informed judgment, section by section, and states which numbers still need confirmation before an award.

We could only write that memo because every number in the pipeline traces back to the line item it came from and the rule that moved it. Try to add that kind of traceability at the end and you’ll be describing a process you can’t actually reconstruct.

WHAT GENERALIZES

Having AI read documents is cheap and easy now. Comparing them fairly isn’t, and a better model won’t fix that. It comes down to rules, and those rules live in the heads of experienced people who’ve usually never been asked to write them down.

If you are pointing a model at a stack of documents and expecting a ranked answer, the question to settle first is what would make any two of those documents comparable, and who decides. Answer it properly and most of the build is mechanical. Skip it and you get a confident ranking of things that were never the same.

The firm noticed the time savings first: leveling that took hours of copying and reformatting now takes minutes. What won them over was the checking, which is slower to appreciate and harder to demo. Being fast got us in the door. Catching a mistake before it reached the board is why they made the tool part of their standard process.

FIELD NOTE 002
5 MIN READ · 2026-07-13

Enablement first, then a team of agents

A small nonprofit team didn’t need another automation plan. They needed hands-on training first, then a set of AI agents in Copilot and Claude to take work off their plate.

Read the note

Picture four people running ad campaigns across a whole region, in several languages, with the workflow they wanted already drawn on a whiteboard.

They had the plan. What they didn’t have was time.

The client is a nonprofit digital-outreach organization whose Asia-Pacific team was carrying all of that. The front half of their loop, the content and the campaign work, consumed everything they had. The back half kept going cold: following up with the people who responded, reconciling campaign data against donor records, reporting on any of it. Their lead put it this way: they had the map but not the hands to walk it.

FIND THE REAL PROBLEM FIRST

From the outside, a team that doesn’t know what’s possible looks a lot like a team that knows exactly what it wants but has no time. Both are stuck, but they need very different help.

The default consulting move is to hand over an automation plan. For this team, that would have been a well-made answer to the wrong problem. They did not need to be told what to build. Telling them again would have added a document to the pile they already had no hands for.

So we started there. The problem was time, not know-how, and everything else followed from that.

AUDIT THE LICENSES BEFORE YOU DESIGN THE CURRICULUM

We audited the licenses the team already held, so every recommendation matched software they could use. We also built a license contingency into the workshop design, so the three days still worked if the premium licenses did not land in time.

This isn’t really about Microsoft. If your plan depends on someone else’s purchasing decision, it has a hole in it. Build around what people have today and treat any upgrade as a bonus.

NAME THE AGENTS IN BUILD ORDER

The opportunity map we produced named the highest-value builds in priority order.

A list of a dozen ideas gets read once and filed away. A list that says “build this first, then this” actually gets used, because a team with no spare time can’t afford to spend it deciding.

WE SPENT THE WORKSHOP BUILDING

Three days, onsite, hands on keyboards. The stack was deliberately plain: Planner as the task backbone, Power Automate as the plumbing, a Copilot agent grounded on the team’s own SharePoint content, and a Power BI builder day.

We kept it simple on purpose. Every new tool is something those same four people have to maintain, and a clever setup that eats an afternoon a week isn’t worth it.

We reworked the material each evening around whatever had come up that day. If your workshop plan doesn’t change once you’re in the room, you probably weren’t listening.

GROUND THE FIRST AGENT IN THEIR OWN MATERIAL

The team came away building. They stood up a communications agent in Copilot, grounded on a SharePoint knowledge library of their own communication style, tone and standard operating procedures. The opportunity map had named that agent; they built it.

Building it on their own material is what makes it more than a generic chatbot. It’s also what won over the skeptics on the first day.

THE PROOF IS WHAT THEY BUILD THAT YOU DIDN’T PLAN

Then they went past the roadmap entirely. They built a subagent team in Claude Code to handle translation, updating large numbers of files across many languages. An operation that took days by hand now takes minutes.

None of that was in the plan, and we weren’t there when they built it. That’s how you know the training worked: the team crossed onto a second platform on their own initiative and got something running.

And they didn’t build just one agent. They built a team of them with the work split up, which is harder to design and does a lot more for four people.

WHAT GENERALIZES

When a team is short on time, handing them a document doesn’t help much, because someone still has to do what it says. Helping them build the thing does.

Using two platforms wasn’t a principle we started with. It’s just where the work led. Copilot made sense because the work already lived in M365. Claude Code made sense because the file-scale language work needed something else. Committing to one vendor’s complete answer would have cost this team the second half of what they now have.

One limit worth being clear about: the team’s biggest challenge isn’t software. It’s finding people to support the work, and no agent fixes that. What the agents do is give the team hours back, and those hours can go toward finding those people. The engagement was pro bono.

FIELD NOTE 003
7 MIN READ · 2026-07-20

Running a company on an AI team

ARA is one person and a team of AI specialists, each with a name and a job description. Here’s how that works day to day, including the parts that don’t.

Read the note

ARA has one human. The rest of the firm is a roster of AI specialists, each with a written job description, a defined scope, and a specific set of tools it is permitted to touch.

This is what it’s like to run, including the parts that don’t work.

THE ORG CHART IS A FOLDER OF FILES

Every member of the team is a file. The file holds a name, a persona, what the role is for, what it must never do, and which tools it can use. There is a chief of staff, a head of people, a researcher, a designer, a strategist, a delivery and business-operations lead, a data scientist, a solution architect who builds, a technical reviewer, a web developer, an education lead, and a handful of domain specialists added for particular kinds of work.

Adding a new role isn’t like hiring. It’s editing a folder, and because the folder is under version control you can read back why a role exists and what it was originally scoped to do.

The part that turned out to matter more than expected is that roles get researched before they get written. The researcher produces an evidence-based brief on what a strong human practitioner in that domain actually knows, uses and produces. Only then does the head of people turn that brief into a role definition. Skip that step and you write a job description out of your own assumptions about a field you do not practice, which is exactly the failure a one-person firm is most exposed to.

ONE RULE DOES MORE WORK THAN ANY OTHER

The orchestrator does not do the work.

The chief of staff routes, briefs, sequences hand-offs, and brings results back, and does not write the deliverable, run the analysis, or make the domain call. We made that a standing rule, because the temptation to just do it is constant and it is strongest exactly when the work is urgent.

This isn’t about being a purist. The moment the orchestrator starts producing, you lose the record of who did what, and you lose the specialist framing that made the specialist worth having.

THE ARCHITECTURE FORCES HUB AND SPOKE

Two constraints shape everything else. Specialists cannot call each other. And they cannot see the conversation that produced their assignment.

So every hand-off passes through the orchestrator, who physically carries the artifact from one specialist to the next. We didn’t pick hub-and-spoke. The setup forced it on us.

So everything depends on the brief. If a file path is not named in the brief, the specialist does not have it. If a hard rule is not carried into the brief, it does not apply. The most common failure mode in this firm is a thin brief, not a weak specialist, and it took a while to stop misdiagnosing the first as the second.

Our fix is a written rule: if a brief is ambiguous, come back and ask instead of guessing. It works, imperfectly. A capable model handed an under-specified task will produce something plausible, and plausible-but-wrong is the most expensive kind of mistake.

NOTHING SHIPS WITHOUT A GATE

One member reviews every deliverable before it goes anywhere. The verdict is approve, approve with conditions, or reject with remediation. Remediation travels back to the original author through the orchestrator, and the work re-enters the gate.

When the reviewer is also the builder, someone else cross-checks the work against the same criteria. Reviewing your own work doesn’t count.

The gate covers more than code. Anything ARA-branded is checked against the brand canon. Anything outward-facing is checked for whether it reads like a person wrote it, because writing that sounds machine-generated isn’t finished.

The downside: the review is slow, and it’s the step we’re most tempted to skip. Its value shows up in the work it sends back, which is hard to appreciate in the moment.

EXECUTION IS RATIONED ON PURPOSE

There are four execution postures, and every member sits in exactly one.

Builders may write and run code, but only inside an isolated worktree, under a sandbox, with a deny list for destructive commands that always overrides whatever permission has accumulated. Authors write text and execute nothing. The reviewer runs validation and writes findings, never the artifact under review. The control-plane engineer writes and runs code in the main tree itself, within a defined scope.

Builders also have to verify their own work before handing it over: turn the task into criteria you can check, then loop until the criteria pass. That does not replace the gate. It just means the gate is not the first place a problem gets found.

THE STANDARDS ARE MOSTLY ABOUT RESTRAINT

Seven rules bind everyone who builds. These are the first four. Think first and state your assumptions; if two readings of the brief exist, present both rather than silently picking one. Ship the minimum that solves the problem, and if two hundred lines could be fifty, rewrite it. Change only what the task requires, so every changed line traces back to the request. Turn the task into verifiable success criteria and work until they are met.

Each of those four exists to stop the same behavior. A capable model will cheerfully produce more than you asked for, and most of what it adds is work someone now has to read, review and maintain.

CONTEXT IS THE SCARCE RESOURCE

This is the part that surprised us most about running the thing.

Hours are not the constraint. What every agent must carry into every session is. So new knowledge starts cheap, as a reference document read on demand. It gets promoted to a skill only when it is a repeatable procedure with a clear trigger. It goes into a specific agent’s definition only when that agent must always know it. It reaches the firm-wide charter only when every agent must. Each rung up is a permanent tax on more sessions.

Demotion is allowed and used. A rule that stops firing goes back down the ladder. Left alone, the rulebook grows until nobody reads it, just like every employee handbook ever written.

WHAT COMPOUNDS

Every gated deliverable produces a short capture file: what we learned, what is reusable and in what form, what idea came out of it worth developing, and what we would scope differently next time. The gate checks that the file exists.

“Nothing reusable came out of this” is a legitimate answer, and the check verifies that the question was answered rather than that an asset was produced. That detail matters more than the mechanism. A capture requirement that demands a reusable asset every time generates fake ones, and a library of fake assets is worse than no library.

We do this because everyone rents the same AI models. What belongs to ARA is the written judgment: the definitions, the rules, the captured patterns, the things we now know not to do again.

THE PARTS THAT DON’T WORK

Debris accumulates. Builder isolation creates working copies, and they pile up. At one point there were dozens of stale ones on disk, enough to start breaking shell commands outright. They now get pruned on a schedule, because nothing about that problem announces itself until it does.

Permissions drift. Every approval click appends to a local allow list. Left alone, that list quietly becomes approve-everything. It gets audited monthly. Two minutes, and the only reason it happens is that it is written down.

For a while, background agents couldn’t ask for permission, so a background agent trying to write a file failed silently. That cost us real time before we understood it. Until Claude Code fixed the problem, anything that wrote had to run in the foreground.

Staleness is structural. Every member’s knowledge is frozen at its model’s training date, so by default the team is confidently out of date about a field that moves weekly. Checking live sources before answering is a written rule rather than a habit, because habits are not something this team has.

WHAT ACTUALLY HOLDS IT TOGETHER

It isn’t the agents. Anyone can rent the same ones, and they’ll be better next quarter no matter what we do.

What holds it together is duller than that. Written definitions of who does what, so the routing question has an answer. Briefs good enough that a specialist who cannot see the conversation can still do the work, because that is the actual ceiling on quality. And a review nobody is permitted to route around.

And the human, who has not been automated out of anything that matters. One person still decides what is worth doing, what is true, what ships, and what the firm will not take on. Producing a draft now costs close to nothing, which means the scarce thing is knowing which draft was worth producing and whether the one in front of you is right.

Doing the work got cheap. Knowing what’s worth doing didn’t.

FIG. 06 — CONTACT

Trust is the architecture.

If you’re thinking about putting AI agents to work, we’d like to hear what you’re working on. Send us a note.

hello@ara-data.com