T-01 · GROUNDED
Is it working from data you can trust? When AI gets something wrong, the real cause is usually bad data underneath it.
ARA — ANALYTICS · RESEARCH · ARCHITECTURE
Getting AI to do impressive things isn’t hard anymore. Getting it to do the right thing every time, in front of your customers, is. That’s where we come in: data and AI architecture for organizations that want to move fast without cutting corners on trust.
A diagram of the four-layer agent stack: L1 Data, L2 Models, L3 Tools, L4 Orchestration. An illustrative request enters at L4, pauses at each layer, then drops to the ground plane below L1, logged at 38 milliseconds. Rotate it by dragging, with the arrow keys, or with the turn buttons.
BIG-COMPANY AI EXPERIENCE, FOR SMALL BUSINESSES AND NONPROFITS.
FIG. 01 — THE THESIS
YOU CAN’T BOLT TRUST ON AT THE END.
Anyone can wire up an AI agent in an afternoon. Building one you’d trust with your customers, your money or your data takes real design work up front. That’s the part we do.
Is it working from data you can trust? When AI gets something wrong, the real cause is usually bad data underneath it.
Does it know what it’s allowed to do, and does it stop there? Clear limits, the right permissions, and a way to hand things off to a person are what keep a helpful assistant from turning into a liability.
Can a person see why it did what it did, and stand behind it? That record should be there from day one, not pieced together after a lawyer asks for it.
Grounded. Bounded. Accountable.
If it doesn’t pass all three, we don’t ship it.
FIG. 02 — TRACK RECORD
30 years
Building enterprise data, analytics and AI systems, first hands-on and later as an advisor to executive teams. ARA was built on that experience.
We go back to when data warehouses were built by hand, through the analytics and machine-learning years, and into today’s AI agents. Our founder spent three decades advising and building for some of the world’s largest media, telecom and tech companies, where a bad call can cost millions and you don’t get many second chances.
C-suite advisory at the world’s largest media & entertainment enterprises.
Analytics and AI for major telecommunications carriers.
Data architecture for global high-tech platforms with audiences in the hundreds of millions.
That’s where we learned the craft. Now we bring the same discipline to small and mid-sized businesses and nonprofits, who usually can’t get this kind of help without a big-company budget.
FIG. 03 — WHAT ARA DOES
We don’t resell anyone’s platform, so we’re free to recommend whatever fits: Databricks, Snowflake, BigQuery, any of the major clouds, open source, or the leading AI platforms.
Dashboards, reporting and analysis built on the tools you already have, or on whatever fits best if you’re starting fresh.
We design the data and AI systems underneath your agents so they’re reliable and safe from the start instead of patched later.
Straight advice on what’s out there: what actually works, what’s still hype, and what’s worth trying now versus waiting on.
FIG. 04 — SUCCESS STORIES
6 photo tiles — the hard-coded contact limit
An iPhone app for an aging parent: six photo tiles, one help button, and a caregiver who manages it all from their own phone. Most of the design work went into deciding what to leave out.
See how it’s builtFIG. 04.1 — CAREDIAL
For a lot of older people, a smartphone causes more confusion than it solves. The contacts app scrolls forever, a mis-tap opens something unrecoverable, and the person who could fix it lives in another city. So the phone stops being used for the one thing that matters most: calling family. CareDial started as a fix for our own family. Elderly parents needed a simple way to call, text and FaceTime a few key people, and we needed a way to manage those contacts remotely. The fix turned out to be less software. We built it for us, then polished it so other families could use it too.
CareDial is an iOS app with two faces. On the parent’s phone it is one screen: up to six large photo tiles under the question “Who do you want to talk to?” Tap a face, the call starts. A caregiver can enable FaceTime or texting per household, and a red Help button calls a designated person, either with a confirm step or by dialing straight through. Settings exist but are deliberately hard to find: a low-contrast gear icon that opens only after a triple tap, so the parent cannot wander into configuration by accident.
The other face is the caregiver’s. On their own phone they add, edit and reorder contacts, set the help contact, and switch features on. Changes sync to the parent’s device through Apple’s CloudKit, over a shared zone the two phones pair with a short code. The parent’s side is read-only, and the app keeps a local copy of the contacts so it still works when sync doesn’t. If the parent does reach the settings sheet, they see one calm message: the contacts on this phone are managed by their caregiver.
A complete, working SwiftUI app: role selection and pairing, the parent’s call screen, the caregiver’s management dashboard, CloudKit sync with local caching, and the help-button safety path. It’s native Swift throughout with no third-party dependencies, so there’s nothing to maintain but the app itself.
CareDial shipped. It is live on the App Store and it is free. It launched recently, so the only users so far are the family it was built for, which is the right place to start. It’s also a good example of how we approach client work. Most apps keep piling on features. This one works because it doesn’t: the contact limit is hard-coded at six, on purpose.
“Intentionally low-contrast so elderly users overlook it. Caregivers triple-tap to open Settings.”
CAPABILITIES
iOS / SwiftUI development · CloudKit sync & device pairing · accessibility-first product design · role-based UX · safety-path design
7 bidders scored — three set aside with documented reasons
An owner’s rep firm was comparing contractor bids by hand. We built a scorecard they can rerun on every project. It does the math, flags anything that needs a human call, and explains its results in plain English for a volunteer condo board.
Read the leveling storyFIG. 04.2 — BID LEVELING
The client is an owner’s representative firm that manages restoration and renovation projects for condominium and HOA boards in Florida. When a project goes out to bid, contractor proposals come back long, inconsistent, and organized however each contractor pleases. The firm leveled them by hand: copying line items into Excel, reformatting, rebuilding formulas. Then it had to explain the result to volunteer board members with no construction background, who carry fiduciary responsibility for multi-million-dollar awards under Florida’s post-Surfside condo statutes. It was slow, it was easy to make mistakes, and it was hard to defend if anyone asked where a number came from.
We didn’t build an AI that reads the bids and picks a winner. We split the work the way a careful estimator would. Everything mechanically present in the bid matrix is computed: totals, price per square foot, price tiers, variance, ranking. Everything that is judgment stays labeled as judgment. The cost baseline and the square-footage basis come from the firm’s estimator each run, and the qualitative ratings are drafted by AI, then reviewed and overridden by a human before any board sees them. Line items map to a CSI MasterFormat taxonomy extended for restoration trades, with leveling rules that keep base bids, alternates, allowances and exclusions in separate buckets. If a required input is missing, the tool stops and asks. It never guesses.
A repeatable bid scorecard the firm runs on each new project, installed on their own machines as a Cowork plugin. They upload the bid files and press play; leveling that took hours of copying and reformatting now takes minutes. The scorecard ranks the bidders and drafts an award recommendation, a capability the firm didn’t have before this project. The handoff package includes an operating guide, a sample output, and a plain-English assumptions and methodology memo written for board members, which separates hard arithmetic from informed judgment section by section. We also delivered a morning business-pulse digest for the firm’s leadership.
It proved itself on the very first project we tested it on. The tool’s independence check found that the supposedly independent cost baseline matched one bidder’s subtotal almost exactly, a strong sign the yardstick had been anchored to that bidder’s own number. It surfaced a duplicate bid from the same contractor with two different totals. Seven bidders were scored and three were set aside with documented reasons. The board received a memo that states plainly which numbers are calculated, which are estimates, and which still need confirmation before an award. The time savings got their attention, but the error-checking is what sold them: the leveling rules catch miscategorized line items, flag entries for review, and find errors with suggested corrections. They are rolling the tool into their standard bid process now.
“The computer does the arithmetic; the human supplies the professional estimates.”
CAPABILITIES
document intelligence · bid leveling & normalization · CSI MasterFormat taxonomy · Python data pipeline · AI-drafted, human-reviewed scoring · board-ready reporting · Cowork plugin packaging
AI agents in Copilot and Claude that take busywork off a small team
A four-person team at a global nonprofit knew exactly what they wanted to automate but had no time to do it. We found where they were stuck, trained them hands-on, and helped them build AI agents to take on the work they couldn’t get to.
Read the enablement storyFIG. 04.3 — AGENTS AND AUTOMATION
The client is a nonprofit digital-outreach organization whose Asia-Pacific team is four people running Meta and Google campaigns to audiences across the region, in multiple languages. The front half of the loop took all their time: content, ad processing, campaign management. The back half kept going cold. Following up with people who respond, reconciling campaign data against donor records and reporting all sat at the end of a queue that never emptied. The team had already drawn its target workflow on a whiteboard.
Before talking tools, we figured out the real problem. They didn’t need anyone to explain what AI could do. They already knew. They just didn’t have the hours. We produced an AI opportunity map that named the highest-value builds in priority order, audited the team’s actual Microsoft licenses so every recommendation matched what they could use, and designed a three-day hands-on workshop around a simple stack: Planner as the task backbone, Power Automate as the plumbing, a Copilot agent grounded on the team’s own SharePoint content, and a Power BI builder day. A license contingency kept the workshop standing even if the premium Copilot licenses fell through. From the first email exchanges to onsite delivery took about a month.
The opportunity map, with three named agents in build order. The full workshop design and facilitator materials. The license and capability audit.
The workshop ran onsite over three days, and each night we reworked the material based on what came up that day. The team came away building. They stood up a communications agent in Copilot, grounded on a SharePoint knowledge library of their own communication style, tone and SOPs. That agent was on the opportunity map, and they built it. Then they went past the roadmap entirely and built a subagent team in Claude Code that handles translation, updating large numbers of files across many languages. We can’t share much about their work, but we can share the time: an operation that took days by hand now takes minutes. Their biggest challenge is still finding people to support the work, but the agents are already saving them real time. The engagement was pro bono.
“The map, but not the hands to walk it.”
CAPABILITIES
AI opportunity mapping · M365 Copilot & Agent Builder · Power Automate & Planner · Power BI · Claude Code subagent teams · workshop & curriculum design · license/capability audit
FIG. 05 — INSIGHTS
We taught AI to compare 100-page contractor bids. Reading them was the easy part. Writing down the rules was the real work.
Read the noteA model can read a hundred-page contractor bid in seconds and give you a clean summary of it. That stopped being interesting a while ago. What it cannot do on its own is tell you whether two of those bids are describing the same building.
That’s the part that takes real work.
We built a bid-leveling tool for an owner’s representative firm in Florida that manages restoration work for condominium and HOA boards. Seven proposals arrive for a project. Each contractor organizes its numbers the way its own estimating shop prefers. The person who eventually has to choose is a volunteer board member with no construction background and a fiduciary duty for a multi-million-dollar award under Florida’s post-Surfside statutes. The job was to give that member a ranking they could defend.
The bids don’t arrive comparable. You have to make them that way.
Take one scope of work. One bidder prices it as a line in the base bid. Another moves it to an alternate, so it only appears if the board elects it. A third carries it as an allowance, which is a placeholder with a number attached rather than a price. A fourth writes it into the exclusions and never mentions it again. Four honest bids, four different totals, and none of the differences are about who is cheaper.
So the first real decision is the taxonomy. We mapped line items to CSI MasterFormat, extended for restoration trades. The standard covers the industry; the extension is where the actual work went, because the mapping is a judgment about what a given line item is, made once and then applied consistently.
The second decision is to keep base bids, alternates, allowances and exclusions in separate buckets and never let them merge. Collapse them and the word “lowest” stops describing price. It starts describing which bidder chose to include the most.
Neither of those is an AI problem. They’re construction know-how, and someone has to write it down before you can automate anything.
We drew that line on purpose and then made the tool enforce it.
Anything mechanically present in the bid matrix is computed: totals, price per square foot, price tiers, variance, ranking. Arithmetic does not need an opinion, and asking a model for one is how you end up with a number nobody can trace.
Anything that is judgment stays labeled as judgment. The cost baseline and the square-footage basis are supplied by the firm’s estimator on every run. They are not inferred, and they are not remembered from last time. The qualitative ratings are drafted by AI and then reviewed, and overridden where needed, by a person before a board sees any of it.
And when a required input is missing, the tool stops and asks. It never guesses.
That sounds like a small detail, but it’s what makes a professional willing to sign off on the output instead of double-checking all of it.
On the validation project, the independence check found that the supposedly independent cost baseline matched one bidder’s subtotal almost exactly. That is a strong signal the yardstick had been anchored to the number it was supposed to be measuring. The tool also surfaced a duplicate bid from the same contractor carrying two different totals. Of the seven bidders scored, three were set aside with documented reasons rather than quietly vanishing from the comparison.
So add validation, but point some of it at your own inputs as well as the bidders’. Those are the mistakes nobody else in the process is looking for.
The reader of this work is a volunteer, not an analyst, and they carry real exposure for the decision. So the handoff package includes a plain-English assumptions and methodology memo that separates hard arithmetic from informed judgment, section by section, and states which numbers still need confirmation before an award.
We could only write that memo because every number in the pipeline traces back to the line item it came from and the rule that moved it. Try to add that kind of traceability at the end and you’ll be describing a process you can’t actually reconstruct.
Having AI read documents is cheap and easy now. Comparing them fairly isn’t, and a better model won’t fix that. It comes down to rules, and those rules live in the heads of experienced people who’ve usually never been asked to write them down.
If you are pointing a model at a stack of documents and expecting a ranked answer, the question to settle first is what would make any two of those documents comparable, and who decides. Answer it properly and most of the build is mechanical. Skip it and you get a confident ranking of things that were never the same.
The firm noticed the time savings first: leveling that took hours of copying and reformatting now takes minutes. What won them over was the checking, which is slower to appreciate and harder to demo. Being fast got us in the door. Catching a mistake before it reached the board is why they made the tool part of their standard process.
A small nonprofit team didn’t need another automation plan. They needed hands-on training first, then a set of AI agents in Copilot and Claude to take work off their plate.
Read the notePicture four people running ad campaigns across a whole region, in several languages, with the workflow they wanted already drawn on a whiteboard.
They had the plan. What they didn’t have was time.
The client is a nonprofit digital-outreach organization whose Asia-Pacific team was carrying all of that. The front half of their loop, the content and the campaign work, consumed everything they had. The back half kept going cold: following up with the people who responded, reconciling campaign data against donor records, reporting on any of it. Their lead put it this way: they had the map but not the hands to walk it.
From the outside, a team that doesn’t know what’s possible looks a lot like a team that knows exactly what it wants but has no time. Both are stuck, but they need very different help.
The default consulting move is to hand over an automation plan. For this team, that would have been a well-made answer to the wrong problem. They did not need to be told what to build. Telling them again would have added a document to the pile they already had no hands for.
So we started there. The problem was time, not know-how, and everything else followed from that.
We audited the licenses the team already held, so every recommendation matched software they could use. We also built a license contingency into the workshop design, so the three days still worked if the premium licenses did not land in time.
This isn’t really about Microsoft. If your plan depends on someone else’s purchasing decision, it has a hole in it. Build around what people have today and treat any upgrade as a bonus.
The opportunity map we produced named the highest-value builds in priority order.
A list of a dozen ideas gets read once and filed away. A list that says “build this first, then this” actually gets used, because a team with no spare time can’t afford to spend it deciding.
Three days, onsite, hands on keyboards. The stack was deliberately plain: Planner as the task backbone, Power Automate as the plumbing, a Copilot agent grounded on the team’s own SharePoint content, and a Power BI builder day.
We kept it simple on purpose. Every new tool is something those same four people have to maintain, and a clever setup that eats an afternoon a week isn’t worth it.
We reworked the material each evening around whatever had come up that day. If your workshop plan doesn’t change once you’re in the room, you probably weren’t listening.
The team came away building. They stood up a communications agent in Copilot, grounded on a SharePoint knowledge library of their own communication style, tone and standard operating procedures. The opportunity map had named that agent; they built it.
Building it on their own material is what makes it more than a generic chatbot. It’s also what won over the skeptics on the first day.
Then they went past the roadmap entirely. They built a subagent team in Claude Code to handle translation, updating large numbers of files across many languages. An operation that took days by hand now takes minutes.
None of that was in the plan, and we weren’t there when they built it. That’s how you know the training worked: the team crossed onto a second platform on their own initiative and got something running.
And they didn’t build just one agent. They built a team of them with the work split up, which is harder to design and does a lot more for four people.
When a team is short on time, handing them a document doesn’t help much, because someone still has to do what it says. Helping them build the thing does.
Using two platforms wasn’t a principle we started with. It’s just where the work led. Copilot made sense because the work already lived in M365. Claude Code made sense because the file-scale language work needed something else. Committing to one vendor’s complete answer would have cost this team the second half of what they now have.
One limit worth being clear about: the team’s biggest challenge isn’t software. It’s finding people to support the work, and no agent fixes that. What the agents do is give the team hours back, and those hours can go toward finding those people. The engagement was pro bono.
ARA is one person and a team of AI specialists, each with a name and a job description. Here’s how that works day to day, including the parts that don’t.
Read the noteARA has one human. The rest of the firm is a roster of AI specialists, each with a written job description, a defined scope, and a specific set of tools it is permitted to touch.
This is what it’s like to run, including the parts that don’t work.
Every member of the team is a file. The file holds a name, a persona, what the role is for, what it must never do, and which tools it can use. There is a chief of staff, a head of people, a researcher, a designer, a strategist, a delivery and business-operations lead, a data scientist, a solution architect who builds, a technical reviewer, a web developer, an education lead, and a handful of domain specialists added for particular kinds of work.
Adding a new role isn’t like hiring. It’s editing a folder, and because the folder is under version control you can read back why a role exists and what it was originally scoped to do.
The part that turned out to matter more than expected is that roles get researched before they get written. The researcher produces an evidence-based brief on what a strong human practitioner in that domain actually knows, uses and produces. Only then does the head of people turn that brief into a role definition. Skip that step and you write a job description out of your own assumptions about a field you do not practice, which is exactly the failure a one-person firm is most exposed to.
The orchestrator does not do the work.
The chief of staff routes, briefs, sequences hand-offs, and brings results back, and does not write the deliverable, run the analysis, or make the domain call. We made that a standing rule, because the temptation to just do it is constant and it is strongest exactly when the work is urgent.
This isn’t about being a purist. The moment the orchestrator starts producing, you lose the record of who did what, and you lose the specialist framing that made the specialist worth having.
Two constraints shape everything else. Specialists cannot call each other. And they cannot see the conversation that produced their assignment.
So every hand-off passes through the orchestrator, who physically carries the artifact from one specialist to the next. We didn’t pick hub-and-spoke. The setup forced it on us.
So everything depends on the brief. If a file path is not named in the brief, the specialist does not have it. If a hard rule is not carried into the brief, it does not apply. The most common failure mode in this firm is a thin brief, not a weak specialist, and it took a while to stop misdiagnosing the first as the second.
Our fix is a written rule: if a brief is ambiguous, come back and ask instead of guessing. It works, imperfectly. A capable model handed an under-specified task will produce something plausible, and plausible-but-wrong is the most expensive kind of mistake.
One member reviews every deliverable before it goes anywhere. The verdict is approve, approve with conditions, or reject with remediation. Remediation travels back to the original author through the orchestrator, and the work re-enters the gate.
When the reviewer is also the builder, someone else cross-checks the work against the same criteria. Reviewing your own work doesn’t count.
The gate covers more than code. Anything ARA-branded is checked against the brand canon. Anything outward-facing is checked for whether it reads like a person wrote it, because writing that sounds machine-generated isn’t finished.
The downside: the review is slow, and it’s the step we’re most tempted to skip. Its value shows up in the work it sends back, which is hard to appreciate in the moment.
There are four execution postures, and every member sits in exactly one.
Builders may write and run code, but only inside an isolated worktree, under a sandbox, with a deny list for destructive commands that always overrides whatever permission has accumulated. Authors write text and execute nothing. The reviewer runs validation and writes findings, never the artifact under review. The control-plane engineer writes and runs code in the main tree itself, within a defined scope.
Builders also have to verify their own work before handing it over: turn the task into criteria you can check, then loop until the criteria pass. That does not replace the gate. It just means the gate is not the first place a problem gets found.
Seven rules bind everyone who builds. These are the first four. Think first and state your assumptions; if two readings of the brief exist, present both rather than silently picking one. Ship the minimum that solves the problem, and if two hundred lines could be fifty, rewrite it. Change only what the task requires, so every changed line traces back to the request. Turn the task into verifiable success criteria and work until they are met.
Each of those four exists to stop the same behavior. A capable model will cheerfully produce more than you asked for, and most of what it adds is work someone now has to read, review and maintain.
This is the part that surprised us most about running the thing.
Hours are not the constraint. What every agent must carry into every session is. So new knowledge starts cheap, as a reference document read on demand. It gets promoted to a skill only when it is a repeatable procedure with a clear trigger. It goes into a specific agent’s definition only when that agent must always know it. It reaches the firm-wide charter only when every agent must. Each rung up is a permanent tax on more sessions.
Demotion is allowed and used. A rule that stops firing goes back down the ladder. Left alone, the rulebook grows until nobody reads it, just like every employee handbook ever written.
Every gated deliverable produces a short capture file: what we learned, what is reusable and in what form, what idea came out of it worth developing, and what we would scope differently next time. The gate checks that the file exists.
“Nothing reusable came out of this” is a legitimate answer, and the check verifies that the question was answered rather than that an asset was produced. That detail matters more than the mechanism. A capture requirement that demands a reusable asset every time generates fake ones, and a library of fake assets is worse than no library.
We do this because everyone rents the same AI models. What belongs to ARA is the written judgment: the definitions, the rules, the captured patterns, the things we now know not to do again.
Debris accumulates. Builder isolation creates working copies, and they pile up. At one point there were dozens of stale ones on disk, enough to start breaking shell commands outright. They now get pruned on a schedule, because nothing about that problem announces itself until it does.
Permissions drift. Every approval click appends to a local allow list. Left alone, that list quietly becomes approve-everything. It gets audited monthly. Two minutes, and the only reason it happens is that it is written down.
For a while, background agents couldn’t ask for permission, so a background agent trying to write a file failed silently. That cost us real time before we understood it. Until Claude Code fixed the problem, anything that wrote had to run in the foreground.
Staleness is structural. Every member’s knowledge is frozen at its model’s training date, so by default the team is confidently out of date about a field that moves weekly. Checking live sources before answering is a written rule rather than a habit, because habits are not something this team has.
It isn’t the agents. Anyone can rent the same ones, and they’ll be better next quarter no matter what we do.
What holds it together is duller than that. Written definitions of who does what, so the routing question has an answer. Briefs good enough that a specialist who cannot see the conversation can still do the work, because that is the actual ceiling on quality. And a review nobody is permitted to route around.
And the human, who has not been automated out of anything that matters. One person still decides what is worth doing, what is true, what ships, and what the firm will not take on. Producing a draft now costs close to nothing, which means the scarce thing is knowing which draft was worth producing and whether the one in front of you is right.
Doing the work got cheap. Knowing what’s worth doing didn’t.
FIG. 06 — CONTACT
If you’re thinking about putting AI agents to work, we’d like to hear what you’re working on. Send us a note.