Project Breakdown

Agent Agency

The AI operating system that runs our own agency work — durable agent workflows, deterministic verification gates, and human approval on every publication

Overview

Agent Agency is the system we use to automate our own agency's workflows with AI — and the clearest demonstration of how we approach AI automation for clients. It is not a chatbot and not a pile of prompts: it is a complete operating system in which AI agents execute multi-stage work inside deterministic, durable workflows, every output is checked by validators before it counts, and a human approves anything irreversible. The first specialist, a blog-content agent, runs a 14-stage pipeline from opportunity qualification through research, drafting, and evaluation to publication — and after one week of cost engineering it produces an evaluated, sourced article for around $0.37 in model spend. Every run records what each stage cost, which models served which role, and why every decision was made.

Project Details

Client
Niro Digital Internal System
Industry
AI Automation / Agency Operations
Service
AI Workflow Automation
Year
2026
Services
AI Workflow AutomationAI Agent DevelopmentCustom Software DevelopmentBusiness Process Automation
Agent Agency control plane — runs dashboard with statuses, stages, costs, and attention flags

The control plane: every agent run with its stage, status, elapsed time, and exact cost — runs that need a human are flagged for attention

Agent output is a claim. Verification and external evidence decide whether work is complete — not the model's own opinion of itself.

A completed agent run — stage timeline with per-stage costs, model usage, and pinned versions

Durable Workflows, Not Chat Sessions

High-level process is deterministic code; models get autonomy only inside bounded stages. The blog agent moves through 14 stages — qualify the opportunity, plan the contribution, research with grounded sources, build an evidence packet, outline, draft, evaluate, revise — and cannot skip evaluation or publish before approval. Every run pins the exact versions of the harness, prompts, model policy, and client configuration it executed with, so results stay reproducible, comparable, and reversible. Each stage records its own cost: the run shown here produced a 3,046-word sourced article for $0.37 across 24 model calls.

What Makes It Reliable

Verification

Output Is Checked, Never Trusted

15 deterministic validators check citations, sources, quotations, links, and prohibited claims, and a weighted evaluation scores 9 quality dimensions. A single critical finding blocks the run no matter how good the score looks.

Roles

Separation of Model Powers

Seven named model roles — researcher, writer, reviewer, editor — with a hard rule encoded in the type system: the reviewer must be a different model than the writer. A model grading its own prose is not an evaluation.

Approval

Humans Approve What Matters

Publishing is gated behind an approvals queue. Each approval binds to one exact artifact revision — approving one revision never authorizes another. The team decides; agents execute.

Cost

Cost Engineered Like a System

One day of cost engineering took an article from $8.40 to $0.37 — by killing self-improvement loops, moving search into the harness, and collapsing redundant evaluator panels. A client on three posts a week costs about $5/month in model spend.

Knowledge

Client Onboarding Is an Audit

A client's writing constitution is built from their own site and data — and building it surfaces contradictions in their existing content, which become machine-enforced prohibited claims.

Delivery

A Central Content Service

Published content lives in one multi-project content service with revisions, preview tokens, and cache-tag revalidation. Client websites fetch published content through an API — they never touch the operations database.

Approvals queue — every decision a workflow is waiting on, with risk level and exact artifact revision

Human-in-the-Loop by Design

Autonomy is granted per action, not globally. Publishing an article, merging code, or raising an ad budget each carry their own policy, and anything high-risk waits for a person. The discipline is real: the first article to reach the approval gate passed the automated evaluation — and the human editor still rejected it, rewrote it, and published their own version. That is the system working as designed: the gate's job was to decide the draft was worth a human's time, and the human's job was the final call.

Why This Matters for Your Business

Agent Agency is the engine behind our AI automation service. The same method we used on ourselves — investigate the workflow, split judgment from repetition, give AI the repetition inside guardrails, keep humans on the decisions — is what we apply to client processes. We run this system in production for our own operations first, which means every recommendation we make about AI automation is backed by scars, run records, and real cost data — not by a demo.

Want workflows like this in your company?

Show us a process that eats your team's hours. We will map it, tell you honestly which parts AI can take over, and build the automation with verification and approval built in.

Previous Project

No previous project

Scroll handle
0