Software tracked the work. Now AI does it.
NowAI does
the work
A ShopOS essay · Notes on work
Software sold
us a cabinet
Software sold a cabinet.
It tracked the workNow it has hands.
For decades, software has helped companies track, route and store their work. A claims team could see every claim and a finance team could see every invoice, but a person still read the file, applied the rules and decided what happened next. The software was the cabinet. The clerk did the work.
That boundary is moving. A model can read an email, pull the facts out of a PDF and draft the reply, which lets a system finish work that used to wait for a person. The opportunity is a company that sells the finished file. AI does most of the labor, and people stay close to the parts where judgment and accountability still matter.
The money
is in the work
A $100M insurance agency spends $1M to $2M on software.
The real cost sits in the work.
Most of the discussion about AI applications stays inside software budgets, which are a relatively small pool of spend. The larger pool sits in the work itself: processing claims, reconciling payments, chasing renewals, preparing compliance files and answering customer requests.
Take an insurance managed general agent with $100M in revenue. It might spend $1M to $2M a year on software, which makes it hard for a software company to build a venture sized business there. Its operating costs tell a different story. Intake, quotes, renewals, customer emails and policy servicing add up to a large share of the cost base. A company that sells the finished workflow sells into that budget.
The line keeps moving
Two axes: input from tidy to messy. Decision from rules to judgment.
Old software lives bottom left. Clean inputs, fixed rules.
Models read messy files and reason. The frontier moves.
Between the lines, AI does the work. People own review.
Every new model pushes the line further.
The top right still needs people. For now.
Picture a map of business work. Left to right, the input goes from structured to messy. Bottom to top, the decision goes from rule based to judgment. Traditional software lives in the bottom left, where the inputs are clean and the next step can be written as a rule.
Language models changed what sits inside the boundary. They read unstructured documents, extract meaning from emails, classify requests, compare records and draft responses. They still struggle when a setting is wide open, but many business workflows are narrower than they look. They are repetitive, bounded and governed by stable rules, even when the inputs arrive in messy form.
The interesting zone sits in the middle. AI does most of the labor while people own review, exceptions and accountability. Each new model generation pushes the line further out, so work that needed a person last year clears itself this year.
Buyers already
hand work out
The vendor starts where work already arrives.
Nobody runs a software rollout.
Over the past decade, companies grew comfortable with distributed teams, offshore operations and outside operating partners. Many functions already run through a patchwork of internal operators, vendors, workflow tools and manual handoffs. The habit of handing work out is already there.
The software shelf is crowded too. A mid size insurer or lender usually owns a core system, a CRM, a ticketing tool, a document store and several workflow products. The bottleneck is the burden of coordinating across all of it. These buyers want lower cost, higher accuracy and faster turnaround, and another application adds to the pile.
A service wrapper fits that reality. The vendor starts in the channels that already exist: the inbox, the portal, the PDF and the spreadsheet. Nobody runs a software implementation or redesigns an operating model first.
A new layer in the middle
Software records and routes the work.
People sit on top and do the work.
A model slots in between. It reads and drafts.
People move up to review and answer for quality.
The old stack had two layers. Software recorded and routed the work, and people did it. The new stack has three. Systems of record stay where they are. A model sits in the middle and reads, extracts, compares and drafts. People move to the top, where they review, handle exceptions and answer for quality.
The human layer should get thinner over time. In the early months it is scaffolding. It lets the company deliver production grade work while the system learns the domain, and it creates the feedback loop through which workflow knowledge turns into software.
Share of work needing a person
30%
Teams like Kim and Pibit report going from over 70% to 30%.
An AI native service starts heavy on people. Early operators learn the workflow, catch failures and handle the odd cases that each customer brings. Every fix then moves into software: prompts, review rules, integrations, evaluation sets and feedback loops.
Kim, which runs customer support for e-commerce companies, and Pibit, which runs insurance underwriting, report cutting the share of work that needs a human from over 70% to about 30%. The work delivered keeps rising while the human effort behind it falls. A traditional services company scales by adding people. This kind of company improves the ratio between work delivered and effort required.
Five lines
worth watching
Each line improves as volume grows.
That learning curve is the part worth watching. Five measures show whether it is real. Revenue per operator rises. Exception rates fall. Turnaround time comes down. Gross margin improves with volume. New customers in the same workflow onboard faster, because the system has already learned the common paths.
Is this just
BPO with AI?
A BPO moves labor somewhere cheaper.
This one moves labor into software.
The category is often described as BPO with AI. A traditional BPO takes an existing process and moves the labor to a cheaper or more scalable place. Technology helps, but labor remains the main unit of scale.
An AI native service can look service heavy at the start, because production quality still needs people. Early operators understand the workflow, catch failures and turn each customer's messiness into repeatable steps. Over time that knowledge moves into software, review systems, integrations and evaluation data.
The test is the shape of the line. If the five measures stay flat, the company is a better services firm. That can be a good business, and it is a different one from the business a venture investor is underwriting.
Which work
to take first
Later, as models improve.
Frequent, painful, bounded.
Skip it.
Fine, and small.
Start where work is frequent, painful and bounded.
Not every workflow suits this model. The best candidates are frequent, painful, bounded and valuable enough to absorb. The quality bar is clear, the steps follow stable rules, and a person can stay close to the parts that need judgment. Open ended or rare work is better left alone until the models improve or the volume justifies the build.
Work that has
no vendor yet
The BPO shelf has its price set.
BPO shelfStranded work fits no vendor.
No software fit · No dealThe obvious places to start are existing services categories: call centers, BPOs, offshore operations teams and back office vendors. The labor is visible and the processes are already outsourced. These markets also carry price benchmarks, set by delivery models that have been tuned for decades, so a new entrant's margin is thinner than it looks.
The more interesting opportunities show up before a formal vendor category exists. Spend a day inside one of these:
A surprising amount of work is repetitive, expensive and still trapped inside the company, because it never cleanly fit into software or outsourcing. The workflow has structure, but the path changes by customer, geography, document type and exception. Nobody could serve it at attractive economics before. It shows up as headcount, backlogs, error rates, slow response times and missed revenue.
One file,
start to finish
cases
Odd cases go to a person. The system keeps the lesson.
The MGA pays per file. It runs no extra team.
Take an MGA, where submissions arrive by email with attachments in every shape. An AI native service can take intake and underwriting preparation as one completed workflow. It receives the submission, extracts the relevant information, checks the file for completeness, enriches it, flags issues and prepares the underwriter's view. Exceptions route to a person, and the system keeps what it learns.
The MGA can pay per file, per policy or per completed workflow. It keeps its core system, manages no additional team and gets faster cycle times and a lower operating load. Pibit runs underwriting workflows for MGAs this way today. The same pattern shows up in healthcare administration, revenue cycle operations, accounting, compliance, logistics, legal operations and customer support.
How it
gets priced
The closer to the result, the more the vendor owns.
Pricing follows the unit of finished work. A vendor can charge per transaction, per file, per policy or per completed workflow, and the strongest companies price against outcomes. The buyer pays for work that is done, which keeps the vendor pointed at speed and accuracy and away from hours billed.
Who's well placed
to build it
The SaaS wave built products. The services wave built delivery.
The edge that lasts is the pairing.
This kind of company needs software talent and operating discipline in the same team. The SaaS wave showed that Indian teams can build product companies for global buyers. The services wave before it built deep capability in process design, offshore delivery, quality management and training. AI native services sit where those two histories meet.
The best early customers are small and mid size companies in the US and other developed markets, with real budgets, real backlogs and little appetite for a software transformation. Lower delivery cost helps from day one. The lasting advantage is a team that combines model engineering with disciplined operations and treats delivery as the product.
Here's how
we run it
Every job feeds
the next one
Each result teaches the next job.
At ShopOS the work runs in a loop. Research benchmarks a model and turns the result into a reusable workflow. The workflow runs as a deployed brand job. A client or a reviewer makes a decision, the result goes live with its cost trace, and the lesson lands in memory, evaluations, a skill or a template. The next comparable job starts from there.
The loop earns its claim only when the next comparable deployment improves on setup, acceptance, review effort, reruns or reliability. A large archive of prompts and outputs shows activity. The curve shows compounding.
A small pod
per brand
A pod of three to five runs each brand.
The nth job should cost less. A target.
ShopOS sells brand outcomes. A brand asks for a catalog, a campaign or a store page and receives the finished work. A pod of three to five people runs each brand: a creative director, an AI expert, a workflow engineer, QA and last mile, and an account lead.
Every engagement leaves something reusable behind, such as a cleaner input contract, a category template, a rejection rule or a test. We compare the first deployment with the nth one for similar work, and the nth should need less setup, fewer reruns and less review. Today that is a target. Setup time, reruns, human hours per outcome and gross margin need two quarters of tracking before it counts as proof.
Press and hold. Or keep scrolling.
What can
go wrong
One person holds it all
The client is happy because one operator remembers everything.
Lots of prompts, no tests
Text saved with no stable input, output or owner.
The same cost every time
Each new account rebuilds the same labor.
Three patterns keep a deployment business stuck as a services business. The first is manual heroics, where the client is happy because one operator remembers everything. The second is a prompt archive, where the team saves text with no stable input, output, test or owner. The third is a flat cost curve, where each new account rebuilds the same labor under a new brand name.
If one of these shows up, the business has revenue and client knowledge and no software yet. Capacity is the other limit, because trained creative people set the pace and hiring takes time. Everything here can look different in six months, so we hold the model loosely and keep the interface simple: a customer can bring in an expert for one task and let them go afterwards.
Questions every
team runs into
1
What work?
Frequent, painful, bounded.
2
Wide, or deep?
Across industries or inside one.
3
How to charge?
File, workflow or outcome.
4
Where's the line?
AI, review, exceptions.
5
When margins rise
They climb with volume.
6
Build or buy?
Or rebuild a services firm.
Teams building in this category run into the same six design choices. Which workflows are frequent, painful, bounded and valuable enough to absorb. Whether to go horizontal across industries or deep into one vertical. How to price: per transaction, per file, per workflow or per outcome. Where the line sits between AI, human review and exception handling. How gross margin should improve as volume grows. And whether to build the operating layer from scratch or buy an existing services business and rebuild it around AI.
Software
tracked it
AI does
the work
Finished
work
most of itPeople keep
it right
The first wave of AI applications helped people work faster. The larger wave takes the work off their plate: a finished file, delivered, with people close enough to keep it right.
A ShopOS
essay
Now it gets
work done
Notes
on work