AGI is here and will enable the first $1T 100 person company
For us, AGI is already here. Depending on your role and expectations, the threshold might have been reasoning with o1-preview in September 2024 or Opus 4.6 in February 2026, which brought a million-token context window in beta and stronger support for long-running agent work. Its benefits are still unevenly distributed, including between people sitting on the same team. At Loop, we decided to reorganize how we work around that fact.
I built a sandbox coding setup over a weekend in early 2026. It was our first taste of agents working away from a laptop, and the productivity boost was significant. By summer, maintaining our own setup stopped being worth it. Today we run three cloud agents, each hired for a different role: Devin, Capy and Vorflux.
Most of the industry's mindshare and usage today is in tools like Claude Code, Cursor and Codex, and they are built to make one engineer more productive. Collective organizational productivity is a different problem. A company of 100 that wants to do the work of 1,000 has to move context between people, keep knowledge out of individual heads, start work without waiting for a person, and learn from the first time something is done so the second time is cheaper. That needs a different class of solutions.
What matters is what the model knows about your company, how quickly it can be taught, how much context it can hold, and what it costs to run at volume. We wanted to build for abundant intelligence over spend cuts. The path of least resistance, unless deliberately thought through, is a company that pays for AI in every seat, speeds up one step of the work, and moves the bottleneck to the other steps.
The $1B 1 person company totally misses the point.
AGI can be blocked too (poor thing)
Agents made the work itself fast. The work still did not get done much faster, as the bottlenecks moved elsewhere.
A capable agent runs into many of the same problems as a capable new colleague. It needs shared context, has to keep up when priorities change, and must coordinate across functions without losing track of an expanding checklist.
- Handoffs between people. Product hands a spec to engineering. Three people split one large feature. A design review happens after the code is written. Scheduling calls became a bottleneck of its own. People needed time together to explain the context, agree on the approach, and review what had been built. People working in the same Slack thread with a shared agent was far more efficient. The PM adds the product requirements and each engineer adds the checks for their own part.
What it takes to move a feature forwardThe same feature, with and without an agent in the thread.BeforeWrite the brieffind a timemeet to align→Buildrepeat contextschedule review→Review the finished codefind a timereview call→ShipAfterA shared threadRequirementsEngineering checksDesign feedbackThe agent works between replies→Review & shipFour of the steps were coordination: finding a time, meeting, repeating the context, scheduling the review. In the thread the context stays with the work, and the agent works between replies.
- Waiting on a specialist. How to read our Salesforce setup and which fields matter. Which promotions were ours and which ones a franchisee ran on their own. How to set up a service, and which database has the data. Not everyone is trained on everything, and the agent stopped where anyone outside that specialty would. Once these are written down as notes, the agent reads them first and the task no longer waits on a specialist.
- Waiting on a human trigger. An error fires, and a root cause write-up and a fix are drafted within a couple of hours. A customer call ends, and the action items are posted. A customer describes a feature on a call, and the transcript feeds a product spec and then a code change. A deal is closed-lost, and the write-up is done. On the other side of the spectrum is work that only makes sense at an agent's cost, such as a recorded demo of each customer's own portal with a narrated voiceover. Remove the human trigger and far more gets done, and the company gets much closer to running itself.
- A laptop's limits. It can do far more powerful things inside a sandbox than a laptop allows, running dozens of tasks in parallel and holding more than a person can follow. Every feature or thread gets its own machine with a reproducible setup, so worktrees and databases stop getting messed up, everything is isolated, and anyone can see the work while it is happening.
- A correction someone had already made. A fix taught in one session died with it, so the same mistake was fixed many times by many people. With a central memory, one person's correction becomes a note every later session reads.
It won't work if it's not top down
We realized that the people hoarding the most context were us. Leaders knew why a KPI mattered, what we'd told the board, and what had changed in the weekly business review. Much of that never reached the people or agents expected to act on it.
AGI-pilling had to start with us writing that context down and sharing it aggressively, from KPIs and board decks to weekly business reviews and the reasoning behind our priorities. Bottom-up experimentation without that direction can consume everyone's time on work the business doesn't need. If we want autonomy, we have to give people and agents enough context to choose the right work.
Web agents in public Slack channels
OpenClaw made event-driven agents click for us. Work could start when something happened and continue without someone prompting every step. We began to see private, single-player agent sessions as a temporary stage in how the tools were built.
In July we mandated web agents for the whole company and stopped using local agent sessions for company work. Requests went into public Slack channels, where the work and its corrections were recorded and visible to the company.
We applied the same rule to ourselves: work conversations had to leave a shared record. Slack DMs were reserved for situations that needed privacy; conversations outside Slack had to be recorded or written down. Sensitive information still needed appropriate access restrictions.
People learned the tools by reading each other's threads. Anyone can see how a colleague asked for something and copy it. When someone prompts badly, people who have done it before see it in the thread and correct it there. Everyone involved in a piece of work can work in the same thread, a salesperson, a support teammate and an engineer together, with the agent doing the work between their messages.



Agent runs per week, June to SeptemberTwo figures from the post, about 500 a week before the July mandate and over 4,000 a week a month later, with an illustrative curve between them.01,0002,0003,0004,0005,0006,000JunJulAugSep≈500 a week4,000+ a weekone month after the mandatestill growingJuly: web agents mandatedevery request goes to a public channelWeekly agent runs across Devin, Vorflux and Capy. A large share are automations with no human trigger. Solid line and both labelled points are as stated in the post; the dotted tail is the direction, not a measurement.
The company brain
By the end of the first month of the mandate the company had written over 200 knowledge notes, over 100 playbooks, over 100 automations and over 300 skills for its agents. We call the collection the company brain.
The brain has six components (credit to Devin for the inspiration).
- Knowledge notes are facts the agent reads before it acts.
- Playbooks are pieces of work a person asks for by name.
- Skills are the written procedures behind them.
- Blueprints describe how a repository and its services start on a fresh computer. The agent keeps them current as the repository changes.
- Automations are work that runs on a schedule or an event.
- Secrets are the personal and company credentials the agent needs to act, scoped to the person and the service.


Knowledge notes
Some examples are:
- Salesforce is the truth for which modules a customer has bought, so never pitch one they do not have.
- Contract terms are in the signed contract, and extensions agreed without an amendment are recorded in a Slack channel the agent reads.
- Checking whether a customer's data is late takes three commands. An engineer used to spend up to an hour on it.

Playbooks and skills
Some examples are:
- A landing page for a prospect, built and published from one Slack message.
- A value deck before a renewal, with everything Loop has recovered or improved for that customer, down to the franchisee.
- A one-page brief before every demo, and a one-pager a week before every quarterly review with the customer.


Automations
Some examples are:
- Every draft invoice is reconciled against our own location records before it is sent.
- Champions and signers who have left a customer are found every week.
- Customer praise in call recordings is clipped and posted every night.
- The recruiting pipeline is posted daily, with anything that looks wrong flagged.
- Every support ticket gets a drafted first reply.
None of these change a customer record without a person approving.





Loop after web agents
- Hiring needs changed. Requests for additional headcount dropped as agents took on more of the work. For every potential increase in workload, our immediate instinct became to ask how the agents could take it before we asked for more people.
- Who we hire. The hires we still make are for taste and end to end ownership of a problem.
- What agents do. They operate our internal admin tools, watch customer health, watch the on-call alerts, build data integrations and build go to market tooling.
- Roles expanded. Engineers took on more product thinking. Product teams shipped basic code and prototypes, taking ideas further into implementation. Customer success and operations ran their own root cause analyses, connected new data sources and built dashboards, with the agent doing the engineering.
- Usage. Agent runs went from about 500 a week before the mandate to over 4,000 a week a month later, and are still growing exponentially. A large share of those runs are automations.
One month after the mandate
| Before | After | |
| Agent runs per week | about 500 | over 4,000 |
| Knowledge notes | in people's heads | 200+ written |
| Playbooks | asked for ad hoc | 100+ named |
| Automations | a handful of scheduled jobs | 100+ on schedules and events |
| Skills | tribal | 300+ written procedures |
Figures as stated in the post. Counts are as of the end of the first month.
We quickly lost interest in the race to a one-person company worth $1B. No one person can be the directly responsible individual for everything: someone has to judge the work, correct the agent and decide what it should do next. Companies, like the economy, depend on layers of specialized work that other people can reliably build on. Agents let a small team do much more while its members retain responsibility for different parts of the business.
Our bet is that small teams will build companies worth $10–100T within our lifetimes. Agents let each person take on far more work, while teammates bring the judgment, expertise and accountability that one person cannot supply across an entire company.
How we think about cost
We look at agent spend the way we look at people spend: there is no bar for the right fit, and each agent is hired for a role. We run three because we need all three.
Where each agent sitsHow structured the work is against how deep it goes. Circle size is what a run costs.STRUCTUREDUNSTRUCTUREDQUICKDEEPCapythe junior engineer · subscriptionVorfluxthe senior engineer · subscriptionDevinthe chief of staff · per unit of computecircle size = cost per runShared by all threeconnectorsknowledge notesskillsplaybookssecretsPositions are qualitative. The cost difference is the billing model: Capy and Vorflux are flat subscriptions, Devin is billed per unit of compute. Together the two subscription agents cost roughly a tenth of Devin.
Devin is the chief of staff. It is built for everyone, gets unstructured work done, and is the most mature all-rounder: computer use, end-to-end testing, exploratory and long-running work, and creating new automations. It is also the one billed per unit of compute, so it has limits and people are warned as they get close to them.
Capy is the junior engineer, built for SDRs, account executives, customer success and operations. Trained well, it does routine work at a very efficient price: running automations, root cause write-ups, lookups, value decks, anything with a skill behind it. It is snappy and has no limits. It still has rough edges on knowledge and on browser use.
Vorflux is the senior engineer, built primarily for engineering. It thinks and designs a lot, writes high-quality code with careful corner-case handling, and reviews pull requests at very little cost. It is slow but meticulous, sometimes working for days to produce the highest-quality output.
All three share the same connectors to our codebase, Salesforce and Drive, the same knowledge notes, skills and playbooks, and the same secrets, so most tasks can move between them. Together the two more economical agents cost roughly a tenth of what Devin does, and Devin costs roughly a tenth of our Series A. Ramp's engineering team made a similar argument about model routing. “The discipline on the cheap side exists to free up unlimited ambition on the expensive side.”
Why this is not more common yet
The tooling to put this intelligence to work across a company has been good enough for us only since around February 2026. Devin was far ahead of the rest for most of that time, and far more expensive. Beyond timing, three things stand in the way of adoption.
- Leaders have to fund the transition. Claude Code makes one engineer faster on day one. A web agent makes the same engineer slower on day one, because the setup is not done and nothing has been taught yet. Left to individual choice, the first user has a worse week than they would with the tool they already have, and they go back. Leadership has to fund the setup and accept that initial productivity dip.
- Setup comes before any benefit. Permissions, codebases, services, secrets and test data all have to be set up for the agent's computer before it can do anything useful. For a new codebase this is small. For an existing one with years of services and secrets it is a project, and the team pays for it before it sees any return.
- It costs more than the laptop tools it replaces. A Claude Code or Cursor seat is one subscription. A web agent is billed per run and per user on top of the seat, and engineers at Loop are on several of them and use them daily. The spend only makes sense as a replacement for, or a complement to, headcount. Blindly measured against last year's AI budget, it is excessive.
Looking forward
Everything above exists so the learning loops can start. Every session and correction is recorded in one place, and from that record:
- Repeated prompts become playbooks and skills.
- The agent learns how the team thinks, such as how a feature request gets scoped and which ones get picked.
- Waits that many agents hit get fixed once, in the setup, by an agent.
- When a person corrects a fix, the agent learns why its reasoning differed from the person's, and that class of disagreement stops recurring.
The next step is a second agent that reads those records and proposes updates to the notes, skills and memory. People still decide what gets built and review how the system learns from their corrections.
We can give an agent a project that runs across days, exploring alternatives and checking results long after any one of us has stopped working. That changes what we're willing to attempt. We want to turn ideas we'd previously dismissed as too ambitious into working prototypes quickly enough for the team to see the vision and build on it.
That speed has to come with deliberate verification. A prototype can move fast; changes to customer data and production systems need clear permissions, review and a way to undo them. The ambition is to take on much larger problems while keeping those checks intact.

