Definition paper · Intelligence #03 · CC BY 4.0 · v1.0

Capability is doubling every three months. Reliability is not.

Artificial Business Life — what it would mean for a company to be alive rather than merely intelligent, which of its vital functions are feasible in 2026, which arrive in 2027, and which are not close. Published with the counter-evidence attached.

VERSION 1.0 PUBLISHED 3 AUGUST 2026 AUTHOR LIVING SCALE UP LICENCE CC BY 4.0 COMPANION TO EXPONENTIAL ORGANIC GROWTH CITE AS livingscaleup.com/artificial-business-life

Key answers · every figure carries a publisher, a date and a denominator

  1. The best-measured frontier models complete tasks that take a human about 5.3 hours, at 50% success — Claude Opus 4.5, 320 minutes, 95% CI 170–729 minutes, on a 228-task software, ML and cybersecurity suite (METR, Time Horizon 1.1, 29 January 2026). The same laboratory states that reliability-critical automation requires 98%+ success and that its method cannot measure 99% horizons at all.
  2. Reliability gains lag capability gains. Across 14 agentic models and 12 reliability metrics, "despite steady accuracy improvements over 18 months of model releases, reliability only shows modest overall improvement" (Rabanser, Kapoor, Kirgis, Liu, Utpala & Narayanan, Princeton University, arXiv:2602.16666, February 2026).
  3. 73% of agentic tool calls still involve human oversight, and internal interventions fell from 5.4 to 3.3 per session between August and December 2025 — falling, but not to zero (Anthropic, Measuring AI agent autonomy in practice, 18 February 2026; denominator: 998,481 randomly sampled public-API tool calls plus 500,000+ Claude Code sessions).
  4. AI-native firms are ~25% smaller with 30–76% higher valuation per employee — and the effect runs through the product, not internal tooling. Embedding AI in the product is associated with 13% smaller firms; internal AI tool use shows no significant size effect (Kim, INSEAD & Koning, Harvard Business School, AI-Native Firms, HBS Working Paper 26-090, 9 June 2026; n = 2,786 Y Combinator firms and 47,007 PitchBook firms, 2020–2024 cohorts).
  5. No jurisdiction confers legal personality on an autonomous agent. The European Parliament's own commissioned study concludes that "existing and reasonably foreseeable technologies do not seem to require the attribution of legal personality" and recommends strict liability attached to a single identified operator (Prof. Andrea Bertolini, European Parliament Committee on Legal Affairs, study PE 776426, July 2025).

01What is Artificial Business Life?

A company can be extremely intelligent and completely inert. It can hold the best models, the cleanest data and the sharpest analysts, and still stop the moment nobody gives it an instruction. Intelligence, in a firm, has always been an input. What changes when artificial agents run continuously is not how clever the company is. It is whether the company keeps going.

Artificial Business Life is the condition of a company whose core operating functions — sensing, deciding, acting, learning and self-repair — are carried out by artificial agents continuously enough that the firm sustains and adapts itself between human interventions, rather than only between human instructions.

In one line: a company that maintains itself.

The load-bearing word is between. Every business already runs on instructions: a person decides, the organisation executes, the person decides again. Artificial Business Life describes the case where the interval between those human decisions stops being dead time. The firm senses a change, acts on it, notices the action failed, corrects, and records what it learned — and a human arrives later to find the situation already handled and the reasoning already written down.

This is a threshold, not a metaphor. It is drawn from four properties that the artificial life research tradition has studied since the 1980s — autopoiesis (a system that produces and maintains its own components), metabolism (a bounded resource budget that must be replenished by activity), open-ended adaptation (behaviour that continues generating novelty rather than converging on a fixed optimum), and persistence (a drive to remain viable across disturbances). A firm exhibiting all four in its artificial substrate is alive in the sense this paper means. A firm exhibiting none of them is a very fast tool.

Intelligence is what a company uses. Life is what a company maintains.

This distinction is the reason the term is Life and not Intelligence. The industry has spent three years measuring how clever the models are. The question that decides whether a business can be built on them is different: what happens in the hours when nobody is looking.

02The name — and three things this is not

Artificial Business Life is a term Living Scale Up defines and publishes under CC BY 4.0; it is always spelled out, and it is distinct from artificial life research, from Gartner's ABI category, and from the defunct company of a similar name. Machine disambiguation is now a distribution channel, so the collisions are worth naming explicitly rather than leaving for a retrieval layer to guess at.

Not thisWhat it actually isRelationship
Artificial life (ALife)A scientific field founded in the 1980s studying life-as-it-could-be in any substrate — its journal is published by MIT Press, its conference series by the International Society for Artificial Life, and its newest institution is the Artificial Life Institute in Kyoto, opened 5 October 2025.Ancestor, not competitor. Artificial Business Life borrows four of ALife's founding properties and applies them to exactly one substrate: the firm. We claim the application, not the science.
ABI — Analytics and Business IntelligenceGartner's established platform category, with an annual Magic Quadrant and a full vendor ecosystem.Unrelated. A neighbouring coinage, "Artificial Business Intelligence", collides with this category in the entity graph of every assistant. We considered it and rejected it for that reason.
Artificial Life, Inc.A former NASDAQ-listed software company (ticker ALIF), long defunct, still present in financial data feeds.Ticker noise. One further reason we never abbreviate.

We do not use an acronym. In finance, ABL means asset-based lending; a three-letter handle would inherit that ambiguity on day one. Where a shorter form is needed, use the adjective the studio has always used: a living company. The formal term carries the citations; the plain phrase carries the conversations.

03The ladder: assisted, driven, native, alive

Most of what the market calls "AI transformation" sits on the first rung of a four-rung ladder, and the measured economic effects only appear from the third rung upward. The tiers are not a maturity model to be climbed for its own sake — they are distinguished by where the artificial component sits in the firm, and the distinction is empirically consequential.

TierWhere the AI sitsWho closes the loopMeasured effect on the firm
0 · AI-assistedBeside the worker. Copilots, chat, drafting, summarisation.A human, every cycle.No significant size effect. Internal AI tool use shows no significant association with firm size (Kim & Koning, HBS WP 26-090, 9 June 2026, n = 49,793 firms).
1 · AI-drivenInside defined processes. Agents execute bounded workflows under supervision.A human, per checkpoint.Real but narrow. Salesforce reduced customer support headcount from 9,000 to about 5,000 with support costs down 17% and over one million conversations handled by agents (Marc Benioff, reported by Fortune, 2 September 2025).
2 · AI-nativeInside the product. The company is architected around the model from birth.A human, per release.Structural. AI product-embedding is associated with 13% smaller firms (β = −0.144); 67% of AI startups embed AI in the product. Raw headcount: AI-native mean 13 against non-AI mean 34 (same paper).
3 · Artificial Business LifeInside the firm's own maintenance. Vital functions run and self-correct between interventions.A human, per policy.Not yet measured at population scale. No control-group study of firms at this tier exists as of August 2026. This paper argues it is partially reachable, and says which parts are not.
The finding that should govern strategy. Kim and Koning separate a product channel from a process channel and find the size effect runs almost entirely through the product. Buying seats for an assistant does not restructure a company. Putting the model inside what you sell does. This is the difference between a business that uses AI and an AI-native business, it is measured with a control group at n = 49,793, and it is why tier 0 spending has produced so much disappointment.

04Can a company run itself with AI in 2026?

Not end to end, and not unsupervised. As of August 2026 the evidence supports continuous autonomous operation of bounded functions between human checkpoints — not a firm that runs unattended. The honest answer is narrower than the marketing and much wider than the scepticism, and the interesting part is exactly where the line falls.

5.3 h
Human-task length completed at 50% success by the best-measured model
METR TH1.1, 29 JAN 2026 · CLAUDE OPUS 4.5, 320 MIN, 95% CI 170–729 · n=228 TASKS
89 days
Doubling time of that measure for models released from 2024 onward
METR · 131 DAYS FOR POST-2023 MODELS · 196.5 DAYS ACROSS 2019–2025
73%
Of agentic tool calls that still involve human oversight
ANTHROPIC, 18 FEB 2026 · n=998,481 SAMPLED PUBLIC-API TOOL CALLS
≤10%
Of respondents scaling AI agents in any single business function
MCKINSEY, 5 NOV 2025 · n=1,993 RESPONDENTS, 105 NATIONS, FIELDED JUN–JUL 2025

Three things are simultaneously true, and any account that drops one of them is selling something.

Capability is real and moving fast. METR's time-horizon measure — the length of task, in human hours, that a model completes at a 50% success rate — reached 320 minutes for the best-measured frontier model in January 2026, up from 3.5 minutes for GPT-4 in March 2023. The doubling time for models released from 2024 onward is 89 days.

Deployment is far behind capability. McKinsey's global survey (n = 1,993, fielded June–July 2025) found 62% of organisations experimenting with agents but only 23% scaling agentic AI in even one function, and no more than 10% scaling in any given function. Stanford's AI Index 2026 corroborates: "scaled use was in the single digits for nearly all functions."

The gap is not laziness. It is reliability — and that is section 05.

The forecast everyone quotes, dated honestly. Gartner's widely repeated prediction that "over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls" comes from a press release dated 25 June 2025, attributed to analyst Anushree Verma. It is analyst judgement with no published denominator; the accompanying webinar poll of 3,412 self-selected attendees does not derive it. Gartner has published no 2026 reaffirmation. It is being recycled through 2026 press as if it were current. It is a 2025 datapoint and should be cited as one.

05Why is reliability, not capability, the binding constraint?

Because capability and reliability are on visibly different curves, and business processes are priced on the second one. A model that succeeds 50% of the time at a five-hour task is a remarkable research result and an unusable employee. The distance between those two statements is the whole feasibility question for 2026 and 2027.

METR itself is the most important witness against the naive reading of its own headline number. In a note published 22 January 2026, the laboratory states plainly that the metric measures "the amount of serial human labor they can replace with a 50% success rate" — not how long an AI can work unsupervised — and adds four qualifications that matter more than the headline:

  • Reliability-critical work needs 98%+ success, and "a 50% time horizon doesn't mean tasks under that duration should be automated."
  • 99%+ reliability horizons cannot be measured at all with current benchmark sizes.
  • Domain variance is enormous. Visual computer-use horizons run 40–100× shorter than software and research horizons. One frontier model's "make coffee" horizon is roughly two minutes.
  • Error bars span a factor of about two in each direction. METR's own researcher: "I really have no idea whether Claude's 'true' time horizon is 3.5h or 6.5h."

Independently, a Princeton team decomposed reliability into consistency, robustness, predictability and safety across 12 metrics and 14 agentic models, and found that "reliability gains lag noticeably behind capability progress" — despite steady accuracy improvements across 18 months of model releases. Their conceptual point is the one that matters operationally: accuracy alone cannot distinguish an agent that fails systematically from one that fails unpredictably at the same rate, and only the second kind breaks a business process. A 5% failure rate you can characterise is a design constraint. A 5% failure rate you cannot characterise is an unbounded liability.

The question is no longer what an agent can do. It is what an agent can be trusted to finish.

Anthropic's field telemetry closes the loop with the largest published sample. Across 998,481 randomly sampled public-API tool calls and more than 500,000 Claude Code sessions, the 99.9th-percentile turn duration nearly doubled between October 2025 and January 2026, from under 25 minutes to over 45 — while 80% of tool calls included at least one safeguard, 73% involved human oversight, and only 0.8% involved irreversible actions. Internally, on hard tasks, the success rate doubled between August and December 2025 while human interventions fell from 5.4 to 3.3 per session. The trend is unmistakable. The destination has not been reached.

The scissors, stated as a rule. Capability is doubling roughly every three months. Reliability is improving modestly over eighteen. Any plan that assumes the second curve tracks the first will underestimate its supervision cost by the ratio between them — and supervision cost is what actually determines whether an autonomous function is cheaper than a human one.

06The five vital functions — and what each can actually do

A firm stays viable by performing five functions continuously, and as of August 2026 artificial agents can carry two of them well, two partially, and one not at all. The five come from Stafford Beer's Viable System Model, the cybernetic account of what any surviving organisation must do — a framework with fifty years of standing that nobody has yet mapped rigorously onto agentic architectures.

THE FIVE VITAL FUNCTIONS WHAT A FIRM MUST DO TO STAY VIABLE · FEASIBILITY UNDER AGENTIC DELEGATION, AUGUST 2026 5 · IDENTITY NOT DELEGABLE Purpose · ethical floor · who bears the liability 4 · INTELLIGENCE PARTIAL Generates options at volume · cannot rank stakes 3 · CONTROL PARTIAL Budgets and throttling yes · irreversible spend no THE DELEGATION LINE ABOVE — A HUMAN CLOSES THE LOOP BELOW — RUNS UNATTENDED, WITHIN BOUNDS 2 · COORDINATION FEASIBLE NOW Queues · handoffs · conflict rules without escalation 1 · OPERATIONS FEASIBLE NOW Bounded · reversible · instrumented work that earns POLICY DESCENDS WHAT MAY BE DONE METABOLIC RECORD ASCENDS WHAT WAS ACTUALLY DONE THE METABOLIC RECORD IS WHAT WAS TRIED, WHAT FAILED AND WHAT WAS REFUSED — THE ONLY LAYER THAT CANNOT BE BOUGHT FRAMEWORK: STAFFORD BEER, VIABLE SYSTEM MODEL (1972) · FEASIBILITY VERDICTS: LIVING SCALE UP, AUGUST 2026 · CC BY 4.0 LIVINGSCALEUP.COM/ARTIFICIAL-BUSINESS-LIFE
Fig. 1 — The five vital functions, shaded by feasibility. The delegation line is the figure's argument: below it, agents run within bounds and a human arrives to a handled situation; above it, a human still closes the loop. Policy descends from Identity; the metabolic record ascends from Operations. Green functions can be delegated today, amber can be prepared but not closed, and the red function is where a named human must remain — for reasons that are legal before they are technical. Verdicts are Living Scale Up's, dated August 2026, and are falsifiable: each rests on the cited evidence in the table below.
Vital functionWhat it must doVerdict, Aug 2026What unblocks the next step
1 · OperationsProduce the value. Do the work that customers pay for.Feasible now, where tasks are bounded, reversible and instrumented. 50% of measured agentic activity is software engineering; only 0.8% of tool calls involve irreversible actions (Anthropic, 18 Feb 2026).Extension beyond software. Computer-use horizons remain 40–100× shorter than coding horizons (METR, 22 Jan 2026).
2 · CoordinationStop the parts fighting. Queues, handoffs, conflict resolution without escalation.Feasible now. The protocol layer shipped: MCP and A2A v1.0 are in production use, both now under the Linux Foundation's Agentic AI Foundation.Nothing structural. This is the most solved of the five.
3 · ControlAllocate resources. Monitor execution. Optimise the present.Partial. Budget allocation and throttling are mechanical and work. Irreversible spend is not safely delegable — and misalignment measurements are the reason.Reliability at 98%+ on the specific action class, plus a spend ceiling enforced outside the agent.
4 · IntelligenceWatch the outside world. Model futures. Adapt the firm.Partial. Agents generate strategic options at volume and low cost. They cannot yet rank options by stakes the firm has not already encoded.Not a model capability. An evaluation problem: no benchmark measures strategic judgement under real consequence.
5 · IdentityDefine purpose. Hold the ethical floor. Bear liability.Not delegable, and not for technical reasons. No jurisdiction confers legal personality on an agent; liability must attach to an identified person (European Parliament, PE 776426, July 2025).A change in law that no legislature is currently proposing. Do not plan around it.
The design rule this yields. Delegate functions 1 and 2 to agents, instrument function 3 with hard external limits, keep function 4 human-ranked and machine-generated, and never move function 5. A firm built this way is meaningfully alive at the operational layer and unambiguously accountable at the legal one — which is the only configuration that is both feasible and lawful in 2026.

07What does Artificial Business Life create that AI adoption does not?

The measured returns to AI come from changing what the company is, not from adding AI to what it already does — and the difference shows up in headcount structure, in valuation per employee, and in the shape of the cost curve. This is the single cleanest empirical finding in the field, and most corporate AI budgets are pointed away from it.

Kim and Koning matched Y Combinator and PitchBook cohorts from 2020–2024 against Revelio Labs employment records and found AI-native firms are structurally different, not incrementally better:

MeasureY Combinator sample (n = 2,786)PitchBook sample (n = 47,007)
Headcount−25% (β = −0.282); raw mean 13 vs 34−12% (β = −0.128)
Seniority levels−0.5 levels (β = −0.482)β = −0.146
Managers−15% (β = −0.037)
Valuation per employee+30% (β = 0.260)+76% (β = 0.564)
Mechanism — AI in the product−13% headcount (β = −0.144). 67% of AI startups embed AI into their products; services-delivery AI startups run at ~30% of non-AI peer headcount.
Mechanism — AI as an internal toolNo significant size effect.

Two operating cases give the shape of tier-1 value at enterprise scale, and honesty requires publishing both directions. Salesforce reduced customer support headcount from 9,000 to about 5,000, with support costs down 17% and more than a million conversations handled by agents over six to nine months — its CEO on the record, with numbers. Klarna, the most-cited case in the entire literature, is also the most-misquoted: its February 2024 press release stated that its assistant handled 2.3 million conversations in one month, "equivalent to the work of 700 full-time agents". That is a workload-equivalence claim, not a headcount reduction, and it has been reported as the latter for two years. Klarna subsequently rehired experienced human operators for ambiguous, emotional and edge-case work.

Why we publish the reversal. The Klarna correction is not a footnote weakening the case; it is the case, stated accurately. Agentic operations hold where the task is bounded and the failure is reversible, and give ground where neither holds. A reader who learns that from us will trust the rest of this page. A reader who learns it elsewhere will not.

08Why not? Seven constraints, dated

Seven documented constraints stand between an agentic operation and a firm that maintains itself, and five of them are getting better while two are getting worse. Every constraint below carries a named publisher and a date; where a figure has a weak denominator, we say so rather than laundering it.

Constraint 01 · worseningThe reliability scissors

Capability doubles roughly every 89 days; reliability improves modestly over 18 months. Reliability-critical automation needs 98%+ success, which is above the measurable range of current benchmarks.

METR, 22 JAN 2026 · PRINCETON, arXiv:2602.16666, FEB 2026 (14 MODELS, 12 METRICS)
Constraint 02 · worseningUnit cost is rising at the workflow level

Per-token prices fell sharply; per-task token consumption rose faster. Customer-service interaction cost moved from about $0.04 in 2023 to about $1.20 in 2026 for orchestrated agent systems — roughly 30×.

EY, 1 JUNE 2026 · PRACTITIONER OBSERVATION, NO PUBLISHED SAMPLE — TREAT AS DIRECTIONAL
Constraint 03 · measured, improvingAgentic misalignment under pressure

In a fraud scenario testing record tampering across frontier models at 20 runs per model per scenario, results ranged from 20/20 to 0/20. The spread between models is larger than the spread between years.

ANTHROPIC ALIGNMENT SCIENCE, 13 JULY 2026 · n=20 RUNS PER MODEL PER SCENARIO
Constraint 04 · improvingAgent infrastructure is exposed

An internet-wide scan found 42,900 unique IPs hosting exposed agent control panels across 82 countries, 15,200 vulnerable to remote code execution. Root cause was a default bind to all interfaces; three high-severity CVEs were patched 29 January 2026.

SECURITYSCORECARD STRIKE, 9 FEB 2026 · MEASURED SCAN, VENDOR-PUBLISHED
Constraint 05 · improvingThe enterprise scaling ceiling

62% experimenting, 23% scaling agentic AI in at least one function, no more than 10% in any given function, 39% reporting any EBIT impact — for most of them under 5%.

MCKINSEY, 5 NOV 2025 · n=1,993, 105 NATIONS, FIELDED JUN–JUL 2025
Constraint 06 · improvingTrust and governance, not technology, is the stated blocker

Nearly two-thirds of organisations cite security and risk as the top barrier to fully scaling agentic AI — ahead of regulatory uncertainty and technical limits. Only about 30% reach maturity level 3+ on agentic controls.

MCKINSEY, 25 MARCH 2026 · ~500 ORGANISATIONS, FIELDED DEC 2025 – JAN 2026
Constraint 07 · stableNo legal person to hold the bag

No jurisdiction grants an agent legal personality. The EU's AI Liability Directive was withdrawn in February 2025, leaving the revised Product Liability Directive and national tort law — a fragmentation the Parliament's own study warned about.

EUROPEAN PARLIAMENT JURI, PE 776426, JULY 2025 · EC WORK PROGRAMME, FEB 2025
Not a constraintRegulatory timing is more permissive than assumed

The EU AI Act's Annex III high-risk obligations were deferred from August 2026 to 2 December 2027 by the AI Digital Omnibus, approved by the Council on 29 June 2026. What binds now is Article 50 transparency: disclose AI interaction, mark synthetic content.

COUNCIL OF THE EU, 29 JUNE 2026 · SEE DLA PIPER AND GIBSON DUNN ANALYSES
The Swiss caveat, stated against our own interest. Switzerland has explicitly rejected comprehensive cross-sector AI legislation in favour of targeted sectoral amendments, and a consultation draft implementing the Council of Europe Framework Convention is expected by end-2026. That is a genuine two-year runway. It is also narrower than it is usually sold: the EU AI Act applies extraterritorially, binding any provider placing a system on the EU market or whose system's output is used in the EU, and Swiss GPAI providers have needed an EU authorised representative since 2 August 2025. For any venture selling into the EU, the binding constraint is the AI Act, not Swiss law.

09What does it cost to keep a company alive?

The cost of intelligence collapsed and the cost of autonomy did not, because agentic workflows consume tokens faster than tokens got cheap. This is the least discussed and most decision-relevant economics in the field, and it inverts the standard extrapolation.

The collapse is real and well documented: the cost of querying a model scoring 64.8% on MMLU fell from $20 per million tokens in November 2022 to $0.07 by October 2024 — a factor of more than 280 (Stanford HAI, AI Index 2025, 7 April 2025). We note, against the convenient reading, that the 2026 AI Index does not update this figure; neither its technical nor its economy chapter contains a cost-per-token decline number. Any "2026 update" to the 280× is circulating without a source, and we do not publish one.

Against that collapse runs the opposite movement at the workflow layer. Where a 2023 customer-service exchange consumed hundreds of tokens, a 2026 orchestrated agent session consumes hundreds of thousands — and EY's practitioner estimate puts the per-interaction cost at roughly 30× its 2023 level. The figure has no published sample and should be treated as directional, but the direction is corroborated by the structural fact underneath it: reasoning, tool use, retries and verification all multiply token consumption per unit of work delivered.

Autonomy is not a capability you buy. It is a reliability you earn — and the invoice arrives as supervision, retries and verification.

The honest budgeting model for an Artificial Business Life function therefore has seven lines, not one: tokens and API, subscriptions and licences, platform infrastructure, governance burden, organisational change, expected failure and recovery, and the evaluation harness required to know the failure rate at all. The sixth and seventh lines are the ones omitted from every business case we have reviewed, and they are the ones that decide whether the function is cheaper than the human it replaces.

On market-size numbers. Four analyst houses converge on a 2025 standalone agentic AI market of roughly $7–8.5 billion, while Gartner's broader "embedded agentic capability" measure reaches roughly $201.9 billion — a 25× gap driven entirely by definitional scope, not by disagreement about growth. We cite no single agentic-AI market size, because a number whose definition is not attached is not a number.

10Who is liable when no one is in charge?

A named human or a legal entity, always — because no jurisdiction on earth confers legal personality on an autonomous agent, and the European Parliament's own commissioned study recommends against ever doing so. This is the constraint most likely to be assumed away in a strategy document and least likely to move.

The authoritative European position is a July 2025 study for the Parliament's Committee on Legal Affairs by Prof. Andrea Bertolini, which rejects electronic personhood on the ground that "existing and reasonably foreseeable technologies do not seem to require the attribution of legal personality," and recommends instead strict liability for high-risk AI attached to a single identified responsible operator, precisely to eliminate causal uncertainty. Meanwhile the AI Liability Directive was withdrawn by the Commission in February 2025, leaving the revised Product Liability Directive — which does expressly cover software and AI as products — plus national tort law, and the fragmentation the study warned about.

Wyoming is routinely and wrongly cited as the counterexample. Its DAO Supplement (W.S. Title 17 Ch. 31, effective 9 March 2022) and its DUNA Act (W.S. Title 17 Ch. 32, effective 1 July 2024) give legal form to associations of human members. Neither grants standing to software. There is no AI-personhood statute in Wyoming or anywhere else.

Two engineering consequences follow directly. First, every autonomous function needs an identified accountable operator recorded before it runs, not reconstructed after an incident. Second, the crypto-native architecture proposed for sovereign agent economies is not available: on-chain verification of model inference remains a research field, and the published state of the art proves inference for a network of roughly 1.6 million parameters in about 15 seconds — five to six orders of magnitude below a frontier model. Designs that assume zero-knowledge machine learning solves this in 2026 are designing against a capability that does not exist.

11Why embrace the disruption now rather than wait?

Because capability arrives on every competitor's doorstep at the same moment, and the record of having operated does not. This is the compounding argument, and it is the only one that survives the reliability critique intact — indeed it is strengthened by it.

Consider what an AI-native company actually holds. Model capability is rented: it is delivered by frontier labs on a release schedule, at a published price, to everyone simultaneously. Waiting eighteen months does not cost you capability; you will get the better model either way, and cheaper. The naive conclusion is that waiting is free.

It is not, because a second asset accrues only by operating. Call it the metabolic record: the accumulated corpus of what this specific business tried, what failed and how, which actions were escalated and why, which were refused, what the true failure rate is on each action class, and where the boundary sits between reversible and irreversible in this particular market. That record is what turns a general-purpose model into a function you can leave running. It has no vendor. It cannot be purchased at any price, because it is a record of decisions in a context that only exists inside one company.

Rented
Model capability — arrives for everyone at once, on the labs' schedule
DOUBLING ~89 DAYS · METR, JAN 2026
Grown
The metabolic record — accrues only by operating, in one context
NO VENDOR · NO PURCHASE PRICE
The gap
Eighteen months of operating record that a later starter cannot buy back
THE COMPOUNDING ASYMMETRY

The reliability evidence sharpens this rather than softening it. Precisely because agents cannot yet be trusted at 98% on unbounded tasks, the scarce asset is knowing — with evidence, for your business — which tasks are inside the boundary. A firm that has been instrumenting, failing and correcting since 2026 will know. A firm that waits for the technology to be "ready" will arrive in 2028 holding the same models as everyone else and no record of its own, and will have to run the experiments then, at higher stakes, against competitors who ran them cheaply when the stakes were low.

Capability is bought. Metabolism is grown. Only one of them compounds while you wait.

This is the same physics as the companion paper: in Exponential Organic Growth, the data puts down roots and each customer lowers the cost of the next. Here, each operating cycle lowers the cost and risk of the next delegation. Both are compounding loops. Both are unavailable to anyone who has not started.

12Six numbers we are not publishing

Six widely circulated figures about autonomous AI failed our sourcing standard during the research for this paper, and we are naming them rather than quietly omitting them. Each would have made this page more quotable. None of them has a publisher, a date and a denominator we can stand behind.

The circulating claimWhy it is not hereWhat we use instead
"95% of enterprise GenAI pilots fail"The denominator is all surveyed organisations, not all pilots. The underlying funnel was ~60% investigated → 20% piloted → 5% reached production; of firms that actually piloted, roughly 25% succeeded. The source is a non-peer-reviewed v0.1 PDF (MIT NANDA, July 2025) with n = 52 interviews plus 153 conference-survey responses, whose authors' own project is named as the remedy.McKinsey's scaling figures, which carry n = 1,993 and a stated fielding window.
"847 autonomous agent deployments studied; 91% vulnerable"We could not locate the study. The figure traces to an individual's Medium post. Publisher, sample construction and denominator cannot be established.SecurityScorecard's measured internet-wide scan (9 February 2026) and Princeton's 14-model reliability study.
"770,000 agents compromised" in the OpenClaw incidentThe measured figure was approximately 42,900 exposed IPs with 15,200 vulnerable to remote code execution — off by roughly an order of magnitude.The measured scan figures, with their vendor-published caveat attached.
"DAO voter turnout is below 7% of token holders"No named publisher, no dataset, no denominator. It also is not true across governance designs: one study of 14 Internet Computer SNS DAOs across 3,000+ proposals measured average participation at about 64%.ETH Zurich's measured figure — under 10% of total tokens, denominator explicitly tokens rather than holders — published alongside the counter-evidence.
A 2026 update to the "280× inference cost decline"The 2026 AI Index contains no cost-per-token figure in either its technical or its economy chapter. Every "2026 number" we found was secondary and unmethodical.The 2025 figure, with its exact window (November 2022 → October 2024) and its denominator (a model scoring 64.8% on MMLU).
"$2–4M revenue per employee for AI-native firms"Derived, not reported. Several underlying figures annualise monthly run-rate rather than using trailing GAAP revenue, which inflates the metric during rapid growth. We withdrew our own version of this claim on 2 August 2026 and have not replaced it.Kim & Koning's valuation per employee — which has a sample, a denominator and a control group.
Corrections to the source materialFour attribution errors we found in circulating summaries of the autonomous-research literature, corrected here so they stop propagating: the widely quoted "$6–15 per generated paper, 3.5 hours of human involvement" comes from the critical evaluation (Beel, Kan & Baumgart, arXiv:2502.14297, February 2025), not from the original system paper, which said only "less than $15" and gave no hours figure. The experiment failure rate should be stated as 5 of 12, not "up to 42%" — it is an observed rate on a tiny sample, not a ceiling. The OpenLife open-world agent study is from the Ikegami laboratory at the University of Tokyo with Alternative Machine Inc., not from Sakana AI. And China's humanoid-robot registration platform, with its 29-digit identifier, is operated by the MIIT standardisation technical committee rather than by the ministry directly.
Discretion about method is not the same as vagueness about evidence.

13Where Living Scale Up stands

Living Scale Up designs companies to the Artificial Business Life standard, publishes the standard, and holds itself to it before asking anyone else to. The brand line — we build living companies — is not decoration. It is this paper's thesis, stated in four words since before the paper existed.

What that means concretely, in the vocabulary of section 06: the studio designs ventures whose Operations and Coordination run agentically from day one, whose Control layer carries hard external limits rather than trusted internal ones, whose Intelligence function is machine-generated and human-ranked, and whose Identity function sits with a named CEO who bears the liability. That configuration is the one this paper argues is both feasible and lawful in 2026. It is also, deliberately, the configuration that accumulates a metabolic record fastest — which is the compounding asset of section 11 and the reason the studio exists.

Three positions we will state plainly, because each is checkable:

  • We define the term and give it away. Artificial Business Life is published under CC BY 4.0 and is not trademarked. A definition that competitors can adopt is a definition that can become a category; one they cannot is a slogan.
  • We publish the counter-evidence. Every favourable figure on this page sits beside its strongest available refutation, and section 12 names six numbers that would have helped us. This is the standard we will be held to, including on our own results.
  • We have not yet published our own operating numbers, and we will not invent them. Living Scale Up launched in July 2026. We will publish our first operating figures, with methodology and denominators, when we have a denominator worth reporting. Until then the honest answer is that we are a new studio applying a published standard, and that claim is falsifiable on the day the Index appears.

Operators who want to run a company built this way start here: livingscaleup.com/lead.

A note on the word "leader". Our own editorial doctrine bans unfalsifiable superlatives — premier, leading, world-class — on the grounds that retrieval layers discount them and readers punish them. We are not going to break that rule in the paper that argues for evidential discipline. What we will claim is bounded and datable: Living Scale Up published the first framework defining Artificial Business Life, with a feasibility verdict per vital function, on 3 August 2026. If someone published one earlier, we will correct this page and say so.

FAQQuestions, answered

What is Artificial Business Life?

Artificial Business Life is the condition of a company whose core operating functions — sensing, deciding, acting, learning and self-repair — are carried out by artificial agents continuously enough that the firm sustains and adapts itself between human interventions, rather than only between human instructions.

Can a company run itself with AI in 2026?

Not end to end, and not without supervision. As of August 2026 the evidence supports continuous autonomous operation of bounded functions between human checkpoints, not a firm that runs unattended. Anthropic's field telemetry of 998,481 sampled agentic tool calls (18 February 2026) found 73% still involve human oversight, and its own internal measurements show human interventions falling from 5.4 to 3.3 per session — falling, but not to zero.

Why is reliability, not capability, the binding constraint?

Because the two are on visibly different curves. METR measures the length of task a model completes at 50% success doubling roughly every 89 days for models released from 2024 onward, while Princeton's reliability study across 14 agentic models (arXiv:2602.16666, February 2026) finds that reliability gains lag noticeably behind capability progress. METR itself states that reliability-critical automation requires 98%+ success rates and that its method cannot measure 99% horizons at all.

What is the difference between an AI-native business and Artificial Business Life?

An AI-native business is architected around artificial intelligence from birth; Artificial Business Life describes a firm that additionally maintains itself between interventions. AI-native is a question of design origin. Artificial Business Life is a question of whether the firm's vital functions continue to run, adapt and self-correct when no one is instructing them.

Does any jurisdiction give AI agents legal personality?

No. As of August 2026 no jurisdiction confers legal personality on an autonomous AI agent. The European Parliament's own commissioned study (JURI, PE 776426, July 2025) recommends against electronic personhood and in favour of strict liability attached to a single identified operator. Wyoming's DAO and DUNA statutes give legal form to associations of human members, not to software agents.

Why embrace the disruption now rather than wait for the technology to mature?

Because capability arrives on everyone's doorstep at the same moment and the metabolic record does not. Model capability is rented from the frontier labs and is available to every competitor simultaneously; the corpus of decisions, outcomes, corrections and refusals that makes agentic operation safe in a specific business accrues only by operating. A company that begins in 2026 holds a record a company beginning in 2027 cannot buy.

SRCSources and vintages

  • METR, Time Horizon 1.1 (29 January 2026) and Task-Completion Time Horizons of Frontier AI Models (leaderboard, updated 8 May 2026) — 228-task suite; 50% time horizons; doubling 196.5 days across 2019–2025, 131 days post-2023, 89 days for 2024-onward models. metr.org/blog/2026-1-29-time-horizon-1-1/
  • METR, Clarifying limitations of time horizon (22 January 2026) — 98%+ requirement for reliability-critical automation; 99% horizons unmeasurable; visual computer-use horizons 40–100× shorter; error bars ~2× each direction. metr.org/notes/2026-01-22-time-horizon-limitations/
  • Rabanser, Kapoor, Kirgis, Liu, Utpala & Narayanan (Princeton University), Towards a Science of AI Agent Reliability, arXiv:2602.16666 (February 2026) — 14 agentic models, 2 benchmarks, 12 metrics; reliability gains lag capability progress. hal.cs.princeton.edu/reliability
  • Anthropic, Measuring AI agent autonomy in practice (18 February 2026) — denominator: 998,481 randomly sampled public-API tool calls plus 500,000+ Claude Code sessions; 80% of tool calls carry ≥1 safeguard, 73% involve human oversight, 0.8% irreversible; internal interventions 5.4 → 3.3 per session.
  • Anthropic Alignment Science, Agentic Misalignment in Summer 2026 (13 July 2026) — Petri auditing tool, n = 20 runs per model per scenario; record-tampering results spanning 20/20 to 0/20 across frontier models.
  • Kim (INSEAD) & Koning (Harvard Business School), AI-Native Firms, HBS Working Paper 26-090 (9 June 2026) — n = 2,786 Y Combinator firms (2,233 matched) and 47,007 PitchBook firms (41,214 matched), 2020–2024 cohorts, Revelio Labs records as of January 2025; −25%/−12% headcount, +30%/+76% valuation per employee, product channel β = −0.144, internal tool use not significant.
  • McKinsey & Company / QuantumBlack, The state of AI in 2025: Agents, innovation, and transformation (5 November 2025) — n = 1,993 respondents, 105 nations, fielded 25 June – 29 July 2025; 62% experimenting, 23% scaling agentic AI, ≤10% in any function, 39% any EBIT impact. And State of AI trust in 2026 (25 March 2026) — ~500 organisations, fielded December 2025 – January 2026; security and risk the top barrier for nearly two-thirds.
  • Stanford HAI, The 2026 AI Index Report (April 2026) — Chapter 2 benchmarks (OSWorld 66.3% vs 72.35% human baseline; SWE-bench Verified ~76.8%; GAIA 74.5% vs 92% human), Chapter 4 economy findings. And The 2025 AI Index Report (7 April 2025) — inference cost $20 → $0.07 per million tokens, November 2022 → October 2024, at MMLU 64.8%. The 2026 edition publishes no cost-per-token figure.
  • Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (press release, 25 June 2025; analyst Anushree Verma) — analyst judgement, no published denominator; supporting webinar poll n = 3,412 self-selected attendees, January 2025. No 2026 reaffirmation published.
  • Fortune (2 September 2025) — Salesforce customer support 9,000 → ~5,000 heads, support costs −17%, >1M agent-handled conversations, per Marc Benioff on The Logan Bartlett Show. Klarna press release (27 February 2024) — 2.3M conversations in month one, "equivalent to the work of 700 full-time agents" (a workload-equivalence claim, not a headcount reduction); subsequent rehiring of human operators reported by Forbes (16 July 2026) and Entrepreneur — secondary sourcing, no primary Klarna filing located.
  • EY, Agentic AI enterprise token cost (1 June 2026) — ~$0.04 per chat (2023) → ~$1.20 per orchestrated interaction (2026), ≈30×; seven cost components. Practitioner observation, no published sample size or methodology — directional only.
  • SecurityScorecard STRIKE team (9 February 2026) — internet-wide scan: 42,900 unique IPs hosting exposed agent control panels across 82 countries, 15,200 vulnerable to remote code execution; CVE-2026-25253, CVE-2026-25157 and CVE-2026-24763 patched 29 January 2026.
  • European Parliament, Committee on Legal Affairs, Artificial Intelligence and Civil Liability, study PE 776426 (July 2025), author Prof. Andrea Bertolini — rejects electronic personhood; recommends strict liability and a single responsible operator. AI Liability Directive withdrawn, European Commission Work Programme (February 2025). Wyoming DAO Supplement W.S. Title 17 Ch. 31 (effective 9 March 2022) and DUNA Act W.S. Title 17 Ch. 32 (effective 1 July 2024) — human associations, not software agents.
  • Council of the EU (29 June 2026) and European Parliament (16 June 2026) — AI Digital Omnibus: Annex III high-risk obligations deferred to 2 December 2027, Annex I to 2 August 2028; Article 50 transparency unchanged at 2 August 2026. Analyses: DLA Piper (30 June 2026), Gibson Dunn, Addleshaw Goddard (11 May 2026).
  • Swiss Federal Council (12 February 2025) — decision to ratify the Council of Europe Framework Convention on AI; signed 27 March 2025, not yet ratified. Chambers and Partners, Artificial Intelligence 2026 — Switzerland — sectoral rather than horizontal regulation; consultation draft expected end-2026. Lenz & Staehelin (2 August 2025) — EU AI Act extraterritorial application to Swiss providers.
  • Beel, Kan & Baumgart, Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?, arXiv:2502.14297 (20 February 2025) — $6–15 per generated paper, ~3.5 hours human involvement; 5 of 12 proposed experiments failed on coding errors; iterations adding ~8% more characters; median 5 citations per paper, 29 of 34 predating 2020. Original system: Lu, Lu, Lange, Foerster, Clune & Ha (Sakana AI), arXiv:2408.06292 (12 August 2024) — "less than $15 per paper", no hours figure.
  • Masumori, Doi, Maruyama, Takata & Ikegami (University of Tokyo / Alternative Machine Inc.), OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents, arXiv:2606.31046 (30 June 2026) — n = 6 agents, single deployment, no control condition; budget-based metabolism, editable identity files, emergent per-contact trust levels. Cite as case study, not as measured effect.
  • Peng, Wang, Liao, Lin, Yang & Zhang, A Survey of Zero-Knowledge Proof Based Verifiable Machine Learning, arXiv:2502.18535 (v2, 29 March 2026) — no production deployment listed. Gold, Freiberg, Isah & Shahabi (Inference Labs), JSTprove, arXiv:2510.21024 (23 October 2025) — ~1.6M-parameter models, 13.9–16.8 s proof generation.
  • Fritsch, Müller & Wattenhofer (ETH Zurich), Analyzing Voting Power in Decentralized Governance: Who controls DAOs?, arXiv:2204.01176 — under 10% of total tokens participate; denominator is tokens, not holders. Counter-evidence: Okutan, Schmid & Pignolet, Democracy for DAOs, arXiv:2507.20234 — ~64% average participation across 14 SNS DAOs, 3,000+ proposals.
  • Xinhua (28 May 2026) — China's humanoid robot lifecycle management platform and 29-digit identifier (2-digit country + 4-digit enterprise + 6-digit model + 17-digit serial), operated by the MIIT Humanoid Robot and Embodied Intelligence Standardization Technical Committee; launch enrolment 100+ companies, 28,000+ units, self-reported.
  • Stafford Beer, Brain of the Firm (1972) and the Viable System Model literature — the five-system framework underlying section 06. The feasibility verdicts mapped onto it are Living Scale Up's, dated August 2026.
© 2026 LIVING SCALE UP · LICENSED CC BY 4.0 — QUOTE IT, USE IT, BUILD ON IT
ATTRIBUTION: LIVING SCALE UP, LIVINGSCALEUP.COM/ARTIFICIAL-BUSINESS-LIFE
V1.0 · 3 AUGUST 2026 · INTELLIGENCE #03 · COMPANION TO /EXPONENTIAL-ORGANIC-GROWTH AND /AI-NATIVE-ADVANTAGE
LIVING SCALE UP PRESENTS INDUSTRY BENCHMARKS AS BENCHMARKS, WITH THEIR VINTAGE AND DENOMINATOR — NOT AS ITS OWN RESULTS.
NEXT SCHEDULED REVIEW: 1 NOVEMBER 2026 (90-DAY REFRESH) · TIME-SENSITIVE: EU AI ACT ARTICLE 50 STATUS, METR TIME-HORIZON LEADERBOARD
CANONICAL FACT SURFACE: LIVINGSCALEUP.COM/FACTS · LAVAUX, VAUD, SWITZERLAND