
廉价的 agentic 劳动,并没有替你回答“我该做什么”。它给你的,是极大变便宜的搜索——而大多数人仍然坐在扶手椅里消耗它。
有一种很现代的烦恼:“我现在几乎什么都能建造——为什么还是想不清什么值得建造?” 那种感觉是,丰裕本应让选择变容易。事实恰好相反。本文解释为什么,并给你一个方法:用廉价、可证伪的实验替代冥想式思考。
这篇文章不是一个自信声音的产物,而是一次对抗性过程的产物:一个第一版框架,一轮严厉批判,后者击碎了框架里几条承重主张;还有带日期的网络研究,用来约束双方诚实。凡是证据和批判赢了的地方,原框架就输了——而那些修正,正是最有用的部分。
换个框架:建造变便宜了,最后一公里没有#
从第一性原理开始。当生产一个东西的边际成本降向零,它的价格也会被竞争压向零——除非 买家与这种商品之间隔着某个不可复制的东西。这是标准经济学,也正是为什么“建造很便宜,所以机会到处都是”这个直觉是反的。让_容易_的部分成本坍缩,只是把全部价值重新搬到那些一直困难的部分。
但真正重要的修正是:AI 让代码和某些认知劳动变便宜了。它没有让交付结果变便宜。 任何真实业务背后的完整生产函数,仍然包括理解没有文档的工作流、获得系统和客户的接入权、接入丑陋的遗留基础设施、处理例外、承担责任、为应收账款融资,以及改变人们的工作方式。这些都没有变便宜。机会很少是“拥有一个稀缺资产”。它通常是:跨过那段所有其他人都拒绝跨、或者跨不过去的最后一公里。
数字让这个缺口变得具体。88% 的调查受访者说,他们在至少一项业务职能中经常使用 AI——一年前是 78%;但只有 39% 说 AI 在企业层面对 EBIT 产生了影响(McKinsey,The State of AI in 2025,2025 年 11 月)。几乎人人都在“使用 AI”,但只有少数人从中得到真实利润——整个游戏都在这一项统计里。区分二者的不是模型接入,而是工作流重设计、组织改变,以及那些毫不华丽却能闭环交付的工作。
问题不是 “我能建造什么?” 建造现在已经是商品。问题是 “我能赢下什么交易,以及当我反复赢下它时,会积累下什么?”
先决定你在玩哪一局#
在选择想法时,最大的单一错误,就是用同一个框架回答三个不同问题。它们不是同一个问题,奖励的市场、证据和防御程度也不一样:
| 这局游戏 |
真正的问题 |
不该借来的错误工具 |
| 现金流——例如每年 50 万美元所有者收益 |
我能不能可重复地赢下一笔付费交易? |
风险投资式的“耐久护城河”要求 |
| 耐久公司 |
我服务每个客户时,什么会复利增长? |
“开始之前我必须先有护城河” |
| 风险投资级——幂律异常值 |
这能不能成为一个定义品类的垄断者? |
Bootstrapper 式的“只跟着发票走” |
一个生意可以靠本地性、便利、关系、运营卓越、市场碎片化和普通的客户惯性,连续多年赚到很好的所有者收益——这些都不是耐久的“护城河”,但全都能还房贷。把风险投资级的耐久性标准搬进普通现金流生意,是有能力的人说服自己放弃好钱的方式。
重要
先做这件事,再做别的。 写下你的目标函数:目标年收入、可接受的首笔收入时间、愿意冒险的资本、最长销售周期、对雇员和运营的胃口、对监管和责任的容忍度。没有它,“我该建造什么?”这个问题在字面上就是不可回答的——而一旦把它说清楚,大多数时髦想法会当场消失。
六个听起来聪明、却会让你亏钱的直觉#
下面每一条,都是一种貌似合理的规则;它们一直能活到遇见市场为止。这里按“听起来聪明 → 实际上”列出,因为诱人的那个版本,才是你默认会伸手去拿的版本。
1 · 建造之前,先命名护城河#
听起来聪明。 “没有防御性就会归零竞争,所以没有护城河我不该开始。”
实际上。 大多数早期公司没有护城河——它们有的是一个 wedge 和一个假设。一开始就要求护城河,会让你发明想象中的数据护城河,拒绝利润可观的失衡型生意,并且过度建造。建造之前真正该问的是:“如果这件事成立,每多一个客户、每多一笔交易、每多运营一个月,什么会在结构上变好?” 那是一个复利向量,不是你必须已经拥有的护城河。
2 · 按结果收费意味着高利润率#
听起来聪明。 “卖最终结果,而不是卖席位,你就能拿下整个劳动力预算。”
实际上。 Outcome pricing 改变的是_谁承担风险_——它并不会消除风险。你继承了坏输入风险、范围模糊、返工、人类例外处理、保证、营运资本和归因争议。只有当结果可归因、可审计、高频、有边界且保险成本低时,它才是高毛利。否则,它只是披着软件倍数外衣的咨询。
3 · 工具位置弱;要拥有结果#
听起来聪明。 “谁都能做工具;钱在交付结果里。”
实际上。 轴错了。真正的区分是:控制点,还是可替代参与者。 一个“工具”如果成了记录系统、授权层、账本,或者工作被_批准_的地方,它就是价值链里最强的座位。一个“结果公司”如果只是调用别人的 API 再交付一个成果,可能什么都没有拥有。
4 · 受众是耐久护城河#
听起来聪明。 “公开建造,积累关注者,你就永远有分发。”
实际上。 大多数 audience 都是从平台那里_租来_的,身份弱、价格敏感,而且附着在一个人格上,而不是一个产品上。Build-in-public 是一种渠道战术:当你的买家活在公共网络里时有效,例如创始人、开发者、创作者;对医院账单团队或市政采购几乎没用。耐久单位不是触达量,而是一个可重复的获客系统:触发器 → 可触达买家 → 可信信息 → 转化 → 留存/转介绍循环。
5 · 从钱已经流动的地方开始#
听起来聪明。 “找到已有支出,把自己插进去——需求已经被证明。”
实际上。 对现金流来说,这是一个强默认选择,但它偏向替代,低估了扩张。既有支出验证了预算,同时也把你锚定在旧工作流上,并邀请竞争者进场。廉价 AI 也可以_创造_需求——“为一个人做的软件”、无处不在的定制、新的消费者行为。既有需求是更高概率的现金路线;新品类是更高方差的异常值路线。不要把这两局游戏混为一谈。
6 · 从你的不公平接入出发#
听起来聪明。 “从你已经有优势的领域、受众或细分市场开始。”
实际上。 这是一个不错的起始启发;但如果你把当前接入当成命运,它就会变成牢笼。优势可以被_制造_:一年偏执地专攻一个领域,收购一个小型在位者,去一线运营岗位干活,和持证运营者合作,或者只是亲手把活干到隐藏结构变得可见。接入应该降低你的搜索成本,而不是定义你的搜索边界。
另一个相关陷阱,是把著名框架当成一套统一教义。它们做的是不同工作。Thiel 讨论的是幂律风险投资回报;Helmer 的 7 Powers 解释的是商业系统形成之后的_持续_回报;Hormozi 是 offer 构造;indie-hacker 实践是低资本搜索;roll-up 讨论的是收购价格和营运资本。没有哪一个是“该建造什么”的通用算法——每一个都只是索引到你上面选择的那局游戏的工具。
方法:付费 wedge 锦标赛#
廉价 agentic 劳动真正给你的礼物,不是更便宜的_建造_,而是更便宜的搜索。所以,不要坐在扶手椅里_选择_一个想法。让三个候选市场用证据竞标你的投入。这个过程会一起发现需求、分发、可行性、例外成本和可能的防御性——而不是试图提前把它们全都想明白。
定义这局游戏。 上一节里的目标函数——现金流目标或风险投资目标、12 个月收入目标、最大投入、最长销售周期、你能接受的运营复杂度。仅这一项就能淘汰掉大多数时髦机会。
寻找交易形问题。 找一种工作流:有可观察触发器、有具名 owner、有可度量完成状态、反复高频出现、失败代价昂贵、输入可获得,并且_错误可检测_。弱问题:“改善营销。”强问题:“在三分钟内回应每一个合格的商业屋顶 lead,生成合规估价,并预约现场检查。”
卖出三个付费 concierge offer。 对每个候选方向,在建产品之前,用人类 + agents 交付结果。只接受强证据——钱、系统接入、数据,或者有约束力的运营承诺。拒绝弱证据:赞美、问卷热情、等待名单、无约束力的意向书。
记录脏活。 跟踪触达 → 会议 → 付费试点转化、成交时间,以及扣除全部人类审核之后的毛利。最重要的是,度量例外集——它是最容易静悄悄杀死大多数“AI 替代劳动”论题的指标。
用四道门选择,然后只构建瓶颈。拉力: 买家是否异常快速地承诺?经济性: 扣除例外、返工和责任之后,利润率是否有吸引力?可重复性: 同一个触发器、买家、offer 和渠道,能不能赢下接下来的十次?复利性: 交付是否积累数据、集成、信任、转介绍或工作流控制?然后,只构建被重复交付暴露出来的那个瓶颈——不要多做。
时钟也支持这个判断。2025 年,20% 的 Stripe Atlas 初创公司在 30 天内向客户收费;2020 年这个比例是 8%——time-to-revenue 正在压缩(Stripe Atlas,year in review)。
杀手指标:例外经济学#
任何 agent 生意的决定性数字,不是 benchmark accuracy。它是残余例外集的成本与后果——也就是 agent 无法干净处理的那些 case。
一个 agent 如果能处理 95% 的 case,并且剩下 5% 明显、容易捕捉、处理便宜,那就很好。如果那些失败安静、灾难性,或者需要资深专家诊断,它就毫无价值。所以要跟踪每笔交易里的例外数、人类审核分钟数、返工率、假阴性成本、升级所需技能级别、责任暴露——最重要的是,例外是否随着规模增长而下降。 如果不会下降,你的“软件”生意就是一个资产负债表更糟的人力生意。
说明
为什么这件事在 2026 年比看起来更重要。 下面的留存和毛利数据不是抽象的市场天气——它们是成千上万家企业的总影子:这些企业的例外经济学从来没有收敛。无声失败最先表现为流失,然后表现为毛利;它们出现时,距离那个完美 demo 已经很久了。
一切都会归零吗?一个诱人的半真相#
只有在严格条件下,价格才会坍缩到边际成本:产品同质、质量可观察、比较便宜、切换便宜、进入自由且即时、产能不受限、没有市场力量,以及需求不会因为价格下降而扩张。真实市场会同时违反其中好几条。软件早就有近零复制成本,Bloomberg、Adobe 和 Salesforce 却从来没有变免费——客户付钱买的是可靠能力、集成和组织合法性,不是复制出来的 bits。
所以,诚实的论题比“一切都会归零”窄得多:通用能力会商品化;但已交付、可信、特定上下文中的表现,不会自动随之归零。 不过,证据确实两边都切;假装不是这样就不诚实了。
怀疑者的证据——AI-native retention 很残酷。 Revenue retention 衡量去年收入中有多少留存下来;越高越健康。
| Cohort |
Revenue retention |
| AI-native——gross |
40% |
| AI-native——net |
48% |
| B2B SaaS benchmark——net |
82% |
ChartMogul 的 2025 SaaS Retention Report 覆盖了大约 2,700 家 B2B SaaS、600 家 B2C SaaS 和 200 家 AI-native 公司,全部 ARR 超过 25 万美元。AI-native 中位 gross retention 从 2025 年 1 月的 27% 升到 9 月的 40%——改善很快,但仍远低于 SaaS 常模。
不同 cohort 的毛利差异极大。 不存在单一的“AI 毛利”。
| 高增长 AI 公司 cohort |
Gross margin |
| “Supernova”(增长最快、agent/compute-heavy) |
约 25%,经常为负 |
| “Shooting Star”(更像 SaaS) |
约 60% |
| LLM-native vertical AI(Bessemer cohort) |
约 65% |
前两个 cohort 及其毛利来自 Bessemer 自己的 State of AI 2025;约 65% 是 Bessemer 在其 vertical-AI thesis 中报告的、2019 年或之后成立的 LLM-native vertical AI 公司总体。按照下面校准说明的精神,这里有一个 caveat:那个很诱人的解释——这种差异跟工作负载有多依赖 inference 有关——是我的解读,不是 Bessemer 的说法。他们报告的是毛利和 cohort 定义;他们并没有把差距归因于 inference intensity。一个相关且被直接报告的事实是:AI coding tools 在转向 usage-based pricing 之前,毛利是零到负数(TechCrunch,2025 年 8 月)。高人均收入不等于高 gross margin。
与怀疑者案例相对的,是一个真实的乐观者案例。2025 年,企业在生成式 AI 应用上的支出估计为 190 亿美元,基础设施为 180 亿美元;其中 application startups 的收入大约是 incumbents 的两倍(Menlo Ventures——它是利益相关方,但这个方向确实反驳了“所有价值都会坍缩到模型层”)。Wrapper 显然也能变成真正的生意:Cursor 大约两年里从约 100 万美元 ARR 到约 20 亿美元 ARR;Harvey 从约 1 亿美元到约 3 亿美元;Sierra 从零到 1.5 亿美元以上。那些 case 里的价值存在于分发、工作流和品牌里——不在它们共同调用的商品化 API 里。
| 怀疑者案例(留存支持它) |
乐观者案例(收入与估值支持它) |
| AI-native retention:40% gross / 48% net,而 SaaS 为 82% net(ChartMogul) |
190 亿美元 app spend > 180 亿美元 infra;app startups 约为 incumbents 的 2:1(Menlo) |
| 面对 AI risk,switching costs、network effects、intangible assets、efficient scale 显示出“几乎没有预测力”;只有 physical cost advantage 仍然保护公司(Helfert / Westwood,Morningstar,2026 年 2 月) |
Cursor、Harvey、Sierra——建立在商品化 API 之上的 wrappers,成为十亿美元级生意 |
| “No inherent endemic moat in the tech stack”(Martin Casado,a16z) |
Casado 自己的更新:“brand is emerging as a real moat”;分发经常胜过 AI 本身 |
真正悬而未决的是: AI-native 的低留存,究竟是一个会成熟过去的生命周期现象(gross retention 确实 在 2025 年 1 月到 9 月之间从 27% 升到 40%),还是对可替代产品的结构性判决。现在还没有多年份 cohort 数据能定案。看起来可能在商品化之后幸存的东西包括:控制点、例外和责任的最后一公里、实体和受监管运营,以及一个能_可度量地_扩大领先优势的数据反馈循环——最后一个也是所有护城河里最常被过度宣称、最少被证明的。
注意
校准说明。 那些病毒式传播的数字——“到 2026 年 80% 的 AI wrappers 会失败”、“OpenAI 吃掉了 200 多家获融资初创公司”、“90 天流失 65%”——在各种博客里反复出现,却找不到可追溯的一手来源。把它们当作民间传说,不要当作证据。本文引用的数字都有来源和日期;这三个数字因为缺乏可追溯一手来源,被有意排除。
专业人士真正分歧在哪里#
“专业人士怎么看这个问题”没有单一答案,因为他们公开分裂。更有用的动作,是看清断层线,然后有意识地下注:
| 断层线 |
最可能的解法 |
| Apps vs. models/incumbents 谁捕获价值 |
都会。通用横向 wrappers 会被碾碎;如果 apps 拥有写入权限、工作流重设计、责任或专有反馈,就能活下来。 |
| Agents replace SaaS vs. strengthen systems of record |
UI 和例行操作会迁移到 agents;systems of record 会继续存在。奖品是成为那个_有权限写回的行动系统_。 |
| Services-as-software vs. high-margin SaaS |
Services-as-software 会在狭窄、可审计、例外率下降的工作流中获胜;在质量主观、失败很晚才浮出水面的模糊知识工作中令人失望。 |
| Existing demand vs. new-category vision |
既有支出 = 更高概率的现金;新品类 = 最大异常值,但需要真正能承受 timing risk。不同游戏——不要混用建议。 |
| Moat-first vs. speed-first |
Speed-first,但带着复利假设。“产品之前先有护城河”太早;“增长却没有走向 power 的路径”同样天真。 |
| Pure software vs. atoms & regulated markets |
风险调整后最好的下注,往往是混合体:AI 实质改善一个实体或受监管运营——是更高利用率的管道_运营者_,不是“给水管工的 AI 软件”。纯软件保留最大上行,但面对最快克隆。 |
真正该记住的一个重构#
停止试图识别一个永久可防御的想法。识别一笔你现在能赢下的交易,证明它在例外和责任之后的全栈经济性,并追问:反复赢下它,是否会导致一个更强的位置积累起来。
利润首先来自解决问题和完成销售。防御性是反复解决问题的_后果_——不是你开始之前必须拥有的入场券。廉价 agents 没有回答“我该建造什么”。它们让搜索变得极其便宜;专业反应,是把这份礼物花在可证伪的付费实验上,而不是更用力地沉思。
周一就做什么#
- 写一段你的目标函数。三局游戏里,你在玩哪一局?
- 列出三个你已经能触达买家的交易形问题。
- 选最强的一个,写出 concierge offer,并在本周向一个真人索要钱、接入或有约束力的承诺。
- 用手工交付它,让 agents 参与其中。从第一笔交易开始记录例外集。
- 只有在拉力 + 经济性 + 可重复性成立之后,才构建那个被反复交付暴露出来的唯一瓶颈。
先写出第一版框架,然后用一次独立的对抗性批判对它做压力测试;那次批判击碎了其中几条主张——“先命名护城河”、“outcome = 高毛利”、“tool = 弱位置”这些反转都来自那一轮。随后,二者都被拿去和三组带日期的网络研究交叉检查:商业模式、防御性、具体 playbooks。凡是证据与框架相冲突的地方,框架都被修正。文中数字都注明了来源和日期;三个广泛流传却找不到一手来源的统计被排除。
发布前,每一个承重数字都被拿回一手来源重读,而不是相信摘要。McKinsey 的 adoption 和 EBIT 数字,ChartMogul 的 retention 数字和样本构成,Bessemer 的 cohort margins,以及 Menlo 的应用与基础设施支出,都与各自来源一致,并且已经按来源自己的措辞在这里重写。那次核对改变了两件事:一个原本标成“inference-light vertical AI”的 margin row,实际是 Bessemer 的整体 vertical-AI cohort;而 inference-intensity 解释原来是我的解释,不是他们的解释。现在两者都已明确标出。Morningstar 的 moats 主张也核对无误,且是按作者自己的话说的:Westwood 多资产策略首席投资官 Adrian Helfert 给 460 家 S&P 500 公司按 moat strength 和 AI risk 打分,发现五个经典支柱中的四个——switching costs、network effects、intangible assets、efficient scale——“在今天的 AI 环境里几乎没有预测力”,只剩 physical cost advantage。依赖它之前,有一点值得知道:采访时,这项研究仍然即将发布在 SSRN,尚未公开。它是本文最醒目的反护城河数字,而它依托的是一项当时尚未公开的工作——这正是读者应该被直接告知,而不是自己事后发现的那类事。
这和我在 No Verification, No Claim 中描述的是同一种纪律:在没有机械检查的领域里,唯一诚实的防线,就是清楚说出哪些数字有来源,哪些只是民间传说,哪些问题仍然开放。这里的一切,都是在同一约束下进行的推理——有论证、有来源、有日期,但在强意义上没有被验证。
商业模式与定价
防御性与护城河
Playbooks、数据与运营
这里引用的投资人文章都是方向性的,并且常常有自利色彩——基金会描述它们投资的品类。诸如“4.6 万亿美元服务”或“13 万亿美元劳动力”这样的 TAM 数字,是框架,不是已测量收入。留存、毛利和支出数字才是更承重的证据,而且都按来源标了日期。把一切都当作 2026 年中期的状态来看待;这是一个快速变化的领域。
What to Build When Building Is Free
Cheap agentic labor didn't hand you an answer to "what should I make." It handed you radically cheaper search — and most people are still spending it in an armchair.
There's a specific modern vexation: "I can build almost anything now — so why can't I figure out what's worth building?" The feeling is that abundance should make the choice easy. It does the opposite. This piece explains why, and hands you a method that replaces contemplation with cheap, falsifiable experiments.
It's the product of an adversarial process rather than one confident voice: a first-pass framework, a hard critique that broke several of its load-bearing claims, and dated web research to keep both honest. Where the evidence and the critique won, the original framework lost — and those corrections are the most useful part.
The reframe: building got cheap, the last mile didn't#
Start from first principles. When the marginal cost of producing a thing falls toward zero, its price gets competed toward zero too — unless something non-replicable sits between the buyer and the commodity. That much is standard economics, and it's why the instinct "building is cheap, so opportunity is everywhere" is backwards. Collapsing the cost of the easy part just relocates all the value to the parts that were always hard.
But here's the correction that matters: AI made code and some cognitive labor cheap. It did not make the delivered outcome cheap. The full production function behind any real business still includes understanding an undocumented workflow, getting access to systems and customers, integrating with ugly legacy infrastructure, handling exceptions, assuming liability, financing receivables, and changing how people work. None of that got cheaper. The opportunity is rarely "own a scarce asset." It's usually "cross the last mile that everyone else refuses or fails to cross."
The numbers make the gap concrete. 88% of survey respondents report regular AI use in at least one business function — up from 78% a year earlier — while just 39% report EBIT impact at the enterprise level (McKinsey, The State of AI in 2025, November 2025). Nearly everyone "using AI," a minority getting real profit from it — that is the whole game in one statistic. What separates the two isn't model access. It's workflow redesign, organizational change, and the unglamorous delivery work that closes the loop.
The question isn't "what can I build?" Building is the commodity now. The question is "what transaction can I win, and what accumulates when I win it repeatedly?"
First, decide which game you're playing#
The single biggest mistake in idea-selection is answering three different questions with one framework. They are not the same question, and they reward different markets, evidence, and levels of defensibility:
| The game |
The real question |
The wrong tool to import |
| Cash flow — e.g. $500K/yr owner earnings |
Can I win a paid transaction, repeatably? |
Venture "durable moat" requirements |
| Durable company |
What compounds with each customer I serve? |
"I need a moat before I start" |
| Venture-scale — power-law outlier |
Can this become a category-defining monopoly? |
Bootstrapper "just follow the invoices" |
A business can earn excellent owner-earnings for years on locality, convenience, relationships, operational excellence, market fragmentation, and plain customer inertia — none of which is a durable "moat," all of which pays a mortgage. Importing venture-scale durability standards into an ordinary cash business is how capable people talk themselves out of good money.
IMPORTANT
Do this before anything else. Write down your objective function: target annual earnings, acceptable time-to-first-revenue, capital at risk, maximum sales cycle, appetite for employees and operations, tolerance for regulation and liability. Without it, "what should I build?" is literally unanswerable — and specifying it eliminates most fashionable ideas on the spot.
Six instincts that feel smart and cost you money#
Each of these is a plausible-sounding rule that survives right up until it meets the market. They're listed as feels smart → actually, because the seductive version is the one you'll reach for by default.
1 · Name the moat before you build#
Feels smart. "No defensibility means a race to zero, so I shouldn't start without one."
Actually. Most early companies have no moat — they have a wedge and a hypothesis. Demanding a moat upfront makes you invent imaginary data moats, reject lucrative disequilibrium businesses, and over-build. The right pre-build question is "if this works, what improves structurally with every customer, transaction, or month I operate?" That's a compounding vector, not a moat you must already own.
2 · Outcome pricing means high margins#
Feels smart. "Sell the finished result instead of a seat and you capture the whole labor budget."
Actually. Outcome pricing changes who bears the risk — it doesn't abolish it. You inherit bad-input risk, scope ambiguity, rework, human exception handling, guarantees, working capital, and attribution disputes. It's high-margin only when the outcome is attributable, auditable, frequent, bounded, and cheap to insure. Otherwise it's disguised consulting wearing software multiples.
Feels smart. "Anyone can build a tool; the money is in delivering the result."
Actually. Wrong axis. The real distinction is control point vs. replaceable participant. A "tool" that becomes the system of record, the authorization layer, the ledger, or the place where work gets approved is the strongest seat in the value chain. An "outcome company" that calls someone else's API and hands over a deliverable may own nothing.
4 · An audience is a durable moat#
Feels smart. "Build in public, grow a following, and you'll always have distribution."
Actually. Most audiences are rented from platforms, weakly identified, price-sensitive, and attached to a personality rather than a product. Build-in-public is a channel tactic that works when your buyer lives in public networks (founders, devs, creators) and is nearly useless for hospital billing teams or municipal procurement. The durable unit isn't reach — it's a repeatable acquisition system: trigger → reachable buyer → credible message → conversion → retention/referral loop.
5 · Start where the money already flows#
Feels smart. "Find existing spend and insert yourself — demand is already proven."
Actually. A strong default for cash flow, but it favors substitution and underweights expansion. Existing spend validates a budget while anchoring you to the old workflow and inviting competitors. Cheap AI can also create demand — "software for one," pervasive customization, new consumer behavior. Existing demand is the higher-probability route to cash; new categories are the higher-variance route to outliers. Don't confuse the two games.
6 · Build from your unfair access#
Feels smart. "Start with the domain, audience, or niche you already have an edge in."
Actually. A fine starting heuristic that becomes a cage if you treat current access as destiny. Advantages can be manufactured: a year of obsessive specialization, acquiring a small incumbent, taking a frontline operating job, partnering with a licensed operator, or just doing the work manually until the hidden structure becomes visible. Access should lower your search cost, not define your search boundary.
A related trap: treating famous frameworks as one unified doctrine. They do different jobs. Thiel is about power-law venture returns; Helmer's 7 Powers explains persistent returns after a business system forms; Hormozi is offer construction; indie-hacker practice is low-capital search; roll-ups are about acquisition price and working capital. None is a general algorithm for what to build — each is a tool indexed to the game you chose above.
The method: a paid-wedge tournament#
The real gift of cheap agentic labor is not cheaper building — it's cheaper search. So don't choose an idea from the armchair. Make three candidate markets bid for your commitment with evidence. This one process discovers demand, distribution, feasibility, exception cost, and possible defensibility together — instead of trying to reason each out in advance.
Define the game. The objective function from the previous section — cash-flow or venture target, 12-month revenue goal, max investment, max sales cycle, operational complexity you'll accept. This alone eliminates most fashionable opportunities.
Find transaction-shaped problems. Look for a workflow with an observable trigger, a named owner, a measurable completion state, frequent recurrence, expensive failure, obtainable inputs, and detectable errors. Weak: "improve marketing." Strong: "respond to every qualified commercial-roofing lead within three minutes, produce a compliant estimate, and book the inspection."
Sell three paid concierge offers. For each candidate, deliver the outcome with humans + agents before building the product. Accept only strong evidence — money, system access, data, or a binding operational commitment. Reject weak evidence: compliments, survey enthusiasm, waitlists, non-binding letters of intent.
Instrument the ugly work. Track outreach → meeting → paid-pilot conversion, time-to-close, and gross profit after all human review. Above all, measure the exception set — the metric that quietly kills most "AI replaces labor" theses.
Choose on four gates, then build only the bottleneck.Pull: does the buyer commit unusually fast? Economics: is margin attractive after exceptions, rework, and liability? Repeatability: can the same trigger, buyer, offer, and channel win the next ten? Compounding: does delivery accumulate data, integrations, trust, referrals, or workflow control? Then build only the bottleneck repeated delivery reveals — nothing more.
The clock supports this. 20% of 2025 Stripe Atlas startups charged a customer within 30 days, against 8% in 2020 — time-to-revenue is compressing (Stripe Atlas, year in review).
The kill-metric: exception economics#
The decisive number for any agent business is not benchmark accuracy. It's the cost and consequence of the residual exception set — the cases the agent doesn't cleanly handle.
An agent that handles 95% of cases is wonderful if the other 5% are obvious and cheap to catch. It's worthless if those failures are silent, catastrophic, or need a senior specialist to diagnose. So track exceptions per transaction, minutes of human review, rework rate, false-negative cost, the skill level required to escalate, liability exposure — and, most importantly, whether exceptions decline as volume grows. If they don't, your "software" business is a staffing business with a worse balance sheet.
NOTE
Why this matters more in 2026 than it looks. The retention and margin data below aren't abstract market weather — they're the aggregate shadow of thousands of businesses whose exception economics never converged. Silent failures show up first as churn, then as margin, long after the demo looked perfect.
Does everything race to zero? A seductive half-truth#
Price collapses to marginal cost only under restrictive conditions: homogeneous products, observable quality, cheap comparison, cheap switching, free and instant entry, unconstrained capacity, no market power, and demand that doesn't expand when price falls. Real markets violate several at once. Software already had near-zero replication cost, and Bloomberg, Adobe, and Salesforce never became free — customers pay for reliable capability, integration, and organizational legitimacy, not copied bits.
So the honest thesis is narrower than "everything races to zero": generic capability commoditizes; delivered, trusted, context-specific performance does not automatically follow. But the evidence genuinely cuts both ways, and pretending otherwise would be dishonest.
The skeptic's exhibit — AI-native retention is brutal. Revenue retention measures how much of last year's revenue survives; higher is healthier.
| Cohort |
Revenue retention |
| AI-native — gross |
40% |
| AI-native — net |
48% |
| B2B SaaS benchmark — net |
82% |
ChartMogul's 2025 SaaS Retention Report covers roughly 2,700 B2B SaaS, 600 B2C SaaS, and 200 AI-native companies, all above $250K ARR. AI-native median gross retention climbed from 27% in January 2025 to 40% in September — improving fast, but still far below the SaaS norm.
Margins vary enormously by cohort. There is no single "AI margin."
| Cohort of high-growth AI companies |
Gross margin |
| "Supernova" (fastest growers, agent/compute-heavy) |
~25%, often negative |
| "Shooting Star" (more SaaS-like) |
~60% |
| LLM-native vertical AI (Bessemer's cohort) |
~65% |
The first two cohorts and their margins are Bessemer's own, from State of AI 2025; the ~65% is Bessemer's aggregate for LLM-native vertical AI companies founded 2019 or later, reported in its vertical-AI thesis. One caveat in the spirit of the calibration note below: the tempting explanation — that the spread tracks how inference-heavy the work is — is my reading, not Bessemer's. They report the margins and the cohort definitions; they do not attribute the gap to inference intensity. The related fact that is directly reported: AI coding tools ran zero-to-negative margins until they moved to usage-based pricing (TechCrunch, Aug 2025). High revenue-per-employee is not the same as high gross margin.
Set against that skeptic case is a real optimist case. Enterprises spent an estimated $19B on generative-AI applications versus $18B on infrastructure in 2025, with application startups out-earning incumbents by roughly two-to-one (Menlo Ventures — an interested party, but the direction contradicts "all value collapses to the model layer"). Wrappers demonstrably become real businesses: Cursor went from ~$1M to roughly $2B ARR in about two years; Harvey from ~$100M to ~$300M; Sierra from zero to $150M+. The value in those cases lives in distribution, workflow, and brand — not in the commodity API they all call.
| Skeptic case (retention favors it) |
Optimist case (revenue & valuation favor it) |
| AI-native retention of 40% gross / 48% net vs. 82% net for SaaS (ChartMogul) |
$19B app spend > $18B infra; app startups ~2:1 over incumbents (Menlo) |
| Switching costs, network effects, intangible assets, and efficient scale show "almost no predictive power" against AI risk; only physical cost advantage still protects (Helfert / Westwood, in Morningstar, Feb 2026) |
Cursor, Harvey, Sierra — wrappers on commodity APIs become billion-dollar businesses |
| "No inherent endemic moat in the tech stack" (Martin Casado, a16z) |
Casado's own update: "brand is emerging as a real moat"; distribution often beats the AI itself |
Genuinely unresolved: whether AI-native low retention is a life-stage artifact that matures out (gross retention did climb from 27% to 40% between January and September 2025) or a structural verdict on substitutable products. No multi-year cohort data settles it yet. What plausibly survives commoditization: the control point, the last mile of exceptions and liability, atoms-and-regulated operations, and a data-feedback loop that measurably widens a lead — the last being the most over-asserted, least-proven moat of all.
WARNING
Calibration note. The viral figures — "80% of AI wrappers fail by 2026," "OpenAI cannibalized 200+ funded startups," "65% churn in 90 days" — recur across blogs without traceable primary sources. Treat them as folklore, not evidence. The numbers cited in this article are sourced and dated; these three are deliberately excluded.
Where the professionals actually disagree#
"How do the pros think about this" has no single answer, because they're openly split. The useful move is to see the fault lines and place your own bet knowingly:
| Fault line |
Most likely resolution |
| Apps vs. models/incumbents capture the value |
Both. Generic horizontal wrappers get crushed; apps survive where they own write-access, workflow redesign, liability, or proprietary feedback. |
| Agents replace SaaS vs. strengthen systems of record |
UIs and routine actions migrate to agents; systems of record persist. The prize is becoming the system of action with permission to write back. |
| Services-as-software vs. high-margin SaaS |
Services-as-software wins in narrow, auditable workflows with falling exception rates; disappoints in ambiguous knowledge work where quality is subjective and failures surface late. |
| Existing demand vs. new-category vision |
Existing spend = higher-probability cash; new categories = the biggest outliers but demand real tolerance for timing risk. Different games — don't blend the advice. |
| Moat-first vs. speed-first |
Speed-first, with a compounding hypothesis. "Moat before product" is premature; "growth with no path to power" is equally naive. |
| Pure software vs. atoms & regulated markets |
Best risk-adjusted bets are often hybrids where AI materially improves a physical or regulated operation — the plumbing operator with better utilization, not "AI software for plumbers." Pure software keeps the biggest upside but faces the fastest cloning. |
The one reframe to keep#
Stop trying to identify a permanently defensible idea. Identify a transaction you can win now, prove its full-stack economics after exceptions and liability, and ask whether winning it repeatedly causes a stronger position to accumulate.
Profit comes first from solving and selling. Defensibility is a consequence of repeated solving — not an admission ticket you need before you begin. Cheap agents didn't answer "what should I build." They made the search radically cheaper, and the professional response is to spend that gift on falsifiable paid experiments instead of contemplating harder.
What to do Monday#
- Write your objective function in one paragraph. Which of the three games are you playing?
- List three transaction-shaped problems you can already reach a buyer for.
- For the strongest one, write the concierge offer and ask a real person for money, access, or a binding commitment this week.
- Deliver it by hand with agents in the loop. Instrument the exception set from the first transaction.
- Only after pull + economics + repeatability hold, build the one bottleneck that repeated delivery exposed.
Method#
A first-pass framework was written, then stress-tested against an independent adversarial critique that broke several of its claims — the "name the moat first," "outcome = high margin," and "tool = weak" reversals all came from that pass. Both were then checked against dated web research across three angles: business models, defensibility, and concrete playbooks. Where evidence contradicted the framework, the framework was corrected. Figures are attributed and dated; three widely-circulated statistics were excluded for lacking traceable primary sources.
Before publication, every load-bearing figure was read back against its primary source rather than trusted from a summary. The McKinsey adoption and EBIT numbers, the ChartMogul retention figures and sample composition, the Bessemer cohort margins, and the Menlo application-versus-infrastructure spending all match their sources and have been rewritten here in the sources' own terms. That pass changed two things: a margin row originally labeled "inference-light vertical AI" turned out to be Bessemer's aggregate vertical-AI cohort, and the inference-intensity explanation turned out to be mine rather than theirs. Both are now marked as such. The Morningstar moats claim also checks out, in its author's own words: Adrian Helfert, chief investment officer of multi-asset strategies at Westwood, scored 460 S&P 500 companies on moat strength and AI risk and found that four of the five classic pillars — switching costs, network effects, intangible assets, efficient scale — "have almost no predictive power in today's AI environment," leaving only physical cost advantage. Worth knowing before you lean on it: that study was, at the time of the interview, still forthcoming on SSRN. It is the single most striking anti-moat number in this article and it rests on work not yet public — which is exactly the sort of thing a reader deserves to be told rather than left to discover.
This is the same discipline I described in No Verification, No Claim: in a field with no mechanical check, the only honest defense is to say plainly which numbers are sourced, which are folklore, and which questions remain open. Everything here is reasoning under exactly that constraint — argued, sourced, and dated, but unverified in the strong sense.
Sources#
Business models & pricing
Defensibility & moats
Playbooks, data & operations
Investor essays cited here are directional and often self-interested — a fund describing the category it invests in. TAM figures such as "$4.6T services" or "$13T labor" are framings, not measured revenue. Retention, margin, and spend figures are the more load-bearing evidence and are dated to their source. Treat everything as of mid-2026; this is a fast-moving field.