Abhay Venkatesh and Emiliano Garcia-Lopez
October 1, 2026.
AI has produced a proposed solution to a Millennium Prize Problem, yet we still cannot rely on it as an effective executive assistant. [2] It can work through a long, hard proof but still miss why a meeting matters or which request deserves priority. In the Valley, we tend to measure intelligence through one kind of technical skill, such as solving programming-contest puzzles. But with AI, strength along one axis does not reliably carry over to others, so it says less about general intelligence than we might expect. We predict that AI will become extraordinarily good at a growing number of things and stay shockingly bad at others, often within the same field. This is "patchwork AGI": AI fills in the problem space patch by patch, while its success changes the work people want it to do. [4]
We forecast that the current paradigm will not become a strict superset of human or all useful intelligence. We do not expect fast takeoff: a compressed transition from roughly human-level AI to broadly superhuman intelligence, often envisioned through recursive self-improvement. We expect progress to be extremely rapid but persistently uneven. Existing gaps can close quickly while new uses create demands that require further learning, training, and deployment. Much of this work will happen in parallel. As one secondhand anecdote has it, what comes after AI is "just 20 more years of AI." [5]
Two forces keep AI uneven. First, some tasks need more than abundant data. Writing has plenty of training material, yet producing fluent text is different from having something new to say. For people, writing carries higher-level ideas: we learn a concept first, then put it into words or code. Models learn the other way, from the words up, and often never reach the concept. A model that writes an algorithm well in one programming language can do much worse in another, even when given that language's rules. [47] [51] It does not understand the algorithm. It knows the language, and what the algorithm should look like written down in it. Models grow strong along the narrow axes that training rewards, the lines in Figure 1. Much of what lies between those lines depends on tacit knowledge, a feel for people, and useful models of how the world works. We expect the strongest skills to keep pulling ahead while others lag.
Second, when AI gets good at something, it changes the work around it. People take on projects that once cost too much and expect more of their tools. [7] Some of that work can draw on existing capabilities; some needs context or judgment the model has not learned.
Take Cognition. Better code generation made its agent, Devin, possible. Companies using coding agents then needed to govern their access and supply company context. Cognition worked with BlackRock on tools for that purpose. [8] Or take Pangram. Cheap AI writing made people want to know whether a person wrote what they read, so Pangram built a model to help tell them. [4] In each case, useful AI created more work around its use. Our prediction is that some of this new work will continue to outrun dependable AI capability.
We expect AI's largest advantages to remain concentrated in particular tasks. In mathematics, a system can search through possibilities and construct proofs at a scale no person could match. That can look like superintelligence. OpenAI's 166-page Navier–Stokes proof passes a Lean check, yet mathematicians say they have so far learned little from it. [50] But extending a line of reasoning and finding a new way to represent a problem are different achievements. General relativity illustrates the importance of the latter. Our prediction concerns how far strength in one carries over to the other.
None of this makes AI small. Peter Thiel has argued that for decades we've had a lot of progress in the world of bits but not much in the world of atoms. [9] The world of bits has already become enormous: all five of the world's most valuable companies, and eight of the top ten, are built on computing. [45]
Fields are the wrong unit for mapping AI. Within one field, some tasks are easy for models and others nearly out of reach, so we sort tasks by what they demand. Drawing on cognitive science and machine learning, we ask five questions of each task. Can a program check the answer? Is a narrow procedure enough to pass? [11] [12] [52] Does it need a model of physics, people, or causes? [13] How much experience beyond text does it need? [14] How much must it keep track of over time? [11] A yes to the first two makes a task easy to train: a narrow procedure can pass, and results can be graded at scale. The last three mark demands that make reliable performance harder. Together, the five questions measure how general a task is. Table 1 applies them to a range of tasks.
Table 1. What a task demands
| Task | Can a program check it? | Is a narrow procedure enough? | World model needed | Experience beyond text | Tracked over time |
|---|---|---|---|---|---|
| Lean proof search | Yes | Often | Formal rules only | None | One proof |
| Coding puzzle | Yes | Often | Formal rules only | None | One problem |
| Maintaining a large codebase | Partly | Sometimes | How the system behaves; what the team wants | Some social | Months |
| Negotiating a deal | Partly | Sometimes | Other minds | Social | Days to years |
| Robot assembly in a new setting | Yes | Rarely | Intuitive physics | Perception, action | Minutes |
| Writing with something new to say | No | No | People and the world | Social | One piece |
| Running a wet-lab experiment | Partly, slowly | Rarely | Physical and biological causes | Perception, action | Days to weeks |
| Running an executive's calendar | No | No | Other minds | Social | Weeks |
| Proposing a new theory | Only in hindsight | No | Causes, at depth | All of the above | Years |
Green: generally favorable for training. Yellow: mixed or conditional. Red: more demanding. The ordering is illustrative, not a measured ranking; difficulty depends on the task, tools, and setting. (see Appendix A.4)
Checkable tasks are the easiest to train. A test suite or proof checker can grade millions of attempts with no person in the loop, so the strongest axes form where a program can check the answer. The same checker keeps them narrow. It rewards whatever passes, and a model can learn a procedure that passes without grasping what the task seems to ask. [11]
Judgment has no checker, so the industry builds one by hand. Companies like Surge and Mercor pay experts to write tasks, compare answers, and draw up grading rubrics, and human data is fast becoming a tens-of-billions-dollar industry. [17] [18] It has worked remarkably well. Yet even Claude Opus 5.5, a strikingly well-aligned model from a lab built around alignment, still can't serve as a dependable executive assistant. [19]
So judgment gets trained the way the Thesis predicts: one field, one contract, and one rubric at a time. Each deployment turns up new failures, and each failure becomes new work for the data vendors. The loop is now routine: a new benchmark comes out, labs buy task data built to look like it, and the next model scores well. [48] Benchmark gains overstate how general a model has become. Table 2 shows why some patches get built first: they are cheap to check, cheap to retry, and easy to sell. (see Appendix A.1)
Table 2. Training conditions by task
| Domain / task | Data | Feedback | Verification | Simulation | Retries | Ease of monetization |
|---|---|---|---|---|---|---|
| Coding Specified function | Often plentiful | Fast tests | Test coverage matters | Executable setting | Usually cheap | Digital delivery; measurable savings |
| Mathematics Formal proof | Varies by topic | Fast checking | Proof checker | Formal setting | Usually cheap | Downstream applications |
| Robotics Grasp an object | Costly across settings | Often immediate | Clear task outcome | Reality gap | Hardware costs | Hardware, deployment, servicing |
| Social understanding Negotiation | Private context | Mixed or delayed | Goals differ | Hard to reproduce | Social costs | Context, trust, integration |
| Judgment Prioritization | Local context | Often delayed | Conflicting criteria | Hidden priorities | Costly errors | Attribution and delegation |
| Taste Product design | Many examples | Mixed timescales | Partly subjective | Partial prototypes | Cheap prototypes | Digital delivery; subjective value |
Green: more favorable. Yellow: mixed or conditional. Red: harder. Colors are illustrative assessments of training and monetization; conditions vary within each capability and depend on the task, environment, and tools.
Most of this progress comes from data. In small-scale pretraining experiments, better training data did more for compute efficiency than better model designs. [20] That points to no special method, only better data, and data has to be built one field at a time. The experiments used small models, so the result may not hold at frontier scale. (see Appendix A.1)
Training on one skill carries over, but only slightly and selectively. In one controlled study, math training raised math accuracy by 3.13 points while lowering average accuracy on other tasks by 1.81; it could help code without helping more distant tasks. [21] In another, math and coding training each lowered the other's score, while puzzle training raised all three (Table 3A). [22] Skill reaches some neighbors, skips others, and can cost ground elsewhere.
Social skills and tool use show the same pattern. Training a small model on negotiation transferred best between the two price negotiations, and weakly or negatively to scheduling and job interviews (Table 3B). [23] Environments built from coding problems raised customer-service accuracy by 8.7 points on average, but they train tool use directly. [24] Where gains reach a new field, training built for that field usually produced them. (see Appendix A.6)
Table 3. Transfer varies by training domain
A. Math, coding, and puzzles [22]
| Training domain | Math | Coding | Puzzles |
|---|---|---|---|
| Math | +25.00 | −3.23 | +13.35 |
| Coding | −3.31 | +6.49 | +13.48 |
| Puzzles | +6.99 | +3.89 | +52.91 |
Qwen2.5-7B-Base; changes in domain-average accuracy, in percentage points. Positive values indicate improvement; negative values indicate decline.
B. Negotiation and scheduling [23]
| Training domain | Craigslist | Marketplace | Job Interview | Calendar |
|---|---|---|---|---|
| Craigslist | +0.265 | +0.317 | −0.032 | −0.035 |
| Marketplace | +0.184 | +0.664 | −0.108 | −0.033 |
| Job Interview | +0.010 | −0.013 | +0.115 | +0.072 |
| Calendar | −0.151 | −0.002 | −0.168 | +0.239 |
Qwen3-4B-Instruct-2507; changes in mean utility (0–1 scale), calculated from SocialRL Table 3. Rows are training environments; columns are evaluation environments. Selected four-domain submatrix; price-negotiation models use supervised training before RL. Panel magnitudes are not comparable across models and measures.
Even success is measured narrowly. In SocialReasoning-Bench, frontier models usually booked the meeting or closed the deal but often chose a poor time or price, even though the benchmark states the user's preferences outright. [25] Finishing a task is not the same as serving the person who asked for it. (see Appendix A.5)
Expecting general intelligence to emerge from a pile of skills asks more of machines than people manage: in humans, training in chess, music, or working memory barely carries over to anything else. [46] People start with a general way of learning; they do not assemble one from skills.
Most of these studies use small models, and transfer at frontier scale may be larger. The next section describes the test that would settle it.
The patchwork follows from how today's systems learn: they fit narrow procedures to what training rewards, and they largely stop learning once deployed. A general learner would break that pattern, and cognitive science gives a fairly concrete picture of one. People build models of objects, other minds, and causes; make new concepts from familiar parts; and learn to learn, so each new task comes faster than the last. [13] [26] These are inductive biases: structure the learner brings to its experience, not facts it reads off the data.
We would change our minds if AI systems began to show these properties in open settings, without a new round of human-built training for each field. The relevant test is whether a capability transfers without additional task-specific training. Astra's driving result does not isolate that effect because its training history is unresolved. [27] [28] Table 4 lists the five signs that would matter most. Each would show a system learning structure rather than habit. (see Appendix A.3)
Table 4. Signs of a general learner
| Sign | Today | Would count against us |
|---|---|---|
| Skill that survives new rules | Given the rules of an unseen programming language, frontier models solve far fewer problems than in Python. [51] | Variants handled almost as well as the originals |
| Learning on the job | Test-time updates help on novel puzzles [29]; models lose track of state over long episodes. [11] | A new workflow learned from a few examples, kept, and updated as things change |
| Parts that travel | Reusable abstractions, learned one domain at a time. [30] [31] | What is learned in one field speeds learning in the next |
| A model of the people it serves | Partial inferences about what people believe and want. [32] | An executive's calendar run well with no training built for the job |
| A new way of seeing | Longer chains of reasoning inside existing frameworks | A theory revised from sparse evidence, and a framework experts adopt [26] [33] |
Each row pairs current evidence with an observation that would count against our forecast.
The economy would show these signs too. AI already helps build training for other skills: a coding model writes reward functions for robot policies, and generated environments teach planning. [34] [35] But people still choose the task, build the simulator, and judge the result, so this speeds up the industrial process rather than replacing it (Figure 5) (see Appendix A.7). Self-improvement still needs people: models are poor judges of their own failures, so a person still has to find the gap and decide what to train. If AI began finding its own gaps and filling them, spending on human data would fall even as capability broadened, and recursive self-improvement would speed things further. [36] Until then, we expect the patchwork to keep renewing itself. (see Appendix A.2)
Persistent unevenness does not mean slow progress. The next ten years could remake how software is built, how science is done, and what work looks like. Many capabilities can advance rapidly and in parallel while substantial gaps remain and new ones appear. The patches change; the patchwork persists.
Patches become an industry. Every field AI enters needs its own data, environments, and graders, and building them is already a multibillion-dollar business. Expect it to become one of the defining industries of the decade, with labs, data vendors, and domain experts working like suppliers on an assembly line.
Software gets cheap, so there is far more of it. AI does not finish software; it speeds it up. When building gets cheaper, more products become worth building, and each one creates demand for more. [37] Teams shrink: Instinct reportedly reached a $10 billion valuation with 14 employees. [38] What those teams still supply is what the patchwork leaves to people: judgment, taste, knowledge of the customer, and distribution. Software comes first because it is where the patchwork turns over fastest.
The digital world keeps its advantage. Software can improve and spread faster than physical systems can be built and changed. We expect the physical-world companies that do best to be those most closely tied to digital progress, whether they supply its infrastructure or use software and AI to improve their products and operations.
Work moves rather than vanishes. Most jobs will change shape within the decade. Some tasks will be automated outright, and new ones will appear around the AI itself, from governing agents to checking what they produce.
Where companies fit. The patchwork creates three kinds of companies (Table 5). The test for each is simple: does the next model make it stronger or obsolete? A gap is not a moat. A product built around one weakness loses its value when the next model fixes it. [39] The durable position is next to the customer, where new gaps show up first.
Table 5. Where the patchwork creates companies
| Company type | Work | Example | When the next model ships |
|---|---|---|---|
| Data | Build the patches: tasks, environments, graders | Mercor, Surge | Stronger, as long as each new field needs its own data |
| Deployment | Fit models to an organization: context, access, governance | Cognition | Stronger: better models mean more to govern |
| Product | Cover what models still get wrong: verification, oversight, detection | Pangram | Depends on whether the gap renews or closes |
We predict that AI will remain extraordinarily good at some things and shockingly bad at others. Some gaps will persist because they are difficult to close; others will emerge as successful automation changes human activity. Individual patches can be completed while the patchwork keeps renewing itself.
There is much more work to be done in AI. Filling the gaps requires either sufficient transfer from existing strengths or better ways to train each domain. AI may increasingly do that work itself, but we do not yet know how far or how quickly that process will reach.
Software will never be finished. Cheaper implementation makes more projects worth attempting, while changing needs create new problems to understand, specify, and solve. [16] "The End of Software" anticipates cheaper creation and disruption of existing business models. [40] Our argument in "The Amplification of Software" is that those changes can increase the importance of judgment, taste, and distribution. [37] Extraordinary coding ability may therefore leave much of the work of building useful software unresolved. Tacit knowing may be inexhaustible: mapping one piece of the world may reveal more that remains unmapped. [41]
As machines take on human activities, we change how we live, relate to one another, and understand ourselves. Formalizing what we do can reveal further distinctions and create new activities. [42] Our prediction is that these changes will repeatedly create demands that outrun dependable AI capability. Transfer and recursive self-improvement may close existing gaps faster; the decisive question is whether they can also keep pace with the new ones. Software keeps extending its reach, and the work of extending it keeps changing.
Thanks to Rajath Salegame for his feedback on the limits of transfer from coding and mathematics to hardware engineering. Thanks to Ronald Qiao for pushing on whether better coding could accelerate the creation of training environments, data, and more general models. Thanks to Zack Baker for the conversation about tacit knowing and the idea that we will not run out of things to do. Thanks to Sami Senapathy for reading and commenting on drafts of this essay. Thanks to Leonard Tang for pointing out how new uses of AI can draw on existing capabilities. Thanks to Rahul Mitra for his feedback on architectural constraints, continual learning, and sample efficiency. Thanks to Vik Pattabi for reading the essay. Thanks to Jacob Thompson for reviewing Figure 1 and the tables, and for reading parts of earlier drafts. Thanks to Turner Merritt, Nelson Arnous, and Charlie Davidmann for providing feedback on an early draft. Thanks to an anonymous reader from a leading neo-lab for their feedback on this piece and alignment with our thesis.
Data and feedback are only part of the bottleneck. Architecture and learning algorithms shape which capabilities a model acquires easily and how well it retains them. Adapting within a conversation differs from learning durably across experiences. Test-time learning and learned memory are active research directions; reliable continual learning remains an open challenge. [44] Sample efficiency matters too: when useful feedback is costly or slow, the amount of experience needed to learn can determine whether a capability is practical to develop. Closing some gaps may therefore require better architectures or learning methods as well as better data. These constraints can change as the methods improve.
Small-scale pretraining experiments comparing 2019–2025 recipes and datasets found larger compute-efficiency gains from data improvements than from model improvements. [20] This supports the importance of the training material. It does not establish limits on transfer: the study used small models and a limited benchmark suite, and its findings may not hold at frontier scale.
Speed alone would not challenge our thesis. Domain-specific training could become cheap and fast enough to close individual gaps almost as soon as they appear, while new demands continue to create others. Models might increasingly identify missing capabilities, build realistic tasks and feedback, and improve performance with little human direction. Recursive self-improvement could accelerate this further if each round made the next round of AI research faster or more effective. [36] The question is whether this process leaves substantial gaps in dependable capability as the work changes.
The decisive test is whether AI keeps pace with the demands its own success creates. We would update if it reliably handled new activities and expectations with little additional human effort, so that substantial capability gaps no longer persisted. Mastering today's tasks is one thing. Keeping up with a changing task frontier is the stronger test.
A more decisive test would compare a model before and after additional coding training, with the same tools and instructions. Does it become a more reliable executive assistant on unfamiliar tasks without training for those tasks? We would test whether it knows when to interrupt a meeting, which conflicting request to prioritize, and when to ask for clarification. Consistent gains large enough to support dependable delegation would weaken our prediction. Small improvements that leave the agent unreliable would still be consistent with patchwork AGI.
Astra's driving result leaves the origin of the capability unresolved. An initial account described it as emerging seemingly without specialist training, but later acknowledged possible driving data and reinforcement learning. [27] DrivingBench reports completion of a fixed cone course in a real car on the second attempt. Models received driving tools and instructions, with up to three attempts in the same conversation. The reported 100% measures course completion, not a general driving success rate. [28]
Prior training, transfer, and learning during the attempts could all contribute. Even game-to-driving generalization would be transfer. The test is whether learning carries over to a new task without additional task-specific training, not whether the model was trained at all. This result does not isolate the contribution of coding training.
Length, interdependence, and elapsed time are separate demands. A hundred unrelated questions differ from a hundred dependent steps. A weeks-long project also requires tracking changes between actions. Each step can look reasonable while the project drifts away from its goal.
It helps to distinguish bounded tasks from open-ended pursuits. Fixing a specified bug has a stopping point; improving a software product can keep creating value. Task horizon is a separate dimension: a bounded task can require a long sequence of dependent actions, while an open-ended pursuit can consist of many short tasks. If further progress in a profitable pursuit keeps paying off, labs have a reason to keep extending that capability. Uneven development may be economically rational.
Useful negotiation is already possible. In Anthropic's Project Deal, 69 agents struck 186 deals worth just over $4,000 in a marketplace with real people and goods. In separate mixed-model runs, Opus obtained better prices than Haiku, but the difference in reported satisfaction was not statistically significant. [43] Even satisfaction can miss differences in how well an agent represents someone. Task completion, satisfaction, and the user's interests are distinct measures. Project Deal did not isolate the contribution of coding training.
CodeGym turns coding problems into tool-use environments with verifiable outcomes. Training Qwen2.5-32B-Instruct in them improved customer-service accuracy on τ-Bench by 8.7 percentage points on average, with different gains in airline and retail service (Table A1). [24] CodeGym explicitly trains tool use and workflows, so it does not isolate the effect of simply becoming better at coding.
Some patches can share a training foundation. Reusable tool-use skills and domain-specific training can coexist; CodeGym does not imply that every capability must be built independently. Targeted training also helps: SocialRL brought a 4-billion-parameter model into the range of much larger frontier models on held-out negotiation scenarios. Transfer was strongest between structurally similar games. Relevant training can close a gap; this result does not establish that better coding produces better negotiation. [23]
Table A1. CodeGym transfer to customer service
| Training environment | Airline service | Retail service |
|---|---|---|
| CodeGym | +4.4 | +13.0 |
Qwen2.5-32B-Instruct; changes in accuracy, in percentage points. CodeGym trains tool use through coding-derived environments. These results use a different model and measure from the math and social panels in Table 3. [24]
In Eureka, GPT-4 writes and revises reward-function code used to train separate robot policies. Those policies outperform policies trained with human-written rewards on 83% of 29 simulated tasks. The coding model helps another system learn a physical skill; its own weights remain unchanged. Humans still supply the simulator, task description, and metric used to judge success. [34]
AgentGen generates planning environments and tasks, then uses a conventional planner to produce successful trajectories for training language models. For Llama-3.1-8B, training raised success on separate formal-planning benchmark tasks from 0% to 15%. Average success across three other environments rose from 6% to 16%, although success on one of them, Jericho, remained at zero. The gains extend beyond the generated environments, though still within planning tasks. [35]
AI can already help build new capabilities. Whether this reaches unstated priorities, social context, and judgment remains untested. That is an open question, not evidence of failure.
[1] Wikipedia contributors. "Inductive bias." Background on the term used in Figure 1. ↩
[2] OpenAI. "On the Navier–Stokes Millennium Prize Problem." OpenAI. OpenAI describes a proposed proof and Lean formalization; this is distinct from formal Clay prize recognition. ↩
[3] The White House. "President Trump at the United Nations: 'While Others Have Talked, I Have Acted.'" September 22, 2026. Documents the proposed name "super intelligence"; it is a naming reference, not evidence for the essay's capability claim.
[4] Pangram Labs. "How Pangram detects AI-generated content." Describes training a language-model classifier to distinguish human-written and AI-generated text. ↩ a b
[5] Peter Thiel, attributed remark. Secondhand anecdote relayed to Abhay Venkatesh by people who know Thiel. Wording as recalled: "What comes after AI? Nothing, just 20 more years of AI." No public recording or transcript has been verified. ↩
[6] Shubham (@sksq96). "what if this is the jagged frontier?" X, September 22, 2026. Source of the illustration adapted in Figure 2.
[7] Marc Andreessen. "Why Software Is Eating the World." The Wall Street Journal, 2011. Reprinted by Andreessen Horowitz. Argues that software companies are transforming industries; this essay extends that thesis to domain-specific AI automation. ↩
[8] Cognition. "Devin is Now FedRAMP High In-Process" and "Governing AI agents at scale with BlackRock." 2026. Describes deployment support and enterprise context management, permissions, and governance. Illustrates additional work around agent adoption; it does not establish that these requirements will remain unautomated. ↩
[9] Peter Thiel. "The End of the Future." National Review, October 3, 2011. Argues that technological progress has stalled outside of computing; the source of the essay's "Thielian stagnation" framing. ↩
[10] Elad Gil (@eladgil). [Chart comparing the five largest companies' combined market capitalization with U.S. nominal GDP]. X, September 28, 2026. Reports roughly $21.1 trillion against $32.5 trillion. Company valuations are a stock and GDP is an annual flow; this comparison does not measure the digital economy's share of output.
[11] Jacob Thompson. "Evaluating World Models in Embodied Question Answering through Computational Primitives and Difficulty Progressions." Master's thesis presentation, Robotics Institute, Carnegie Mellon University, 2026. Reports weaknesses in spatial memory and state tracking, and substantial improvement from an explicit memory method. Cited from the public thesis-presentation abstract. ↩ a b c d
[12] Zhaofeng Wu et al. "Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks." NAACL, 2024. arXiv:2307.02477. Performance drops when familiar tasks are changed from their default assumptions, such as arithmetic in other bases or code with 1-based indexing, suggesting partial reliance on narrow, non-transferable procedures. ↩
[13] Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman. "Building Machines That Learn and Think Like People." Behavioral and Brain Sciences 40, e253, 2017. Argues that human learning rests on intuitive physics and psychology, causal and compositional models, and learning to learn. ↩ a b
[14] Yonatan Bisk et al. "Experience Grounds Language." EMNLP, 2020. arXiv:2004.10151. Proposes five "World Scopes," from text corpora to perception, embodiment, and social interaction, and argues that text alone cannot supply what the higher scopes give. ↩
[15] YipitData (@yipitdata). [Post on Anthropic customer spending concentration]. X. Reports that the top 1% of customers account for 46% of spending, up from 25% the previous August. The post does not identify the largest customers or break spending down by use case.
[16] Giovanni Cattani. "Nobody is talking seriously about AI demand." X. Distinguishes bounded tasks from open-ended pursuits and proposes that coding, AI research, and trading drive much frontier demand. The demand shares are conjectural; the distinction informs the economic argument here. ↩
[17] Billy Perrigo. "How Meta's $14 Billion Deal Upended the AI Data Industry." Time, June 16, 2025. Reports that each leading AI company spends around $1 billion a year on human data and that data budgets are rising. ↩
[18] The Information, via Dealroom. "Mercor doubles to $2B gross revenue run rate as AI labs buy expert data." June 2026. A gross figure; Mercor pays 60–70% of it to its expert contractors. ↩
[19] Anthropic. "Introducing Claude Opus 5.5." September 22, 2026. On Anthropic's automated behavioral audit, its most comprehensive alignment test, Opus 5.5 is the strongest-performing model it has tested. ↩
[20] Dwarkesh Patel and Jerry Han. "Pretraining progress is mostly coming from data." Dwarkesh Podcast, September 8, 2026. Small-scale pretraining experiments found larger compute-efficiency gains from data improvements than from model changes. The results do not establish limits on transfer at frontier scale. ↩ a b
[21] Chuxuan Hu et al. "Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?" ICLR, 2026. arXiv:2506.19733. Controlled experiments and comparisons across 16 benchmarks find selective transfer from reinforcement post-training; the numerical example above comes from Table 3. ↩
[22] Yu Li et al. "Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning." arXiv preprint arXiv:2507.17512, 2025. Table 9 reports changes in domain-average accuracy for Qwen2.5-7B. Includes positive transfer from puzzle training; peer-reviewed acceptance has not been verified. ↩ a b
[23] Microsoft Research. "From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models." 2026. SocialRL improves negotiation through targeted training and finds uneven transfer between games. It does not isolate transfer from coding training. The full SocialRL paper (arXiv:2608.13787), Table 3, supplies the cross-domain utility results. ↩ a b c
[24] Weihua Du et al. "Generalizable End-to-End Tool-Use RL with Synthetic CodeGym." ICLR, 2026. arXiv:2509.17325. Converts coding problems into interactive tool-use environments and reports transfer to unfamiliar workflows. The 8.7-point gain is an absolute accuracy improvement on τ-Bench; the training explicitly targets tool use. ↩ a b c
[25] Microsoft Research. "SocialReasoning-Bench: Measuring whether AI agents act in users' best interests." 2026. Evaluates calendar coordination and marketplace negotiation, distinguishing task completion from outcome quality and due diligence. Preferences are explicit; long-term relationships are not modeled. ↩
[26] Tomer D. Ullman and Joshua B. Tenenbaum. "Bayesian Models of Conceptual Development: Learning as Building Models of the World." Annual Review of Developmental Psychology 2:533–558, 2020. Frames children's learning as building and revising intuitive theories of the world. ↩ a b
[27] Benjamin Todd (@ben_j_todd). [Post and follow-up on Astra's driving performance]. X, September 23, 2026. Todd initially described the driving ability as seemingly untrained, then acknowledged possible driving data and reinforcement learning. ↩ a b
[28] DrivingBench. "Report." DrivingBench. The reported 100% is completion of a fixed cone course on the second attempt, not a general driving success rate. ↩ a b
[29] Ekin Akyürek et al. "The Surprising Effectiveness of Test-Time Training for Few-Shot Learning." arXiv:2411.07279, 2024; revised 2025. Updating model weights at test time substantially improves performance on ARC; ensembled with program synthesis, the method reaches 61.9%, matching average human performance. ↩
[30] Kevin Ellis et al. "DreamCoder: Bootstrapping Inductive Program Synthesis with Wake-Sleep Library Learning." PLDI, 2021. Grows a library of reusable program abstractions that supports generalization from few examples. ↩
[31] Gabriel Grand et al. "LILO: Learning Interpretable Libraries by Compressing and Documenting Code." ICLR, 2024. arXiv:2310.19791. Combines language models with symbolic compression to learn reusable, documented abstractions. ↩
[32] Jacob Andreas. "Language Models as Agent Models." Findings of EMNLP, 2022. arXiv:2212.01681. Argues that language models can infer and represent some of the beliefs, desires, and intentions of the agents who produce text, though only partially. ↩
[33] Joshua B. Tenenbaum, Charles Kemp, Thomas L. Griffiths, and Noah D. Goodman. "How to Grow a Mind: Statistics, Structure, and Abstraction." Science 331(6022):1279–1285, 2011. Describes how structured, hierarchical knowledge lets people generalize far beyond sparse data. ↩
[34] Yecheng Jason Ma et al. "Eureka: Human-Level Reward Design via Coding Large Language Models." ICLR, 2024. A coding model generates and revises rewards used to train separate robot policies. Humans supply the simulator, task description, and evaluation metric. ↩ a b
[35] Mengkang Hu et al. "AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation." KDD, 2025. Generated planning environments supply training trajectories. Tables 1 and 3 report improvements on separate evaluation tasks, including tasks outside the training formalism. ↩ a b
[36] Anthropic. "When AI builds itself." Anthropic Institute. Reports AI-assisted gains in AI development while identifying gaps in research judgment; full recursive self-improvement has not been achieved. ↩ a b
[37] Abhay Venkatesh. "The Amplification of Software." 2026. Argues that cheaper implementation can amplify advantages in judgment, taste, speed, and distribution. A related argument about competition, rather than empirical evidence of limits on AI capability. ↩ a b
[38] Andrew Ross Sorkin (@andrewrsorkin). [Report on Instinct's funding and team size]. X / DealBook, September 28, 2026. Reports a $10 billion valuation and 14 employees. An example of a highly valued software company with a small team; valuation is distinct from revenue. ↩
[39] Abhay Venkatesh. "The AI Economy." 2026. Argues that products built around a moving frontier can be absorbed, displaced, or commoditized as labs expand their capabilities and product surfaces. ↩
[40] Chris Paik. "The End of Software." 2024. Argues that falling creation costs will proliferate software and disrupt traditional software businesses. The claim concerns software economics, not the exhaustion of software work. ↩
[41] bouvard (@bouvard38829538). "The ultimate substrate of machine intelligence is tacit knowing, which is inexhaustible and replenished by every attempt to map it." X, September 22, 2026. ↩
[42] Dennis Bouvard. "Infra-Humaning." 2024. Describes how attempts to define and automate human characteristics reveal further distinctions and create new human activities. A philosophical basis for an open-ended human–machine relationship, rather than empirical proof that it continues forever. ↩
[43] Anthropic. "Project Deal: our Claude-run marketplace experiment." 2026. Agents negotiated transactions involving real people and goods. Mixed-model comparisons favored Opus over Haiku on several outcome measures; the experiment does not identify which training produced the advantage. ↩
[44] Google Research. "Titans + MIRAS: Helping AI have long-term memory." Describes architectures with neural memory that updates at test time, illustrating research on learning and retention beyond a fixed context window. ↩
[45] CompaniesMarketCap. "Largest Companies by Market Capitalization." Accessed October 1, 2026. The five most valuable companies were Nvidia, Apple, Alphabet, Microsoft, and Amazon; eight of the top ten were computing companies. ↩
[46] Giovanni Sala and Fernand Gobet. "Does Far Transfer Exist? Negative Evidence From Chess, Music, and Working Memory Training." Current Directions in Psychological Science 26(6):515–520, 2017. Finds that far transfer from these trainings rarely occurs, and that effects shrink as studies are better controlled. ↩
[47] Federico Cassano et al. "MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation." arXiv:2208.08227, 2022. Translates two Python code-generation benchmarks into 18 other programming languages. Performance tracks how common a language is, though some rare languages do as well as popular ones. ↩
[48] Member of technical staff at a leading synthetic-data company that sells training data to all frontier labs. Conversation with the authors, 2026. Describes labs commissioning task data built to resemble new benchmarks. ↩
[50] Geoff Brumfiel. "AI solved one of math's hardest problems. Humanity learned nothing (so far)." NPR, September 22, 2026. Mathematicians, including James Maynard, find it hard to extract human understanding from OpenAI's 166-page, Lean-checked Navier–Stokes proof. ↩
[51] Vinayshekhar Bannihatti Kumar, Disha Makhija, Manoj Ghuhan Arivazhagan, and Rashmi Gangadharaiah. "Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language." COLM, 2026. arXiv:2605.15607. Given the rules of an unseen language, frontier models including GPT-5.4 solve far fewer problems than in Python, though they often pick the same algorithm. ↩ a b
[52] Alex Bie, Travis Dick, Alex Kulesza, Prabhakar Raghavan, Vinod Raman, and Sergei Vassilvitskii. "AI-rithmetic." arXiv:2602.10416, 2026. GPT-5, Claude Opus 4.1, and Gemini 2.5 Pro get worse at adding two integers as the numbers get longer; most errors are misaligned digits or missed carries. ↩