Desk Trials

SOURCE READER · 14–15 SEP 2026

Assistants, copilots and agents

Read access, write access and the work of review.

Original research edition

This is the supplied September 2026 report, including its caveats and source links. Its prices, vendor events and verification claims have not all been independently rechecked for this preview. Consult the original correction log alongside the chapters.

Market Definition: What Counts as an "AI Productivity Product" in 2026

The category boundaries in this market are genuinely unstable — vendors relabel the same product from "copilot" to "agent" as marketing fashion shifts, and the four big platforms keep absorbing what were standalone categories a year earlier. The definitions below reflect how the terms are actually used in 2025–2026, with the caveat that almost every real product now straddles at least two of them.

The Core Taxonomy

AI assistants are conversational, general-purpose systems a person addresses directly: you ask, it answers or drafts, and you take the output elsewhere. ChatGPT (roughly 700–800 million weekly users by early-to-mid 2026, per widely circulated OpenAI figures) and Claude are the archetypes. The assistant is destination software — you go to it. The boundary blurs because the leading assistants have absorbed agent capabilities (ChatGPT's agent mode, Claude's computer use and file creation), so "assistant" now names the interface more than the capability level.

AI copilots embed the model inside an existing tool, working alongside a human who stays in the driver's seat, in the tool's own context. GitHub Copilot (autocomplete and chat inside the IDE) and Microsoft 365 Copilot (drafting inside Word, summarizing inside Teams and Outlook) defined the pattern. The distinguishing feature is that a copilot sees what you are working on without being told. The blur: GitHub Copilot itself now ships an autonomous "coding agent" mode that takes an issue and opens a pull request unattended — the flagship copilot brand has become an agent product.

AI agents pursue a goal over multiple steps with limited supervision: planning, using tools, checking results, iterating. 2025–2026 produced four important sub-species. Computer-use agents operate a browser or full GUI the way a person would — OpenAI's Operator (folded into ChatGPT agent mode in mid-2025) and Anthropic's Claude computer use are the reference examples; reliability improved sharply into 2026 but these remain the least dependable agent class on long tasks. Deep research agents (OpenAI Deep Research, Gemini Deep Research, Claude's research mode) browse for tens of minutes and return cited reports — arguably the first agent category with mass consumer adoption. Coding agents are the commercial breakout: Claude Code, OpenAI Codex, Cursor's agent mode, and Devin run multi-file changes, tests, and PRs with minutes-to-hours of autonomy, and by 2026 they are a primary interface for many professional developers (JetBrains' 2026 research tracks rapid adoption). Finally, agent protocols matter because they turned agents from demos into infrastructure: Anthropic's Model Context Protocol (MCP, open-sourced November 2024) became the de facto standard for connecting agents to tools and data after OpenAI, Google, and Microsoft all adopted it in 2025, with Google's Agent2Agent (A2A, donated to the Linux Foundation) covering agent-to-agent communication. The assistant/copilot/agent distinction is best read as a spectrum of autonomy per invocation, not three product types — most 2026 products expose a slider across all three.

AI-native applications were architected around the model from day one; remove the AI and no product remains. Cursor (an editor rebuilt around AI pair-programming, one of the fastest-growing software companies ever by revenue) and Perplexity (search rebuilt as answer-plus-citations) are clean examples. AI-enhanced software is the inverse: an established product that bolted on AI features — Notion AI inside Notion, Zoom AI Companion, Adobe Firefly inside Creative Cloud. The boundary blurs when incumbents rebuild rather than bolt on: Grammarly acquired Coda and the email client Superhuman, then in October 2025 rebranded the whole company as Superhuman, repositioning a 15-year-old grammar checker as an "AI-native productivity suite" with a cross-app agent (Superhuman Go). Whether that constitutes AI-native or aggressive re-labeling is exactly the kind of judgment buyers now have to make.

Automation tools predate the AI wave: deterministic, trigger-action systems like Zapier and UiPath (RPA). Their defining trait was reliability through rigidity — they broke when the interface changed but never improvised. Both have now injected LLM steps (Zapier Agents, UiPath's agentic automation), converging on the agent category from the opposite direction: automation tools are adding judgment, agents are adding reliability, and the two meet in the middle as "workflow platforms with AI steps."

Productivity apps (docs, tasks, calendars, email — Notion, Todoist, Superhuman Mail) and knowledge-management tools (Obsidian, Confluence, and AI-first entrants like Mem or Notion's Q&A) are user-facing categories defined by job-to-be-done rather than technology. The material change since 2024 is that retrieval-augmented search over your own notes and company wiki went from differentiator to table stakes; Glean built a company on enterprise search-plus-assistant and now competes directly with Microsoft Copilot's and Gemini's built-in versions of the same thing.

Workflow platforms (n8n, Make, Airtable, ServiceNow) orchestrate multi-step business processes across systems; they are the natural substrate for deploying agents, and n8n's 2025–2026 growth came almost entirely from being the place non-developers wire LLM calls into processes. Personal-assistant apps — consumer schedulers and life managers like Reclaim.ai or Motion, plus the device-level assistants (Gemini on Android replacing Google Assistant) — remain the weakest standalone category, because the platform assistants absorb their features fastest; the AI-hardware attempts (Humane AI Pin, Rabbit r1) failed outright, with Humane shut down and sold for parts to HP in early 2025.

Enterprise AI platforms sell the substrate for building and governing AI inside large organizations: Microsoft Copilot Studio and Azure AI Foundry, Google Vertex AI, AWS Bedrock, Palantir AIP, Salesforce Agentforce, plus OpenAI's and Anthropic's enterprise tiers. The buyer is IT, and the pitch is governance, security, and model choice rather than any single workflow. Vertical AI applications go the opposite direction — deep in one regulated domain: Harvey in legal (reported at an $8–11B valuation in 2026 with revenue variously reported between roughly $190M and $350M ARR — figures differ across sources and could not be verified exactly) and Abridge in clinical documentation (ambient note-taking from doctor-patient conversations, one of health systems' most-adopted AI purchases) are the standard examples. Vertical AI is where 2025–2026 venture conviction concentrated, on the thesis that workflow depth plus proprietary domain data resists platform absorption better than horizontal features do.

What Separates Genuinely Useful from "Chatbot Bolted On"

The consistent lesson of the 2024–2026 evidence is that value tracks workflow integration depth and context access, not model quality. A chat panel that cannot see your data, write into your systems of record, or act without copy-paste adds a step instead of removing one. MIT's much-cited mid-2025 "GenAI Divide" report found roughly 95% of enterprise generative-AI pilots produced no measurable P&L impact — with the failures concentrated in generic, un-integrated tools that don't retain context or fit existing workflows, while the successful 5% were narrow, deeply embedded, and often bought rather than built. That report's methodology drew criticism (small interview base, loose definition of "failure"), so treat the 95% as directional, not precise.

The measured productivity record is genuinely mixed, and any honest market definition should carry all of it. On the positive side: Brynjolfsson, Li, and Raymond's study of ~5,000 customer-support agents (NBER, later published in the Quarterly Journal of Economics, 2025) found an AI assistant raised issues resolved per hour by about 14–15% on average — 30%+ for novice workers, near zero for the most skilled. A large multi-firm RCT of GitHub Copilot (Cui et al., ~4,900 developers at Microsoft, Accenture, and a Fortune 100 firm) found a ~26% increase in completed tasks, again concentrated among less-experienced developers. Microsoft's own RCT of M365 Copilot across 6,000+ workers in 56 firms (Dillon et al., 2025) found real but modest effects: regular users cut email-reading time ~30 minutes a week and finished documents about a day faster, while meeting time didn't move at all.

On the other side sits METR's July 2025 randomized trial: 16 experienced open-source developers working on their own mature codebases were 19% slower with early-2025 AI tools — while believing afterward that AI had sped them up by 20%. The sample was small and deliberately adversarial to AI's strengths (experts, unfamiliar-tooling learning curves, high-context repos), but it established the most important consumer-facing fact in this market: perceived time saved is not measured time saved, and self-reported productivity gains should be discounted heavily. The mechanism METR identified — time shifts from writing to prompting, waiting, and reviewing — generalizes: the review burden is the hidden cost line of every AI product, and output quality only creates value when it exceeds the cost of checking it. This is also why hallucination remains the canonical failure mode for text products and why "shallow wrapper" products (thin UI over an API, no proprietary context, no workflow lock-in) churned out of the market in waves once ChatGPT or Copilot shipped the same feature natively.

Pricing-to-value is the practical test. Seat-based AI add-ons ($20–30/user/month for M365 Copilot or similar) require roughly an hour of genuinely saved time per month to break even at typical loaded salaries — a low bar the RCT evidence suggests most regular users clear, but which the MIT pilot data suggests many deployments don't, because usage never becomes habitual. The 2026 trend is a shift from per-seat to usage- and outcome-based pricing (per agent task, per resolved ticket), which forces the value question into the invoice. Durable products share observable signals: they own a system of record or proprietary data loop, they sit in the workflow rather than beside it, retention and usage depth (not sign-ups) grow, and their job is one the platform vendors are structurally unwilling to do (regulated verticals, cross-vendor neutrality, on-prem deployment). Features get absorbed; workflows with data gravity survive.

Market Size and Platform Dynamics

Headline numbers are large and moving fast. Gartner's May 2026 forecast puts total worldwide AI spending at $2.59 trillion in 2026 (up 47% from $1.76T in 2025), though over 45% of that is infrastructure — chips and AI-optimized servers — not software anyone uses. The relevant slice for this report is AI software: $282.9B in 2025, forecast to reach $453.2B in 2026 and $638.4B in 2027, with generative-AI model spending roughly doubling to $32.6B in 2026. Productivity/collaboration SaaS as a whole is far smaller and slower — business-productivity software is generally estimated in the ~$80–110B range with high-single-digit growth (Mordor Intelligence and similar analyst houses; no single authoritative figure could be verified, and estimates vary widely by category definition). The strategic point is the ratio: AI software spending now dwarfs the traditional productivity-software market it is remaking, and Gartner notes most 2026 spending still comes from vendors and hyperscalers themselves rather than enterprise ROI — a fact that fuels ongoing bubble debate.

Platform dynamics dominate everything above. OpenAI (ChatGPT apps SDK and AgentKit from DevDay 2025, agent mode, Deep Research), Microsoft (Copilot everywhere, agents in M365), Google (Gemini across Workspace and Android), and Anthropic (Claude Code, MCP, enterprise Claude) each expanded from model vendor into direct competitor with their own ecosystems. Each major model release triggers a "kill zone" discussion: OpenAI's DevDay 2025 announcements alone overlapped with standalone startups in agents, search, meeting notes, and app-building. The counter-moves that have worked are the ones the taxonomy predicts: go vertical (Harvey, Abridge), go deep into workflow and data (Cursor, Glean), consolidate into suites (Grammarly/Coda/Superhuman), or become infrastructure the platforms adopt rather than replace (MCP itself). For buyers, the practical implication is to assume any horizontal, single-feature AI product has an expected lifespan measured in platform release cycles.

Unverified / flagged items: Harvey's exact revenue and valuation (sources conflict: ~$190M vs ~$350M ARR; $8B vs $11B); precise 2026 ChatGPT weekly-user counts (secondary aggregators only); the single "productivity SaaS market size" figure (analyst estimates diverge by definition); MIT's 95% figure (methodology contested).

Sources