Gemini Spark and Google's AI Agent Ecosystem: The Platform Bet

Contents

Google does not have one AI agent. It has a family of them, and in 2026 it gave that family a centre. On 19 May 2026, at I/O, Google introduced Gemini Spark, a personal agent that runs on Google’s cloud, keeps working after you close your laptop, and reaches into Gmail, Docs, Sheets, Calendar, Chrome and a growing list of third party services. Spark is the renamed and rebuilt successor to Gemini Agent, which launched as an experiment for US AI Ultra subscribers alongside Gemini 3 on 18 November 2025. Since May it has moved from trusted testers to a beta for AI Pro and Ultra subscribers in more than 160 countries, learned to drive your own Chrome browser, and been wired into Gemini Live so you can delegate by voice.

Around Spark sits everything else: agentic booking and checkout in AI Mode, auto browse in Chrome, Gemini UI automation on Android, Workspace Studio and Gemini Enterprise for companies, and a developer layer of computer use models, Managed Agents and open protocols. This article treats that ecosystem as the product. Google’s advantage is not the smartest single Gemini AI agent. It is the number of places where an agent can already act for a user without asking anyone new for permission, and that creates design and trust problems a standalone agent does not face.

Why Google’s AI agent story is an ecosystem story

The industry’s path reads like a staircase. Chatbots answered questions. Copilots drafted inside a document. Task agents such as Deep Research took a goal and returned a report. Computer-use agents started clicking through real websites. The current step is persistent agents: software that accepts a standing instruction, runs on someone else’s computer for hours or days, and returns only when it needs a decision.

Every major lab is climbing, but from different floors. OpenAI starts from a chat product with enormous reach. Anthropic starts from developers and knowledge workers and brings the agent to tools through connectors and a Chrome extension. Perplexity starts from research. Google starts from the places where work and errands already happen: the dominant browser, the dominant mobile operating system, a search engine whose AI Mode had a billion monthly users by Sundar Pichai’s count at I/O, a productivity suite, a payments wallet, a local business graph and a cloud platform.

So the interesting question is not whether Spark can book a restaurant. Plenty of agents can. It is what happens when one agent identity, memory and permission model appears in your phone’s status bar, your browser’s side panel, a search result and your company’s documents, and whether users can tell where one ends and the next begins.

What Google launched: the Gemini agent stack in September 2026

This is a direction rather than one product, so each piece is labelled by its status on 30 September 2026.

Gemini Spark, the personal agent at the centre

Timeline. Announced 19 May 2026. Trusted testers that week, then a beta for US AI Ultra subscribers. US AI Pro subscribers on 24 July; AI Pro in more than 160 additional countries from 30 July.

What it is. A cloud-based agent that Google says keeps working “even when you close your laptop or lock your phone.” It uses the Antigravity agent harness and launched on Gemini 3.5 Flash; Google’s July country posts say Gemini 3.6 Flash.

Building blocks. Google’s help centre defines Tasks (the goal), Schedules (when it runs, by time or event) and Skills (reusable instructions, invoked with a slash). Up to 15 tasks can run at once.

Integrations. Native Gmail, Calendar, Drive, Docs, Sheets, Slides, Keep, Tasks, Contacts, Photos, YouTube and Maps, plus Model Context Protocol (MCP) connections; Canva, OpenTable and Instacart were the launch partners. Since 30 July, Spark can use your Chrome with logged-in accounts and saved passwords, with permission (US only).

Access. AI Pro or Ultra, 18 or over, personal accounts (the help centre excludes work and school accounts, though Google’s overview page mentions select business users), not available in the EEA, Switzerland, the UK or Nigeria. Available on the Gemini mobile apps, Mac and the web. Status: beta, with compute-based usage limits.

Chrome: auto browse

Auto browse began rolling out on desktop to US AI Pro and Ultra subscribers in late January 2026. It fills forms, researches travel, manages subscriptions and can use Chrome’s password manager with permission, pausing before purchases or social posts. Gemini in Chrome reached all US Android users by 18 August 2026, with auto browse limited to Pro and Ultra. WebMCP, a proposed standard for sites to expose structured tools to browser agents, is in an origin trial from Chrome 149. Status: auto browse released to paying US users; WebMCP experimental.

Search: AI Mode learns to transact

Agentic restaurant booking launched for US Ultra subscribers in August 2025 and went global without a subscription on 10 April 2026. Agentic checkout, which buys a tracked item when it hits your price, and “Let Google Call” for store inventory arrived on 13 November 2025. On 11 January 2026 Google introduced the Universal Commerce Protocol (UCP), built with Shopify, Etsy, Wayfair, Target and Walmart, plus a buy button in AI Mode and Gemini. A cross-retailer Universal Cart followed on 20 May. Agentic hotel booking reached US English users around 1 September 2026, with hotels and travel sites as merchant of record. Information agents that monitor the web launched for subscribers in June; on 28 September Search VP Robby Stein announced monitoring for all users globally. Status: released, mostly US first.

Android: from app automation to a home for agents

On 25 February 2026 Google announced Gemini UI automation, which operates apps for you, as a beta on the Galaxy S26 and select Pixel 10 phones in the US and Korea, starting with food delivery, grocery and rideshare. The same post introduced AppFunctions, which lets apps expose callable functions to agents; it is in experimental preview. Android Halo, a status bar surface where agents report progress, was teased at I/O and Google says it is coming later this year. Status: UI automation limited beta; Halo announced, not shipped.

Workspace, Enterprise and developers

Workspace Studio, a no-code agent builder across Gmail, Drive and Chat, launched in December 2025 and added Gmail reply and Drive steps in September 2026. Gemini Enterprise launched on 9 October 2025; at Cloud Next on 22 April 2026 Vertex AI became the Gemini Enterprise Agent Platform, with agent identity, a gateway, a registry and observability. For developers, the Gemini 2.5 Computer Use model entered public preview on 7 October 2025, computer use came to Gemini 3.5 Flash on 24 June 2026, and Managed Agents, which provision an agent with a remote sandbox in one API call, launched in preview on 19 May alongside Antigravity 2.0. Jules, the asynchronous coding agent, left beta on 6 August 2025.

Names that changed

The brief’s topic, “Google Gemini’s agentic ecosystem,” is accurate: no single product carries that name. Agent Mode (previewed May 2025) became Gemini Agent (November 2025), which became Gemini Spark (May 2026). Project Mariner, the browser agent prototype from December 2024, was shut down on 4 May 2026 according to multiple reports, its technology absorbed into Gemini, Chrome and the API. Project Astra was a research programme whose camera, screen sharing and memory work surfaced in Gemini Live, which since 26 August can hand long jobs to Spark. Models moved monthly: Gemini 3.5 Flash (19 May), 3.6 Flash (21 July), 3.7 Flash (13 August), 3.8 Flash (2 September). Gemini 3.5 Pro, promised for June, had not shipped as of the latest reporting we found.

Why the Gemini AI agent strategy matters

What changed. Google Assistant ran fixed intents. Early Gemini could answer questions about your inbox. Spark accepts a goal, plans, runs remotely, crosses services and resumes on a schedule. The unit of interaction is no longer a turn. It is a task with a lifespan.

New behaviour. Standing instructions. “Every Friday, collect invoices from my inbox into a sheet and flag anything over budget” is something an agent can own indefinitely, which shifts the user’s job from doing to supervising.

Why now. Flash-class models became fast and cheap enough for long agent loops (Google claims 3.5 Flash is four times faster than other frontier models). Computer use matured into a preview API. And rivals made it urgent: Anthropic turned its Chrome side panel into a Cowork session on 12 August, and OpenAI launched dots, its always-on agents, on 29 September. Moving Spark from Ultra to AI Pro within ten weeks suggests Google judged reach more valuable than exclusivity.

The category. The product appears designed to make Google the default place where delegation happens, as it became the default place where searching happens. If people hand errands to an agent, the one that already holds their email, calendar, passwords, payment methods and location starts with a structural lead.

How Gemini’s agents work

In plain terms, Spark is an assistant with a desk in Google’s data centre. You give it a job; it writes a plan, opens the tools it needs, works through the steps, and asks you when it needs a signature.

Tools before clicks. Spark prefers structured routes. Creating a Calendar event is an API call, not a simulated click. Third party services arrive through MCP. Only when no structured path exists does it browse.

Two browsers. The help centre describes remote browsing, a separate browser on Google’s infrastructure that keeps going when your laptop sleeps (its cookies persist until you delete them), and local browsing in your own Chrome with your signed-in sessions, which is more capable but needs your device online.

The computer use loop. The model receives the task, a screenshot and recent actions, proposes a click or keystroke, or requests confirmation for sensitive steps. The environment executes and returns a new screenshot. Google’s developer version adds a per-step safety service that screens each action before it runs.

Protocols for commerce. UCP gives merchants a standard way to expose catalogues, carts and checkout to agents, and Google says it is compatible with Agent2Agent (A2A), the Agent Payments Protocol and MCP. On a supporting merchant, “buy this” becomes a structured Google Pay transaction rather than an agent typing into a form.

Memory. Spark draws on Personal Intelligence, Google’s opt-in feature that connects past chats and apps such as Gmail and Photos, plus saved instructions. Keep Activity must be on, a reminder that this memory is stored.

The user experience of a Gemini AI agent

This is a reading of Google’s documentation and published coverage, not hands-on testing.

Interaction model. Text and, increasingly, voice. Spark has its own section of the Gemini app, separate from chat; Gemini Live can pass it spoken tasks.

Task initiation. Spread across surfaces: the Spark tab, a Live conversation, a Chrome Skill, an AI Mode monitoring suggestion, a Workspace Studio flow. Each is low friction. Together they make one mental model of “my agent” harder to form.

Delegation. Unusually broad. Workspace data plus Chrome credentials is the key difference from standalone agents, which need accounts connected one by one.

Visibility. A work panel shows completed, current and planned steps and the files touched. Treating the plan as a visible object, rather than a spinner, is a strong pattern. Android’s live view and the coming Halo put status in the system chrome, where persistent processes belong.

Control. Users can stop tasks, pause schedules and “take control” of a browser; for passwords the agent hands the browser back. Documentation says little about editing a running plan.

Trust. The language is consistent across surfaces: the agent asks before spending money, sending messages, modifying data or submitting forms. The weakness is that confirmations carry the whole load, and consent prompts decay into reflexive taps.

Feedback and completion. Notifications, the work panel and Daily Brief. Alerts from Search monitoring, Spark and Chrome arrive through different channels, which risks fragmentation. For scheduled tasks, “done” becomes a recurring heartbeat.

Errors. Tasks return to the user when manual action is needed. Google warns that “Gemini can make mistakes and do unexpected things” and that offline schedules may complete unintended actions before you can intervene. Recovery is the weakest link.

Memory and permissions. App connections are off by default; Chrome asks once, then confirms per task. The gap is granularity: “access Gmail” is a very large grant for a job that only needs receipts, and users cannot easily see which memories shaped a task.

Real-world use cases for Google’s AI agents

Subscription audit. Task: find recurring charges. Agent: Spark parses Gmail receipts into a sheet. Human: decides what to cancel. Result: an evening’s chore done in minutes.

School deadlines. Task: never miss a form. Agent: a weekly schedule scans school emails and adds dates to Calendar. Human: approves early runs, then spot checks. Result: a standing process replaces a habit.

Weekend trip. Task: a hotel under budget. Agent: AI Mode compares options and books through a partner; Spark drafts the itinerary. Human: chooses and confirms payment. Result: discovery to booking in one flow.

Price-triggered purchase. Task: buy an item below a set price. Agent: agentic checkout buys through Google Pay when the price drops. Human: sets the ceiling and shipping in advance. Result: a purchase at the right moment, unwatched.

Competitor tracking (marketing). Task: watch three rivals’ launches and pricing. Agent: a Search monitor or Spark schedule summarises weekly into a Doc. Human: judges significance. Result: a living brief.

Research synthesis (design). Task: cluster interview notes in Drive. Agent: Spark drafts themes and a slide outline. Human: checks themes against raw notes. Result: a first draft in an hour, judgment still human.

Interview scheduling (recruiting). Task: book panels. Agent: a Workspace Studio flow reads replies, proposes slots, drafts confirmations. Human: approves outgoing messages. Result: less back and forth.

Vendor portal chores (operations). Task: download monthly reports from a portal with no API. Agent: auto browse signs in with saved credentials. Human: handles two-factor prompts, checks files. Result: a recurring manual job automated.

Checkout regression testing (development). Task: verify a web checkout each release. Agent: a Managed Agent with computer use runs the flow. Human: triages failures. Result: UI coverage without brittle scripts, with nondeterminism to manage.

Becoming agent-purchasable (e-commerce). Task: sell inside AI Mode. Agent: other people’s agents buy; the merchant implements UCP and conversational product attributes. Human: owns catalogue quality and returns. Result: the store becomes a destination for agents.

What changes for product designers

The Gemini ecosystem sharpens one question: what happens when the agent reaches your product through Chrome, Search, Android or a protocol rather than your interface? Our broader take is in designing products for AI agents.

Your UI becomes one of several front doors. A booking may start in AI Mode and a reorder may run through auto browse. Design the flows agents traverse with the care you give the human path.

Dashboards shift from operating to overseeing. If agents do the routine work, the dashboard explains what happened, what is pending and what needs a decision. The ideas in SaaS dashboard design still hold, but the key widget may be an activity log with approvals.

Forms must be legible to machines. Clear labels, predictable validation and standard controls help computer use models as much as people. The basics of form design are now agent readiness work.

State beats navigation. Agents do not browse menus. They need stable URLs and actions for “change delivery address.”

Show autonomous activity honestly. Halo’s pattern, a small persistent indicator that expands into detail, generalises well: show who acted, on whose authority, and how to undo it.

Scope permissions to jobs. Ask for “read receipts from the last 90 days to build this sheet,” and let the grant expire with the task.

Design recovery first. Make agent-triggerable actions reversible, with undo windows, and mark irreversible ones in both the interface and the action schema.

Conversation and GUI work together. Google’s generative UI in Search points to the pattern: conversation to state intent, interface to review and confirm. See our notes on LLM UX patterns.

What changes for developers

Expose structured actions. Google now offers at least four agent channels: MCP (used by Spark), WebMCP in Chrome, AppFunctions on Android and UCP for commerce. Build one internal action layer and project it into each.

Authentication moves to the centre. An agent using Chrome passwords acts as the user. Prefer delegated, scoped tokens, detect agent traffic, and support step-up confirmation instead of blocking automation outright.

Speak agent to agent. The Linux Foundation reported A2A at a 1.0 specification with more than 150 supporting organisations in April 2026, with support in Azure AI Foundry and Amazon Bedrock AgentCore. An agent card and A2A endpoint may matter as much as a REST API for enterprise buyers.

Plan for agent-generated UI. Google’s September 2026 Android series introduced A2UI renderers for Jetpack Compose, letting agents describe components for the client to render. Decide which components an agent may compose.

Reliability and safety. Agents retry, so make writes idempotent. Log the actor type on every action. Treat content agents will read as untrusted input, because prompt injection is an input validation problem.

The competitive picture

Documented capabilities as of 30 September 2026. Not a ranking; each product changes monthly. For the other side of the comparison, see our analysis of OpenAI’s always-on dots agents and of Anthropic’s Claude computer use and Cowork.

Dimension Google Gemini Spark and surfaces OpenAI dots Anthropic Claude (Cowork) Perplexity Computer
Launch Gemini Agent Nov 2025; Spark 19 May 2026 29 Sep 2026, DevDay Cowork preview 12 Jan 2026; merged into Claude 16 Sep 2026 Computer 25 Feb 2026 (Max); Personal Computer on Mac for Pro and Max 7 May 2026
Computer use Remote browser, local Chrome; Android beta Own cloud computer and browser Clicks, types and fills forms through Claude in Chrome Cloud sandboxes with browser; local Mac files and apps via Personal Computer
Background execution Cloud, schedules, 15 concurrent tasks Runs continuously on its cloud computer Cloud tasks and schedules; local files need the desktop app Long-running cloud tasks and schedules
Integrations Native Workspace, Maps, YouTube, MCP 4,000+ apps via ChatGPT plugins; messages in Slack and Teams Connectors and skills Hundreds of connectors; 400+ apps on Enterprise
Commerce UCP checkout, AI Mode booking, Google Pay Not a documented focus Not a documented focus Not a documented focus
User control Confirms sends, buys, submissions; take control Messages its owner when a decision is needed Asks approval by default; autonomy configurable Inline confirmations and sign-in prompts
Availability AI Pro and Ultra, 18+, not EEA or UK Pro and Business Premium, rolling out Paid plans; merged app rolling out to Pro and Max first Pro, Max and Enterprise, credit metered

Two observations follow from the evidence rather than the marketing. Google is the only company here with first party distribution in the browser, a phone operating system and search at once. And that breadth carries a constraint the others face less: Spark’s help centre still excludes work and school accounts (Google’s overview mentions only select business users), so for many professionals the personal agent and the work agent are still different agents.

The business model behind Google’s agent ecosystem

Subscriptions are the gate. At I/O Google added a $100 a month Ultra tier beside the $200 tier, then opened Spark to AI Pro (reported at $20 a month in the US) in July. That speed suggests Spark is meant to be a reason to subscribe.

Agent compute is different. A chat turn is short; an agent task can hold a virtual machine and make dozens of model calls. Google’s “compute-based usage limits” wording hints the constraint is infrastructure. Owning TPUs, the cloud and the models is a cost advantage, and the Flash line is aimed explicitly at long-horizon agentic work.

Commerce is the long game. UCP, Universal Cart and partner booking put Google in the transaction path without becoming a retailer. One plausible implication is that agent-completed commerce could offset some value if users click fewer links. Google has not detailed how it monetises agentic checkout.

Enterprise. Gemini Enterprise sells by seat and consumption; Google reported 40 percent quarter-on-quarter growth in its paid monthly active users in Q1 2026.

Limitations and risks

Incorrect actions (vendor acknowledged). Google says Spark can “make mistakes and do unexpected things,” and warned that Gemini Agent could “expose data unintentionally.”

Prompt injection (documented risk). Google claims protections but says user supervision is the most important safeguard. We found no independent evaluation of Spark’s defences.

Offline autonomy (documented). Schedules act while you are away. That is the feature and the risk.

Data concentration (structural). The integration that makes the agent useful concentrates email, location, payments and browsing in one account, and remote browser cookies persist until deleted.

Regional gaps (documented). No Spark in the EEA, UK, Switzerland or Nigeria; auto browse and several Search agent features are US only.

Model volatility (documented). Monthly releases change behaviour, and the 3.5 Pro delay shows roadmaps slip. Businesses pay re-validation costs.

Naming confusion (observed). Agent Mode, Gemini Agent, Spark, Mariner, auto browse, Studio flows, Enterprise agents, information agents. Accountability depends on users knowing which agent did what.

Lock-in and over-automation (plausible, not yet demonstrated). An agent that works best with Google services nudges users toward them, and one-sentence buying makes impulse and error cheaper.

What Google’s agents reveal about the future of AI agents

From answering to acting: shipping. Search books tables and hotels and completes checkouts today.

From sessions to continuity: supported, within limits. Schedules, monitoring and Personal Intelligence assume persistence, but continuity lives inside Google’s account boundary.

From apps to agents: partly. UI automation operates other companies’ apps, yet AppFunctions, WebMCP and UCP keep apps as the source of truth. Apps become capability providers, not relics.

From interfaces to capabilities: the strongest signal. Nearly every Google developer announcement in 2026 is about machine-readable actions.

From prompts to outcomes: forward-looking. Spark accepts goals, but its reliance on confirmations shows users still approve the path.

What founders should pay attention to

Customers may arrive as agents. Being callable through UCP, MCP or WebMCP could matter the way SEO did. Our piece on AI search and SEO covers discovery.

Vertical agents still have room. Domain rules, regulated workflows and specialised data beat generalists in narrow lanes, and can still plug in through A2A or Gemini Enterprise’s catalogue of partner agents.

Cross-vendor oversight is open. Google builds identity and observability for its own platform. Auditing what any agent did across ecosystems remains an opportunity.

Avoid features Google will absorb. Mariner’s fate shows capabilities near Google’s core surfaces tend to be folded into them.

What designers should start doing now

  • Map agent entry points: browser, Android, Search, protocols, API.
  • Write an action inventory with inputs, outputs and reversibility. It doubles as your agent schema.
  • Design approval states and their copy; UX writing and microcopy matters more when a prompt guards a payment.
  • Log actors distinctly: human, agent for human, agent on schedule.
  • Make undo real and mark what cannot be undone.
  • Test key flows with a computer use agent and watch where it stalls.
  • Write status messages that survive being shown in a status bar or side panel you do not control.
  • Earn trust visibly, following the principles in designing AI features people trust.

Final takeaway: what fundamentally changed

In 2026 Google stopped treating agents as experiments at the edge of its products and made delegation a property of the platform. Spark gives users one agent to talk to. Chrome, Android, Search and Workspace give it places to act. UCP, A2A, WebMCP and AppFunctions give other companies a way to be acted upon. No competitor can currently assemble that combination from first party parts.

That does not make Google’s agent the best at any single task, and its own documentation is candid about mistakes, supervision and regional limits. The shift is architectural. When the browser, the phone and the search box carry the same agent, the question for everyone else changes from “should we build an agent?” to “what should another company’s agent be allowed to do in our product, and how will our users see it happen?”

Watch four things next: whether Android Halo ships with third party agents and real isolation, whether Spark reaches work accounts, whether WebMCP and UCP spread beyond launch partners, and whether confirmation-based trust holds as tasks grow longer.

Building a product that people and their agents both need to operate? hello@beconfidency.agency, we design agent ready products with clear actions, visible autonomy and recoverable workflows.

This thinking shapes our AI design and integration service.

Sources

Next project

Have an ideaworth raising?