← Articles
Read in another language
Technology

AI Browsers Are Here: What They Do and What They Risk

Jayden

Maintains the wage calculators and public-data regional information at 생활데이터랩, and analyzes technology, industry, and policy issues.

Published

Key points

  • Within a single stretch of 2025, ChatGPT Atlas (Oct 21), Edge's expanded Copilot Mode (Oct 23), Gemini in Chrome (Sep 18) and a free Comet (Oct 2) turned the browser itself into the AI battleground.
  • The capability splits into two tiers with very different weight: a low-risk reading tier that summarizes and reasons across tabs, and an acting tier — "agent mode" — that clicks, books and fills forms on your behalf.
  • Indirect prompt injection is unsolved and category-wide: Brave showed hidden Reddit-comment instructions extracting an email address and a one-time login code from Comet, then showed injections a user cannot even see across Comet, Fellou and Opera Neon.
  • The most credible admission came from inside a vendor — OpenAI's CISO wrote that prompt injection "remains a frontier, unsolved security problem" while the feature shipped to millions.
  • Vendors and researchers do not even agree on what counts as a flaw: Perplexity classified LayerX's CometJacking report "Not Applicable," a judgment the researchers rejected.

In a single stretch of autumn 2025, the web browser stopped being a window and started trying to be an assistant. OpenAI launched ChatGPT Atlas on October 21, Microsoft expanded a "Copilot Mode" inside Edge on October 23, and Google had already begun putting Gemini directly into Chrome for free U.S. desktop users on September 18 [source: OpenAI, 2025; source: Microsoft, 2025; source: Google, 2025]. Perplexity, which had shipped its Comet browser in July, made it free to everyone on October 2 [source: Perplexity, 2025]. The pitch across all of them is roughly the same: a browser that reads pages for you, and — in "agent mode" — clicks, types, and acts on your behalf. In the very same weeks, security researchers published working demonstrations of how to hijack these agents using instructions hidden inside ordinary web pages [source: Brave, 2025]. That collision, between a genuinely useful new interface and a genuinely unsolved security problem, is the real story.

Why this is suddenly everywhere

The idea of an AI that browses for you is not new, but 2025 was the year it shipped from nearly every major company at once. Perplexity described Comet, launched on July 9, 2025, as the first "agentic AI browser" [source: Perplexity, 2025]. Within months the rest of the field arrived: Gemini in Chrome, Edge's Copilot Mode, ChatGPT Atlas, and Dia from The Browser Company, the startup behind the Arc browser [source: Google, 2025; source: Microsoft, 2025; source: OpenAI, 2025; source: TechCrunch, 2025]. The competitive logic is straightforward. The browser is where people already spend their working day, and whoever owns the assistant layer inside it owns an enormous amount of attention and data. What was marketed as convenience is also a land grab for the most valuable surface on the internet.

What actually changed for users

Strip away the marketing and there are really two tiers of capability, and they carry very different weight.

The first tier is reading. These browsers can summarize a long article, answer questions about the page in front of you, and reason across multiple open tabs at once — Google says Gemini can draw context from up to ten tabs in a window [source: Google, 2025]. This is a real, incremental convenience, and it is mostly low-risk, because the AI is only looking, not acting.

The second tier is doing. In "agent mode," the browser navigates and takes actions: Perplexity lists booking a hotel, comparison shopping, scheduling a meeting, and filling in forms; Microsoft's Copilot Actions include bulk-unsubscribing from newsletters and making reservations; Google says agentic features such as booking a haircut or ordering groceries are "coming months" away [source: Perplexity, 2025; source: Microsoft, 2025; source: Google, 2025]. This is the genuinely new thing. It is also where the marketing gets ahead of the evidence. OpenAI itself cautions that agent mode "may make mistakes on complex workflows," and independent verification of how reliably these agents complete real tasks remains thin [source: OpenAI, 2025]. A demo that books a hotel on stage is not the same as one that books the right hotel, on the right dates, without a costly error, every time.

A quick tour of who's who

ChatGPT Atlas (OpenAI) launched October 21, 2025 on macOS, with agent mode in preview for Plus, Pro, and Business tiers and Windows, iOS, and Android promised later [source: OpenAI, 2025].

Comet (Perplexity) launched July 9, 2025, initially behind a $200-a-month subscription and a waitlist, then made free to all on October 2, 2025, with mobile apps following [source: Perplexity, 2025].

Gemini in Chrome (Google) began rolling out to free U.S. desktop users on September 18, 2025, with cross-tab context and voice via Gemini Live, and agentic actions framed as still to come [source: Google, 2025].

Copilot Mode in Edge (Microsoft) expanded on October 23, 2025 with Copilot Actions, "Journeys," multi-tab reasoning, and opt-in access to your browsing history [source: Microsoft, 2025].

Dia (The Browser Company) opened a public beta on June 11, 2025 for members of its earlier Arc browser, offering a context-aware assistant and "skills"; the company has since wound down Arc to focus on Dia [source: TechCrunch, 2025].

The problem nobody has solved: prompt injection

Here is the core security issue, and it is not a bug in one product — it is a property of the whole category. Large language models cannot reliably tell the difference between instructions from you and instructions embedded in the content they read. When an agentic browser loads a web page and feeds it to its model, any text on that page — including text a malicious author planted — can be interpreted as a command. This is called indirect prompt injection.

Brave's security team published a concrete demonstration on August 20, 2025. They hid instructions inside a Reddit comment; when a user asked Comet to "summarize this webpage," the browser passed the page to its model without separating the user's request from the untrusted page content, and the hidden instructions took over [source: Brave, 2025]. In their proof of concept the injected commands could pull the user's email address and even a one-time login code from a connected account — enough, in principle, for account takeover [source: Brave, 2025]. Brave's blunt conclusion was that for these agents the classic browser defenses do not apply: "the same-origin policy or cross-origin resource sharing are all effectively useless," because "traditional web security assumptions don't hold for agentic AI" [source: Brave, 2025]. The reason is simple: the agent is acting as you, inside your logged-in sessions, so the walls that normally keep one website from touching another are irrelevant when the attacker is telling your own trusted assistant what to do.

It gets harder to defend against. In a follow-up published October 21, 2025, Brave showed injections that a user cannot even see: faint text that a browser's screenshot-and-OCR feature reads back as commands in Comet, navigation-based tricks in the Fellou browser, and hidden HTML in Opera's Neon [source: Brave, 2025]. Their assessment was that "indirect prompt injection is not an isolated issue, but a systemic challenge facing the entire category of AI-powered browsers" [source: Brave, 2025]. Other security firms reached similar conclusions from different angles. LayerX disclosed a technique it called "CometJacking" on August 27, 2025, in which a single crafted URL turns query parameters into commands that reach the agent's memory, Gmail, and calendar, using base64 encoding to slip past exfiltration safeguards [source: LayerX, 2025]. And Trail of Bits, in an audit published in 2026, catalogued several methods — from fake CAPTCHAs to threatening "system" messages — that extracted private Gmail data from Comet, tracing the root cause to the same flaw: "external content isn't treated as untrusted input" [source: Trail of Bits, 2026].

Company claims versus independent verification

The most honest voice here came from inside a vendor. OpenAI's chief information security officer, Dane Stuckey, wrote on October 22, 2025 that "prompt injection remains a frontier, unsolved security problem, and our adversaries will spend significant time and resources to find ways to make ChatGPT agent fall for these attacks" [source: OpenAI, 2025]. That is a striking admission from a company shipping the feature to millions. OpenAI's stated defenses are layered rather than absolute: a "logged-out mode" so the agent acts without your credentials, a "watch mode" for sensitive sites, plus red-teaming, model training, and rapid response [source: OpenAI, 2025]. Perplexity, for its part, says it fine-tuned an open model, Qwen3-30B, to scan raw HTML for injection attempts before the agent acts on a page [source: Perplexity, 2025].

But the gap between claim and verification is exactly where users should focus. Defenses that detect known attack patterns are in an arms race with attackers who invent new ones — which is what Brave's screenshot and OCR examples illustrate. And vendors do not always agree that a reported flaw is real: Perplexity classified the LayerX CometJacking report as "Not Applicable," a dispute the researchers rejected [source: LayerX, 2025]. When a company that builds the product and a firm that attacks it disagree on whether a vulnerability even counts, the marketing claim and the independent evidence are not describing the same world. The pattern to watch for is whether a defense holds up when someone outside the company tries to break it, not whether it sounds reassuring in a launch post.

Weighing productivity against risk, fairly

None of this means agentic browsers are useless or that the danger is hypothetical in both directions. The productivity case is real for low-stakes, tedious tasks: summarizing a dense report, pulling the same fact out of ten tabs, or unsubscribing from a pile of newsletters saves genuine time, and the downside of a mistake is small. The problem is that the features companies most want to sell — the agent that shops, books, and manages your accounts — are precisely the ones that combine three risky properties at once: broad access to your logged-in data, the power to take actions that are hard to undo, and an unsolved injection problem that can turn a booby-trapped page into a command channel. The convenience and the exposure grow together, not separately.

A fair reading, then, is neither "this changes everything" nor "this is all hype." The reading tier is a useful upgrade available today. The acting tier is a promising capability that is not yet trustworthy enough to hand your most sensitive accounts, and the most credible people saying so include a vendor's own security chief. The reasonable posture is to use the assistant freely for reading, and to treat agent mode the way you would treat handing a stranger your logged-in laptop: fine for something trivial, not for your bank.

What to watch next

The question that decides this category is architectural. Today's agents mostly pour the user's request and the untrusted page into the same model with no firm boundary between them; the durable fix is a genuine trust boundary that keeps page content from ever being read as a command. Watch for whether vendors move toward that, and whether their injection defenses survive independent testing rather than internal assurances. Watch, too, for whether regulators and standards bodies start requiring things like scoped permissions and audit logs for agents that act with your identity. And in the meantime, the practical move for a cautious user is simple: enjoy the summaries, keep the agent logged out of anything you would not want a stranger touching, and give agent mode access to your money, health, and primary email only once the evidence — not the marketing — says it has earned it.

Timeline

  1. The Browser Company teases Dia, an AI-centric successor to its Arc browser.

  2. Dia opens a public beta for existing Arc members, offering a context-aware assistant and "skills."

    TechCrunch — Dia public beta (opens in a new tab)
  3. Perplexity launches Comet, describing it as the first "agentic AI browser" — initially behind a $200-a-month subscription and a waitlist.

    Perplexity — Introducing Comet (opens in a new tab)
  4. Brave publishes a working indirect prompt injection against Comet: instructions hidden in a Reddit comment reach the user's email address and a one-time login code.

    Brave — Comet prompt injection (opens in a new tab)
  5. LayerX discloses "CometJacking" — one crafted URL turns query parameters into agent commands reaching memory, Gmail and calendar.

    LayerX — CometJacking (opens in a new tab)
  6. Google begins rolling Gemini into Chrome for free U.S. desktop users, with cross-tab context and Gemini Live voice; agentic actions framed as still to come.

    Google — Gemini in Chrome (opens in a new tab)
  7. Comet becomes free to everyone, removing the paywall that had limited the agentic browser to subscribers.

    Perplexity — Comet free rollout (opens in a new tab)
  8. OpenAI launches ChatGPT Atlas on macOS, with agent mode in preview for Plus, Pro and Business tiers.

    OpenAI — Introducing ChatGPT Atlas (opens in a new tab)
  9. Brave follows up with "unseeable" injections — faint text read back through screenshot-and-OCR in Comet, navigation tricks in Fellou, hidden HTML in Opera Neon — calling it a systemic challenge for the whole category.

    Brave — Unseeable prompt injections (opens in a new tab)
  10. OpenAI CISO Dane Stuckey states publicly that prompt injection "remains a frontier, unsolved security problem"; Atlas ships with layered mitigations rather than a fix.

    OpenAI — Atlas and prompt injection (opens in a new tab)
  11. Microsoft expands Copilot Mode in Edge with Copilot Actions, "Journeys," multi-tab reasoning and opt-in access to browsing history.

    Microsoft — Copilot Mode in Edge (opens in a new tab)
  12. Comet arrives on Android, extending the agentic browser beyond the desktop.

  13. Trail of Bits publishes a Comet audit cataloguing several extraction methods — from fake CAPTCHAs to threatening "system" messages — and traces them to one root cause: "external content isn't treated as untrusted input."

    Trail of Bits — Auditing Comet (opens in a new tab)
  14. Comet reaches iOS, completing the mobile rollout that followed the free release.

Analysis

Two tiers, two entirely different risk profiles

Reading — summarizing a page, answering questions about it, pulling one fact out of ten open tabs — is an incremental convenience with a small downside when it goes wrong. Acting is a different category: the agent navigates, clicks, books and fills forms inside your logged-in sessions. Almost every disputed claim in this field collapses once you ask which tier it belongs to.

The land grab explains the timing

The idea of an AI that browses for you is not new, but nearly every major company shipped it within a few months of each other in 2025. The browser is where people already spend the working day, and whoever owns the assistant layer inside it owns an enormous amount of attention and data. What was marketed as convenience is also a fight for the most valuable surface on the internet.

Classic browser defenses do not transfer

Brave's blunt conclusion was that for these agents "the same-origin policy or cross-origin resource sharing are all effectively useless." The reason is structural rather than incidental: the agent is acting as you, inside your own sessions, so walls designed to stop one website from touching another are irrelevant when the attacker is simply telling your trusted assistant what to do.

The attack got harder to see, not easier to block

The August demonstration used text a user could in principle have found. The October follow-up used text a user cannot see at all — faint characters recovered through a screenshot-and-OCR feature, navigation-based tricks, hidden HTML. That the same class of failure appeared in three different browsers is what moves this from a product bug to what Brave called "a systemic challenge facing the entire category."

The most credible warning came from inside a vendor

A security researcher saying a product is unsafe is expected. A shipping company's own chief information security officer writing that prompt injection "remains a frontier, unsolved security problem" — and that adversaries will spend significant time and resources on it — is a different kind of evidence, because it runs against the speaker's commercial interest.

Mitigations are layered, and every one of them is company-stated

Logged-out mode, watch mode for sensitive sites, red-teaming, rapid response, and Perplexity's fine-tuned Qwen3-30B model scanning raw HTML are all defenses described by the companies that built them. None has been shown here to survive an outside attempt to break it, and defenses that recognize known attack patterns are structurally behind attackers who invent new ones — which is exactly what the screenshot-and-OCR case demonstrated.

The parties do not agree on what counts as a vulnerability

Perplexity classified LayerX's CometJacking report as "Not Applicable"; the researchers rejected that judgment. This is the sharpest illustration of the article's thesis — when the company that builds a product and the firm that attacks it cannot agree that a flaw exists, a reader cannot treat the marketing claim and the independent finding as two views of the same fact.

Why this article carries no chart

Only three quantities appear in the reporting at all: the ten tabs Gemini can draw context from, Comet's original $200-a-month price, and the 30-billion-parameter size of the model Perplexity fine-tuned as a scanner. They come from three different companies, describe three unrelated things — a capability ceiling, a since-removed price, and a model size — and share no unit or measurement basis, so no honest axis can hold two of them. Charting them side by side would manufacture a comparison the evidence does not support. The table below lists each figure with its qualifier instead.

Comparison

Who shipped what, and how far each one goes — dates and tiers as stated by each company.
Product (company)DateReading tier / acting tier, as stated
Comet (Perplexity)2025-07-09 launch; free 2025-10-02Both — self-described first "agentic AI browser"; lists hotel booking, comparison shopping, meeting scheduling, form filling
Gemini in Chrome (Google)2025-09-18 rollout beginsReading now — cross-tab context, Gemini Live voice; agentic actions such as booking a haircut framed as "coming months" away
ChatGPT Atlas (OpenAI)2025-10-21 launch (macOS)Both — agent mode in preview for Plus, Pro and Business; the company cautions it "may make mistakes on complex workflows"
Copilot Mode in Edge (Microsoft)2025-10-23 expansionBoth — Copilot Actions (bulk unsubscribe, reservations), "Journeys," multi-tab reasoning, opt-in history access
Dia (The Browser Company)2025-06-11 public betaReading — context-aware assistant and "skills" for former Arc members; Arc has since been wound down
Claim, speaker, and evidence tier. Company-stated defenses and independently demonstrated attacks are not interchangeable.
ClaimWho is making itEvidence tier
Agent mode books hotels, shops, schedules meetings and fills formsPerplexityCompany-stated capability
Copilot Actions bulk-unsubscribes from newsletters and makes reservationsMicrosoftCompany-stated capability
Agent mode "may make mistakes on complex workflows"OpenAICompany-stated limitation — against its own interest
Prompt injection "remains a frontier, unsolved security problem"OpenAI CISO Dane StuckeyCompany-stated admission — against its own interest
Hidden Reddit-comment instructions extracted an email address and a one-time login codeBraveIndependently demonstrated proof of concept
Injections a user cannot see work across Comet, Fellou and Opera NeonBraveIndependently demonstrated across three browsers
One crafted URL reaches the agent's memory, Gmail and calendar ("CometJacking")LayerX; Perplexity classified it "Not Applicable"Independently demonstrated — disputed by the vendor
Every quantity this story reports, with the qualifier attached to it. Each is a lone figure on its own basis, which is why none of them sits on a chart axis.
FigureWhat it actually measuresWhy it stays off an axis
Up to 10 tabsContext Gemini can draw from in one Chrome window (Google)A ceiling on one product's capability, not a measured outcome; a ceiling must not be redrawn as a point value
$200 a monthComet's original subscription price before the October 2 free release (Perplexity)A single price with no comparable second value; it was removed rather than changed, so there is no before/after series
Qwen3-30BSize of the open model Perplexity says it fine-tuned to scan raw HTML for injectionsA model parameter count from a third company — a different unit and a different subject entirely
No figure reportedHow reliably agents complete real tasksIndependent verification is described as thin; no party publishes a completion rate, and absence of a number cannot be charted

Process

  1. Ask which tier the claim belongs to

    Reading is a modest, low-risk upgrade available today; acting is the promising, unproven part. Marketing blurs the two deliberately.

  2. Identify who is making the claim

    A capability described in a launch post is company-stated. An attack reproduced by an outside firm is independently demonstrated. They are not the same kind of evidence.

  3. Check whether the defense survived outside testing

    The pattern to watch for is whether a defense holds up when someone outside the company tries to break it, not whether it sounds reassuring in a launch post.

  4. Look at what credentials the agent can reach

    The risk comes from acting inside your logged-in sessions. A logged-out agent that only reads carries a fraction of the exposure.

  5. Ask whether the action can be undone

    Summarizing a report is reversible. A booking, a purchase or an email sent as you is not — and irreversibility is where an injected command does its damage.

  6. Grant access last, and from the bottom up

    Enjoy the summaries; keep the agent logged out of anything you would not want a stranger touching; give agent mode your money, health and primary email only once the evidence, not the marketing, says it has earned it.

Sources

  1. OpenAI — Introducing ChatGPT Atlas (2025-10-21).View source (opens in a new tab)
  2. OpenAI (Dane Stuckey, CISO) — Statement on Atlas and prompt injection (2025-10-22).View source (opens in a new tab)
  3. Perplexity — Introducing Comet and free rollout (2025).View source (opens in a new tab)
  4. Google — Gemini in Chrome and new AI features (2025-09-18).View source (opens in a new tab)
  5. Microsoft — Copilot Mode in Edge (2025-10-23).View source (opens in a new tab)
  6. The Browser Company / TechCrunch — Dia public beta (2025-06-11).View source (opens in a new tab)
  7. Brave — Comet prompt injection demonstration (2025-08-20).View source (opens in a new tab)
  8. Brave — Unseeable prompt injections across AI browsers (2025-10-21).View source (opens in a new tab)
  9. LayerX — CometJacking (2025-08-27).View source (opens in a new tab)
  10. Trail of Bits — Auditing Comet for prompt-injection data leaks (2026).View source (opens in a new tab)

Tags

  • #ai-browser
  • #agentic-ai
  • #prompt-injection
  • #chatgpt-atlas
  • #perplexity-comet
  • #browser-security
AI Browsers Are Here: What They Do and What They Risk | 생활데이터랩