Over the past two years, one of the most visible shifts in large models has been the move from "answering questions in a chat box" to "agents that can get things done." The browser is one of the most natural battlegrounds for this shift. The so-called agentic browser embeds an AI agent directly into the browser: it does not merely answer questions, but reads the page you are on, performs clicks and form filling, and completes a whole task across multiple websites.
How is it different from a traditional browser, or from ordinary "AI search"? A traditional browser hands full control to the user, with AI acting only as a plug-in; AI search mainly helps you "find information and summarize it." The agentic browser is more ambitious—it wants to be your "operating proxy": when you say "compare prices and order from the cheapest one," it has to open several tabs, read prices, handle logins, and ask for your confirmation at key steps. The core difference is whether it directly changes the state of a website and presses the buttons for you.
To actually work, such products usually need several key capabilities. First, page understanding: turning complex page structure, dynamic content and visual information into something the agent can "read," judging which elements are clickable and which are input fields. Second, action execution: reliably clicking, typing, scrolling and switching tabs, and coping with pop-ups, CAPTCHAs and loading delays. Third, multi-step planning and memory: breaking a big task into steps, remembering intermediate results, and retrying or switching paths when something fails. Fourth, human-in-the-loop: pausing to ask for your confirmation before high-risk actions such as paying, deleting or submitting forms.
Several directions are already being explored: placing the agent in a browser sidebar, building a standalone browser that can "take over tabs," or using extensions to bring "automation" to existing pages. It is fair to expect that "whose agent can reliably complete real tasks" will be the focus of competition for some time.
But the greater the power, the more concrete the risks. The most prominent is prompt injection: a web page may hide malicious instructions that trick the agent into doing things the user never intended, such as leaking email contents or making a transfer. Next comes permission and privacy: letting AI "operate for you" often means it can see your logged-in sessions and sensitive data, and abuse would be far more serious than "a wrong answer." Third is reliability: real web pages are full of uncertainty, and success rates are far from reassuring.
As a result, security boundaries are likely to be the watershed for whether agentic browsers go mainstream: sandbox isolation, action allow-lists, second confirmation for high-risk operations, and "down-weighting" untrusted content. These engineering practices will decide whether users dare to hand important accounts and tasks to such agents.
For ordinary users, the more realistic stance in the short term is to treat an agentic browser as "a capable intern who still needs supervision." Let it handle low-risk tasks such as looking things up, filling forms and organizing information, while keeping payment, authorization and deletion decisions in your own hands.
[Reference] Compiled from publicly released industry information.