RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Practice

Practice / From the record · 23 January 2025 event · prepared 16 September 2026

A browser agent shipped with confirmation prompts built in

OpenAI's Operator launch materials and system card describe a browser-driving research preview built around confirmation prompts and site restrictions.

Visual for this record: A browser agent shipped with confirmation prompts built in
Visual published by docs.livevox.com, shown for identification of the record. Credit: docs.livevox.com · source page ↗ Rights: owner-review-pending.

A browser agent as a research preview

On 23 January 2025, OpenAI introduced Operator, describing it as a research preview available to Pro-tier ChatGPT users in the United States, built to complete browser-based tasks such as filling in a form or placing an order by typing, clicking and scrolling like a person would. It is powered by a separate model OpenAI calls Computer-Using Agent (CUA), which combines GPT-4o's vision with reinforcement-learning-trained reasoning over screenshots, and which OpenAI reported scoring 38.1% on the OSWorld computer-use benchmark, 58.1% on WebArena and 87.0% on WebVoyager, against stated prior best results of 22.0%, 36.2% and 56.0% respectively. Human performance on the same benchmarks was reported at 72.4% and 78.2%, a gap the announcement did not describe as closed.

Confirmation built into the design

The system card discloses that Operator asks for explicit approval before finalising a consequential action, such as submitting a purchase or sending an email, and that this confirmation step achieved 92% recall on risky actions across 607 internally constructed test tasks; a related watch mode pauses execution on sensitive sites such as webmail if the user appears to have stopped paying attention. The card also reports a prompt-injection monitor, intended to catch instructions hidden in a webpage rather than typed by the user, at 99% recall, and describes Operator as trained to proactively refuse categories of task such as unsupervised stock trading.

What the disclosed numbers leave open

None of these figures describe a system without errors. The same document reports a 13% baseline error rate across 100 representative user tasks, including five errors it classifies as irreversible or severe, and states that resistance to novel, previously unseen prompt-injection techniques remains an ongoing challenge rather than a solved one. The benchmark comparisons are OpenAI's own measurements against baselines it also selected, and a recall rate measured on a fixed internal test set is a statement about that test set's coverage of risky actions, not a guarantee that every risky action in the wild will be caught the same way.

Questions to carry into your own evaluation

  • Which actions in your own workflow would need a confirmation step before a browser agent is allowed to take them unattended?
  • Does the agent's disclosed error rate include the specific task category you plan to use it for?
  • How would a prompt-injection attempt reach the agent through content on a page it visits, rather than through your own instructions?

Operator's own documentation is unusually specific about where it still fails. That specificity is a more useful signal than the headline benchmark numbers on their own.

Sources & reading trail

Introducing Operator ↗

Describes Operator as a research preview with a takeover mode, order-confirmation prompts, and refusal of sensitive tasks such as banking.

Source published: Not established · Retrieved: 16 September 2026

Computer-Using Agent ↗

Reports CUA's benchmark results against stated prior state-of-the-art and human baselines, and discloses UI and text-editing weaknesses.

Source published: Not established · Retrieved: 16 September 2026

Operator System Card ↗

Discloses the preparedness scorecard, the 92% recall confirmation-prompt measurement across 607 test tasks, and a 13% baseline error rate.

Source published: Not established · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.