
A browser agent as a research preview
On 23 January 2025, OpenAI introduced Operator, describing it as a research preview available to Pro-tier ChatGPT users in the United States, built to complete browser-based tasks such as filling in a form or placing an order by typing, clicking and scrolling like a person would. It is powered by a separate model OpenAI calls Computer-Using Agent (CUA), which combines GPT-4o's vision with reinforcement-learning-trained reasoning over screenshots, and which OpenAI reported scoring 38.1% on the OSWorld computer-use benchmark, 58.1% on WebArena and 87.0% on WebVoyager, against stated prior best results of 22.0%, 36.2% and 56.0% respectively. Human performance on the same benchmarks was reported at 72.4% and 78.2%, a gap the announcement did not describe as closed.
Confirmation built into the design
The system card discloses that Operator asks for explicit approval before finalising a consequential action, such as submitting a purchase or sending an email, and that this confirmation step achieved 92% recall on risky actions across 607 internally constructed test tasks; a related watch mode pauses execution on sensitive sites such as webmail if the user appears to have stopped paying attention. The card also reports a prompt-injection monitor, intended to catch instructions hidden in a webpage rather than typed by the user, at 99% recall, and describes Operator as trained to proactively refuse categories of task such as unsupervised stock trading.
What the disclosed numbers leave open
None of these figures describe a system without errors. The same document reports a 13% baseline error rate across 100 representative user tasks, including five errors it classifies as irreversible or severe, and states that resistance to novel, previously unseen prompt-injection techniques remains an ongoing challenge rather than a solved one. The benchmark comparisons are OpenAI's own measurements against baselines it also selected, and a recall rate measured on a fixed internal test set is a statement about that test set's coverage of risky actions, not a guarantee that every risky action in the wild will be caught the same way.
Questions to carry into your own evaluation
- Which actions in your own workflow would need a confirmation step before a browser agent is allowed to take them unattended?
- Does the agent's disclosed error rate include the specific task category you plan to use it for?
- How would a prompt-injection attempt reach the agent through content on a page it visits, rather than through your own instructions?
Operator's own documentation is unusually specific about where it still fails. That specificity is a more useful signal than the headline benchmark numbers on their own.
Sources & reading trail
Describes Operator as a research preview with a takeover mode, order-confirmation prompts, and refusal of sensitive tasks such as banking.
Source published: Not established · Retrieved: 16 September 2026
Reports CUA's benchmark results against stated prior state-of-the-art and human baselines, and discloses UI and text-editing weaknesses.
Source published: Not established · Retrieved: 16 September 2026
Discloses the preparedness scorecard, the 92% recall confirmation-prompt measurement across 607 test tasks, and a 13% baseline error rate.
Source published: Not established · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.