Skip to content
Everruns Cloud is open in early access. Run agents without operating the platform.

Computer Use

IDcomputer_use
CategoryBrowser
FeaturesLeased resources
Dependenciessession_storage
RiskHigh (the model sees and acts on whatever the screen shows)
AvailabilityExperimental

Computer use lets an agent work a screen the way a person does. It takes a screenshot, decides where to click or what to type, and gets a new screenshot after each action. It works on pages that selector-based browser tools cannot handle, such as canvas apps, heavy single-page apps, and custom widgets.

The capability is provider-neutral. It is a regular function tool that returns images, so it works with OpenAI, Anthropic, Gemini, and any other model that accepts images in tool results.

The display is a Chromium page on a Browserless browser. Connect Browserless in Settings > Connections > Browserless first. The agent shares one persistent browser per session with the Browserless tools, so a page opened by browserless_open_browser is the page computer sees.

Performs one action and returns a screenshot of the display afterwards. Coordinates are pixels in the latest screenshot, with the origin at the top left.

actionParametersWhat it does
screenshotnoneCapture the display
left_click, right_click, middle_click, double_click, triple_clickcoordinate (optional [x, y], defaults to the cursor), text (optional modifiers such as shift or ctrl+shift)Click
left_click_dragstart_coordinate, coordinatePress, drag, and release
mouse_movecoordinateHover without clicking
scrollscroll_direction (up, down, left, right), scroll_amount (wheel clicks, 1 to 50), optional coordinateScroll
typetextType text at the keyboard focus
keytext (a key or combo such as Return, Tab, ctrl+a), optional repeatPress keys
waitduration (seconds, up to 30)Pause
navigateurlLoad a page

Example call:

{ "action": "left_click", "coordinate": [412, 230] }
FieldDefaultDescription
display_width1280Display width in pixels (320 to 1920)
display_height800Display height in pixels (320 to 1200)
screenshot_after_actiontrueReturn a screenshot after every action. When off, only screenshot returns an image.
max_actions_per_session300Hard cap on actions in one session, screenshots included

Screenshots are billed as image tokens. A smaller display, or turning off screenshot_after_action, lowers the cost per step.

  • The model sees everything on the screen. Text on a page can try to steer the agent. The capability tells the model to treat screen contents as untrusted and to stop and ask before typing credentials, making purchases, sending messages, or confirming irreversible actions. Add the soft_approval capability when you want those confirmations recorded.
  • Keep credentials out of reach. Do not give a computer-use agent a browser that is signed in to accounts it should not use.
  • Egress. navigate refuses private and internal addresses and follows the session’s network access list. If a page navigates somewhere blocked on its own, the page is reset to a blank page and the action reports an error.
  • Budget. max_actions_per_session stops runaway loops.

Native OpenAI and Anthropic computer tools, and desktop displays on sandboxes, are planned.