User guide

Page tools

A page tool is an action a website offers your agent directly. This is how to find out what is on offer, and how to ask for it.

A page tool is an action a website offers directly to your agent.

Ordinarily an assistant working with a web application has two bad options: read the screen and guess, or drive the mouse and hope. A page tool is the website skipping both — publishing, in machine-readable form, here are the things you can ask me to do: search these tickets, set this status, refund this order. When the agent uses one, it calls that action directly rather than simulating somebody clicking.

This is better for you in three ways, and worse in one.

Better:

  • It is precise. A tool takes named arguments. There is no misclicked button and no half-filled form.
  • It is legible. You can see exactly which action ran and with what — which is not true of an assistant clicking around a page on your behalf.
  • It is classified. A site marks its actions as read-only, additive or destructive, and XataWorks treats the three differently.

Worse:

  • It runs as you. A page tool executes in your browser, in your session, with your permissions. The website is not asking XataWorks for permission to do something — it is offering the agent a lever that is already connected to your account.

That last point is the whole reason for Approvals and permissions. Everything XataWorks does around page tools follows from it.

Seeing what is on offer

You never have to guess. Two controls show you, and neither of them runs anything.

The page-tools control, for one tab

The browser toolbar carries it, with a count. Click it for the list of what the current tab publishes: each operation’s friendly name, its description, and its class. A filter box narrows the list by keyword, and chips along the top narrow it by class.

The page-tools list open from the browser toolbar, headed Page tools (6), with a filter box and chips reading All (6), Read only (2), Mutating (2) and Destructive (2), above six named operations each showing its class.
One tab’s operations. The classes beside them were not chosen for this picture: they are what XataWorks worked out from what the site declared about each operation.

The Tools palette, for everything this chat can reach

The page-tools control shows one tab. The chat header’s ⋮ menu has a Tools entry that shows everything this chat can currently reach — every open tab, grouped by tab, with a friendly label for each.

This is the one to open when the agent says it cannot do something and you are certain it should be able to, because the usual answer is that the tab it needs is open in a different chat.

  • It is read-only. Nothing in it runs a tool.
  • It shows this chat’s tabs only. A family of chats sharing one browser shares its tabs, so what you see is what this conversation can reach and nothing else.
  • A filter box narrows as you type, matching the friendly name, the technical id, the description and the tab label. It is display-only and resets each time you open the palette.

If the palette is empty, that is a real answer: either no tab in this chat is offering anything, or the chat has no browser at all — a sub-agent dispatched to do a job on its own does not have one.

A dialog headed Available tools, with a filter box, then a group headed Larkspur Helpdesk — Tickets with its host beneath. Six operations are listed, each with a description and its technical name: Search Larkspur tickets and Read a Larkspur ticket are badged Read only; Add a note to a Larkspur ticket and Set a Larkspur ticket's status are badged Mutating; Close a Larkspur ticket is badged Destructive and off-site.
Everything this chat can reach, grouped under the tab it came from. Each row carries the operation’s own description — the same text the agent reads when it decides which one to use.
The same two routes, in order: click the page-tools control in the browser toolbar to see what this tab offers; close it; open the chat header’s ⋮ menu and choose Tools to see everything this chat can reach; type in the filter box to narrow it. This recording has no sound.

Using one

You do not invoke a page tool. You ask for the outcome, in your own words, and the agent picks the tool.

The skill, such as it is, lies in two habits.

Say which page you mean

The agent can only use tools from tabs this chat has open. If you have several, say which — “on the support board”, “in the billing tab” — rather than relying on it to guess. If the site is not open at all, ask it to open the site first; it can.

Name the operation when precision matters

Most of the time describing the outcome is better: it lets the agent choose, and it recovers when your idea of the name is wrong. But when a site publishes two similar actions and you care which one runs, name it. The names in the page-tools list are the names to use.

Reading what happened

A run in the transcript names the operation that ran and the tab it ran in. This matters more than it sounds: a run labelled only as “browsing” tells you nothing about what was done, and XataWorks names the specific action instead — so a call you did not expect is visible as a call you did not expect.

Fold a run open to see what it was given and what came back.

A chat transcript, cropped to the chat column, over the helpdesk. The question reads “Read Larkspur ticket LK-4417 and tell me who raised it.” Below it a run reading “Reading skill”, the assistant saying it will look the ticket up, a folded group reading “4 tools”, and then the answer: the ticket was raised by Brightwood Joinery, a table of its fields, and two notes the assistant drew from the team's triage skill. At the bottom a meter reading 4% of context with a Compact button.
One answered turn, from a question in ordinary words. A page-tool call reads Calling <operation> via WebMCP in the transcript — there is one in the figure in Approvals and permissions. The answer is a real one from a real model reading the site’s own data, so its wording would differ if the picture were taken again.

Tools it remembers

Some sites publish their operations only after you have interacted with the page — which would mean the agent could not plan around one until after it had stumbled into it. XataWorks remembers what a site published last time you were there, so the agent knows what is likely to be available before the page has finished telling it.

Remembered tools are a hint, not an authority. What actually runs is what the page publishes now, and permissions are unchanged by memory: a remembered tool is still asked about exactly as it would have been the first time.

When a result is too big

Page tools frequently answer with far more than anyone needs. Ask a ticketing system for open tickets and it may hand back every field of every ticket — tens of thousands of words to answer a question about six of them.

XataWorks does not put all of that into your conversation. A result past a sensible size is trimmed before it reaches the agent, and the agent is told it was trimmed and roughly by how much.

  • The answer can still be right. Trimming removes bulk, and the agent knows the bulk existed.
  • The better move is to narrow the question. Just the ticket numbers and their ages produces a smaller result, and a smaller result is an undiminished one.
  • It matters for questions about completeness. How many are there in total cannot be answered from a trimmed result. Ask the site for a count rather than for the things being counted; a count is never too big.

Sites XataWorks already understands

Most websites publish nothing at all. For a number of widely used systems, XataWorks brings its own understanding, so the agent has real actions to work with whether or not the site has done anything to support assistants.

You do not turn these on. Open the site in the chat’s browser and its operations appear, exactly as if the site had published them — listed the same way, and approved the same way.

SystemWhat the agent gains
GmailSearching the mailbox, and working from the threads currently visible. Composing is confirmed: the agent fills a message in front of you and you send it.
Google DriveFinding files and working with what is on screen.
Google DocsReading and working within the open document.
JiraIssues — finding them, reading them, and the ordinary updates a board needs.
ZendeskTickets: searching, reading, and the fields a triage pass touches.
SlackChannels and messages in the workspace you have open.
ShopifyStore data — orders, products, customers — from the admin you are signed in to.
GrafanaDashboards and panels: what is on them and what they currently say.
LaunchDarklyFlags and their state. Toggling a flag is confirmed every time, whatever else you have granted the site.
AWS ConsoleWhat the console can see from the account and region you are signed in to.
draw.ioBuilding and editing the diagram in the open editor.
Mermaid LiveBuilding and editing the diagram in the open editor.
PlexThe library you are signed in to.

Confirmed actions stay confirmed. Some of the actions above always ask — sending mail, toggling a feature flag. That confirmation is not a permission you have not granted yet; it is a property of the action, and a broad grant on the site does not remove it.

The agent still needs the site open, and still needs permission. XataWorks understanding a system gives the agent actions to use. It does not sign you in, it does not open the site, and it does not decide for you that the agent may act there.

Back to all user-guide articles.