User guide
Page tools
A page tool is an action a website offers your agent directly. This is how to find out what is on offer, and how to ask for it.
A page tool is an action a website offers directly to your agent.
Ordinarily an assistant working with a web application has two bad options: read the screen and guess, or drive the mouse and hope. A page tool is the website skipping both — publishing, in machine-readable form, here are the things you can ask me to do: search these tickets, set this status, refund this order. When the agent uses one, it calls that action directly rather than simulating somebody clicking.
This is better for you in three ways, and worse in one.
Better:
- It is precise. A tool takes named arguments. There is no misclicked button and no half-filled form.
- It is legible. You can see exactly which action ran and with what — which is not true of an assistant clicking around a page on your behalf.
- It is classified. A site marks its actions as read-only, additive or destructive, and XataWorks treats the three differently.
Worse:
- It runs as you. A page tool executes in your browser, in your session, with your permissions. The website is not asking XataWorks for permission to do something — it is offering the agent a lever that is already connected to your account.
That last point is the whole reason for Approvals and permissions. Everything XataWorks does around page tools follows from it.
Seeing what is on offer
You never have to guess. Two controls show you, and neither of them runs anything.
The page-tools control, for one tab
The browser toolbar carries it, with a count. Click it for the list of what the current tab publishes: each operation’s friendly name, its description, and its class. A filter box narrows the list by keyword, and chips along the top narrow it by class.
The Tools palette, for everything this chat can reach
The page-tools control shows one tab. The chat header’s ⋮ menu has a Tools entry that shows everything this chat can currently reach — every open tab, grouped by tab, with a friendly label for each.
This is the one to open when the agent says it cannot do something and you are certain it should be able to, because the usual answer is that the tab it needs is open in a different chat.
- It is read-only. Nothing in it runs a tool.
- It shows this chat’s tabs only. A family of chats sharing one browser shares its tabs, so what you see is what this conversation can reach and nothing else.
- A filter box narrows as you type, matching the friendly name, the technical id, the description and the tab label. It is display-only and resets each time you open the palette.
If the palette is empty, that is a real answer: either no tab in this chat is offering anything, or the chat has no browser at all — a sub-agent dispatched to do a job on its own does not have one.
Using one
You do not invoke a page tool. You ask for the outcome, in your own words, and the agent picks the tool.
The skill, such as it is, lies in two habits.
Say which page you mean
The agent can only use tools from tabs this chat has open. If you have several, say which — “on the support board”, “in the billing tab” — rather than relying on it to guess. If the site is not open at all, ask it to open the site first; it can.
Name the operation when precision matters
Most of the time describing the outcome is better: it lets the agent choose, and it recovers when your idea of the name is wrong. But when a site publishes two similar actions and you care which one runs, name it. The names in the page-tools list are the names to use.
Reading what happened
A run in the transcript names the operation that ran and the tab it ran in. This matters more than it sounds: a run labelled only as “browsing” tells you nothing about what was done, and XataWorks names the specific action instead — so a call you did not expect is visible as a call you did not expect.
Fold a run open to see what it was given and what came back.
Tools it remembers
Some sites publish their operations only after you have interacted with the page — which would mean the agent could not plan around one until after it had stumbled into it. XataWorks remembers what a site published last time you were there, so the agent knows what is likely to be available before the page has finished telling it.
Remembered tools are a hint, not an authority. What actually runs is what the page publishes now, and permissions are unchanged by memory: a remembered tool is still asked about exactly as it would have been the first time.
When a result is too big
Page tools frequently answer with far more than anyone needs. Ask a ticketing system for open tickets and it may hand back every field of every ticket — tens of thousands of words to answer a question about six of them.
XataWorks does not put all of that into your conversation. A result past a sensible size is trimmed before it reaches the agent, and the agent is told it was trimmed and roughly by how much.
- The answer can still be right. Trimming removes bulk, and the agent knows the bulk existed.
- The better move is to narrow the question. Just the ticket numbers and their ages produces a smaller result, and a smaller result is an undiminished one.
- It matters for questions about completeness. How many are there in total cannot be answered from a trimmed result. Ask the site for a count rather than for the things being counted; a count is never too big.
Sites XataWorks already understands
Most websites publish nothing at all. For a number of widely used systems, XataWorks brings its own understanding, so the agent has real actions to work with whether or not the site has done anything to support assistants.
You do not turn these on. Open the site in the chat’s browser and its operations appear, exactly as if the site had published them — listed the same way, and approved the same way.
| System | What the agent gains |
|---|---|
| Gmail | Searching the mailbox, and working from the threads currently visible. Composing is confirmed: the agent fills a message in front of you and you send it. |
| Google Drive | Finding files and working with what is on screen. |
| Google Docs | Reading and working within the open document. |
| Jira | Issues — finding them, reading them, and the ordinary updates a board needs. |
| Zendesk | Tickets: searching, reading, and the fields a triage pass touches. |
| Slack | Channels and messages in the workspace you have open. |
| Shopify | Store data — orders, products, customers — from the admin you are signed in to. |
| Grafana | Dashboards and panels: what is on them and what they currently say. |
| LaunchDarkly | Flags and their state. Toggling a flag is confirmed every time, whatever else you have granted the site. |
| AWS Console | What the console can see from the account and region you are signed in to. |
| draw.io | Building and editing the diagram in the open editor. |
| Mermaid Live | Building and editing the diagram in the open editor. |
| Plex | The library you are signed in to. |
Confirmed actions stay confirmed. Some of the actions above always ask — sending mail, toggling a feature flag. That confirmation is not a permission you have not granted yet; it is a property of the action, and a broad grant on the site does not remove it.
The agent still needs the site open, and still needs permission. XataWorks understanding a system gives the agent actions to use. It does not sign you in, it does not open the site, and it does not decide for you that the agent may act there.
Back to all user-guide articles.