Setting up the agent
The demo needs an AI client with two things wired up: a model reached through Cloudflare AI Gateway, and the five MCP servers reached through the MCP server portal. These are the two places the controls live.
- LLM / model traffic goes to
, the AI Gateway's own custom domain which is behind Cloudflare Access. On a custom domain, AI Gateway accepts a valid Access JWT as the request credential, so the client sends no gateway token and no API key at all. - MCP / tool traffic goes to the MCP portal at
, which fronts all the deployed MCP servers.
Access secures both traffic paths and therefore every request is attributed to the person who made it.
Sign in as the demo user
Every demo is run as Alice Watson — alice.watson@company.com,
password Savetheinternet!1. When the agent connects to the portal you will be sent
through Cloudflare Access and then FlareID.
admin@company.com is the system administrator, and it exists only for the
/admin page on each app.
First: authenticate each MCP server (one browser login each)
The script registers five MCP servers with Access and creates an Access application for each, but it cannot perform the upstream OAuth login those servers require. Until it is done each server sits in Waiting and the portal has nothing to offer your agent.
- In the Cloudflare dashboard, go to Zero Trust → Access controls → MCP Portals
→ MCP servers. All five (
hr,crm,work,wiki,fin) should be listed, each showing Waiting. - Click on the three dots next to each server → Authenticate server.
- Sign in as
nikita.chapman@company.com(passwordSavetheinternet!1). Cloudflare fetches the server's tools and the status becomes Ready. - Repeat for
crm,work,finandwiki.
admin@company.com exists in FlareID but is not an employee in any of the
apps, and every MCP server maps the Access identity to an employee record before it will
issue a token.
Setting up the clients
Three are documented below and you only need one. All three reach the same two endpoints, so every demo script works with any of them.
- Goose — use this one. Desktop app, native tool calls, and the only client of the three that shows the audience why something was blocked: it surfaces the gateway's own message instead of a bare status code. One OAuth covers every server behind the portal, for good.
- opencode — also a desktop app with native tool
calls, and what the walkthrough was originally written against. Workable, but it reports a
refusal as
Provider request failed with HTTP 424and nothing else, so every block becomes a trip to the dashboard to explain. - Cloudflare OS — runs in the browser, nothing to install. Its agent works by writing code against its bindings, so expect two steps where the others take one, and a grant per server in each chat. What you get for that: every model call routes through the AI Gateway rather than only the leg a desktop client chooses to send.
Six of this walkthrough's eight steps end in something being refused, so how a client reports a refusal is not a detail — it is most of what the room sees. Goose prints the gateway's own reason, and in two steps the agent reads our policy's block message and explains the refusal in its own words. opencode shows a 424 and makes you narrate the rest.
A local VM or a cloud one, either is fine — but do use one. This is the recommendation that saves the most grief, and the reason is not tidiness.
Enrolling a device in this demo's Zero Trust organisation is not a small act.
wire-access.sh deliberately turns on install CA to system certificate store,
because Gateway cannot inspect HTTPS unless the device trusts its certificate — so the
Cloudflare One client puts this tenant's root certificate into your trust store, and from
then on the demo account can decrypt that device's traffic. On a disposable VM that is exactly what
you want. On your work laptop you have handed a demo tenant your browsing.
Three more reasons, each of which costs an afternoon the first time:
- The DLP policies apply to everything the device does, not just the demo. The rule that blocks a paste into Gemini does not know which browser tab is yours. Your own AI tooling stops working while you are enrolled.
- You have to be Alice. The device session is the credential for the model
endpoint, so the client has to be signed in as
alice.watson@company.com— which is not a thing you want your real machine to be. - Resetting between runs is a snapshot rather than an archaeology exercise. Client config, MCP OAuth tokens and chat history all accumulate, and a stale portal registration fails in a way that does not look like auth. A VM you can roll back makes that a non-problem.
A cloud VM has one extra advantage worth knowing: you can hand the whole thing to a customer or a colleague to drive, which is the difference between showing a demo and running a lab.
Cloudflare OS is the exception and needs none of this — it runs in a browser and authenticates through Access, so there is no device to enrol and no certificate to install. If you cannot make a VM, that is the client to use.
Client 1 — Goose (preferred)
Verified on Goose Desktop 1.53.0. Two things to add, in this order: a custom provider pointing at the AI Gateway, and a Streamable HTTP extension pointing at the portal.
- Add the provider. At first launch choose Connect to a provider →
+ Add a custom provider → Configure manually. Leave the
provider type as OpenAI Compatible, name it anything (Cloudflare does),
and set the API URL to
with/compaton the end. Leave the API base path empty. - List one model under available models:
@cf/google/gemma-4-26b-a4b-it. No API key — the Access application in front of the gateway accepts your Cloudflare One client's device session as the credential, which is why there is no key to paste. - Add the portal. Extensions →
+ Add custom extension, type Streamable HTTP, name it
Company AI Portal, endpoint
with/mcpon the end. Add extension. - Disable every other extension so the portal is the only one left enabled under Default Extensions. This matters as much here as the deny list does in opencode: Goose ships with developer tooling, and an agent that can read your filesystem will answer from it instead of from the company's systems.
- Send any prompt. You are redirected to the portal once, where you authorise the servers — and then never asked again.
This is the practical difference from Cloudflare OS, and worth knowing before you choose. Goose completes a single OAuth flow against the portal and can then call every server behind it, for the life of the token — no per-server grant, no per-chat grant, nothing to redo before a second run. Cloudflare OS instead issues a capability per server, scoped to the chat it was granted in.
Both are defensible, and the per-capability model is arguably the better security story. For presenting the same eight steps repeatedly, one login is the one you want.
Client 2 — opencode
Read this before choosing opencode, because it costs you something in every step that matters. When the gateway refuses a call, opencode reports the whole of it as:
Provider request failed with HTTP 424
No reason, no policy name, nothing to read out. The gateway does send one — "Request content blocked due to DLP policy violations" — and opencode discards it. Goose prints it:
Request failed with status 424 Failed Dependency at
https://aig.<your-zone>/compat/v1/chat/completions:
Request content blocked due to DLP policy violations.
Six of the walkthrough's eight steps end in something being refused, so this is not a rough edge in a corner — it is most of what the room sees. On opencode every one of those becomes you switching to the AI Gateway dashboard to explain what just happened, which works but turns a thirty-second beat into a two-minute one and moves the audience's attention off the transcript.
Nothing here is fixable from our side: the gateway sends the reason, the client drops it. If you are already set up on opencode it is perfectly usable, and the walkthrough tells you where to look for each block. If you are choosing now, choose Goose.
opencode runs on your machine, and that machine has to be enrolled in the Cloudflare One
client and authenticated as alice.watson@company.com. Not optional, and not
just for tidiness — it is what makes the configuration below work at all.
The device session is the credential. The Access application in front of
is configured to accept it, which is why the provider config has no API key in it: nothing in a file,
nothing in an environment variable, nothing on disk. Without the client, every model call is turned
away by Access before it reaches the gateway.
Signed in as the wrong person is worse than not signed in, because it looks like it works: the requests succeed and every one of them is attributed to whoever that is. The walkthrough's argument is that the agent acts with Alice's identity, so check the client's account before you present, not after.
The same session is what makes step 1 happen: the
browser-side DLP policy and the desktop notification explaining the block both come from this client.
On a machine without it, cloudflared can fetch a short-lived Access token instead, and
opencode can run that for you — but you lose step 1, so for a demo it is not a real
substitute.
One deploy-time caveat, because it fails the same way as a missing client. Accepting a device
session needs an account-wide Cloudflare One Client Authentication duration to exist first, and
wire-access.sh degrades rather than failing if it cannot set one — it prints
"Without client-session auth: clients authenticate with 'cloudflared access login' instead".
If you saw that, the flag is not on the application and no device session will authenticate, however
correctly the client is signed in. Turn it on under Zero Trust → Access controls →
Access settings and re-run.
Cloudflare OS needs none of this: it authenticates in the browser, so a signed-in Access session in any tab is enough.
Then two things to configure: a provider against AI Gateway's OpenAI-compatible endpoint, and an MCP server against the portal. There are two configuration shapes below, and do not pick between them by version number — that is the trap.
Verified on opencode Desktop 2.0.22: the 1.x configuration works and the 2.x one does not. The desktop app's version number is its own, and it does not tell you which config schema it reads — a "2.0.x" desktop build still wants the 1.x shape.
So: Desktop → the 1.x card. Reach for the 2.x card only if you are running
the 2.x CLI, where opencode --version does mean what it looks like:
opencode --version
Getting this wrong is quiet, which is why it is worth stating. 1.x accepts keys it does
not understand without complaining, so a 2.x config on a 1.x-shaped client validates, loads,
and does nothing: the agents block with its permissions, system prompt and step cap is
read and discarded. The first you know of it is an agent spending seventeen steps guessing at tool
names in front of an audience.
The one error it does raise is the good case, and worth memorising:
Configuration is invalid at ~/.config/opencode/opencode.json
V2 permissions are not supported by OpenCode V1.
Use V1 "permission" rules or run opencode2. agents.employee.permissions
Either way, confirm rather than assume: opencode agent list should show
employee with "permission": "execute" denied. And during a run, a single
Execute row in the transcript means the deny is not in force — nothing else needs
checking first.
opencode 1.x — and the Desktop apptop-level permission map · verified on Desktop 2.0.22 and on the 1.18.x CLI
In ~/.config/opencode/opencode.json.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"cf-ai-demo": {
"npm": "@ai-sdk/openai-compatible",
"name": "Cloudflare AI Gateway (demo)",
"options": {
"baseURL": "https://aig./compat"
},
"models": {
"workers-ai/@cf/google/gemma-4-26b-a4b-it": {
"name": "Google Gemma 4 (Workers AI)"
}
}
}
},
"mcp": {
"servers": {
"ai-demo": {
"type": "remote",
"url": "https://mcp./mcp",
"codemode": false
}
}
},
"permission": {
"execute": "deny",
"bash": "deny",
"edit": "deny",
"read": "deny",
"glob": "deny",
"grep": "deny",
"webfetch": "deny",
"websearch": "deny",
"subagent": "deny",
"skill": "deny"
}
}
opencode 2.x CLI onlyagents with a permissions array · rejected by Desktop 2.0.22
In ~/.config/opencode/opencode.json. Only for the 2.x CLI —
the Desktop app rejects this shape, including Desktop 2.0.22, so use the card above for it.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"cf-ai-demo": {
"npm": "@ai-sdk/openai-compatible",
"name": "Cloudflare AI Gateway (demo)",
"options": {
"baseURL": "https://aig./compat"
},
"models": {
"workers-ai/@cf/google/gemma-4-26b-a4b-it": {
"name": "Google Gemma 4 (Workers AI)"
}
}
}
},
"mcp": {
"servers": {
"ai-demo": {
"type": "remote",
"url": "https://mcp./mcp",
"codemode": false
}
}
},
"default_agent": "employee",
"agents": {
"employee": {
"description": "An ordinary employee's assistant - company tools only",
"mode": "primary",
"steps": 10,
"system": "You are the assistant of an employee at this company. Answer using only the company tools available to you. You have no filesystem, no shell and no access to the public internet. Call the tool that matches the question directly - do not search for tools and do not write code. If the tools cannot answer the question, say so plainly rather than guessing.",
"permissions": [
{ "action": "execute", "resource": "*", "effect": "deny" },
{ "action": "shell", "resource": "*", "effect": "deny" },
{ "action": "edit", "resource": "*", "effect": "deny" },
{ "action": "read", "resource": "*", "effect": "deny" },
{ "action": "glob", "resource": "*", "effect": "deny" },
{ "action": "grep", "resource": "*", "effect": "deny" },
{ "action": "webfetch", "resource": "*", "effect": "deny" },
{ "action": "websearch", "resource": "*", "effect": "deny" },
{ "action": "subagent", "resource": "*", "effect": "deny" },
{ "action": "skill", "resource": "*", "effect": "deny" }
]
}
}
}
Client 3 — Cloudflare OS
./deploy.sh deploys
Cloudflare OS
at , and the five apps carry a Company AI tile in their launcher
pointing at it. Set DEPLOY_CLOUDFLARE_OS="false" to skip it, in which case the tile is
left out rather than linking somewhere that does not resolve.
Worth knowing why you might prefer it: every model call it makes routes through the AI Gateway named in its config, so the DLP profiles and guardrails apply to every agent turn rather than only the leg a desktop client chooses to send, and it authenticates users through Access, so the agent runs as the signed-in person.
It needs three things done once per person, in the browser, and no API can do any of them — a model and a connector both belong to a user account and are configured over that user's own session:
- Add the model. Settings → Choose your model → Add new model →
Other Cloudflare Workers AI…, then the bare id
@cf/google/gemma-4-26b-a4b-it. No key is requested: gateway mode authenticates on the Workers AI binding. The id takes noworkers-ai/prefix here — that belongs to other clients' naming. - Connect the portal account. On
/gatekeepers, add the portal. It runs an OAuth flow against Access. This connects the account and grants no servers — see the next step, which is not on that page. - Grant the servers you need, one at a time. Open a chat, click the
+ in the message composer (the control labelled "Files, connections and skills"),
then Add a new connection. Choose Company AI Portal, keep the
account already connected, pick a Server, choose All tools, and
confirm. Repeat per server. The picker lists five similar invented product names, so pick by what
the step needs rather than by what sounds right:
- WorkBox MCP — inbox and calendar. Step 4 needs this one.
- Pipeline MCP — the CRM. Step 8 needs this one.
- WorkWeek MCP — HR. Only for the per-app scripts.
- Nexus MCP — the wiki. Only for the per-app scripts, and the one most easily picked by mistake when you wanted the calendar.
- Ledger MCP — finance. Do not grant it; Alice is not entitled to it and its absence is the point of step 8.
This is the one that makes everything else look broken, and no prompt or instruction fixes it. A granted MCP resource becomes a binding in the chat where you granted it. A new chat is seeded from gadgets and nothing else, so it starts with no portal binding at all — and the agent has no choice but to ask for one again, which is most of the steps you were trying to avoid.
So the pre-grant is not a global setting you do once. Open the chat you are going to demo in, grant WorkBox and Pipeline there, and run all eight steps in that same chat. Switching chats between steps puts you back to an eight-step step 4.
A tell-tale: a grant you made yourself from the composer is named after the system —
WORKBOX_MCP — and one the agent requested is named whatever the agent chose, so
COMPANY_PORTAL or COMPANY_AI_PORTAL. If you see the latter, the chat had
no grant and the agent asked for one. A name that changes between runs means a new grant each
time.
If you will run this more than once, there is a way out of the per-chat grant. A chat's starting
bindings come from the workspace's gadgets, so a gadget holding the portal bindings seeds
every chat after it. cloudflare-os-instructions.md has the recipe — grant the
servers, ask the agent to create a holder gadget and wire them in with
setGadgetBinding, then accept the changes, which is the step that
makes it permanent.
Two caveats. It is per workspace, so per person, like the model. And it is an
off-label use: setGadgetBinding says to use it only when a gadget's code needs the
resource, whereas we want the propagation side effect — which rests on how
defaultBindingList reads gadget bindings, an implementation detail rather than a
contract. Good enough for a demo; worth knowing before building a workshop on it.
If you are deciding whether Cloudflare OS can replace a desktop client, these are the numbers rather than an impression. Same prompt, step 4 of the walkthrough.
Granted in advance — two steps, about a minute. Measured:
describeBinding on the WorkBox binding, then one executeCode that reads
the calendar and answers. Note that the prompt asks for two search terms and the tool takes one
substring per call, and the agent still did both inside that single run — so two steps is the
floor and it is reachable, not theoretical.
Granted during the turn — eight steps, about two minutes. Observed:
an exploratory code run; describeBinding WORKBOX, which errors because the agent is
guessing a name; describeBinding GIT, a Cloudflare OS built-in that has nothing to do
with anything; listConnectableResources github, a guess; then the real
listConnectableResources mcp_portal, requestConnection, your acceptance,
a describeBinding on the new binding, and finally the code that answers.
Six of those eight steps exist only because the capability was missing when the turn started. The agent is not being slow, it is shopping for a tool it has not been given — so the fix is the pre-grant, not a better prompt.
opencode, for comparison: one step. A native tool call.
Two of the eight are irreducible; the rest are worth attacking.
describeBinding plus executeCode is the architecture — the agent has
no per-tool functions at all, so writing code against a binding is how it calls a tool, and
no setting changes that. Everything above those two is the agent not knowing which binding serves
which system, and that is fixable: see the next card.
Cloudflare OS takes deployment-wide agent instructions, appended to the system
prompt for every user. Signed in as admin@company.com, go to /admin and
paste the block from cloudflare-os-instructions.md in the repository root. One action,
once, for everybody who uses the deployment — unlike the model, which each person adds to
their own account.
It is worth the two minutes because it targets the wasted steps directly. The block tells the agent not to invent binding names, that this deployment has no GitHub integration, that query strings are single substrings and limits cap at 25, and that a policy refusal means stop rather than try another system. Every line of it is one detour observed on a live run.
It also carries the tool API itself, which is the part that removes discovery
rather than shortening it. A portal binding exposes each tool as a camelCase method keeping the
portal's server prefix — work_list_company_calendar becomes
workListCompanyCalendar — and the block lists the dozen the demo uses with their
parameters, plus a worked executeCode example. An agent that already knows the method
and its arguments has no reason to call describeBinding at all, which is the difference
between two steps and one.
The limit is 8000 characters and the block is about 4100, so there is room to add anything your own audience keeps tripping over.
Worth knowing what it cannot do: the two irreducible steps stay, and instructions are guidance rather than enforcement — a model can still ignore them. The controls in this demo are deliberately not of that kind, which is a contrast worth drawing if someone asks why you do not simply instruct the agent not to leak things.
If you skip the grants and let the agent ask mid-chat, it calls
requestConnection and you get an accept card. Accept it and the grant is made —
but the server it ends up scoped to may not be the one the card describes. A card reading
"I need access to the Company AI Portal to search the WorkBox calendar" has been observed producing
a binding onto Nexus, the wiki.
That failure is expensive and does not look like a failure. The agent holds a binding it believes
is the calendar, so it writes code against it, finds nothing, assumes it searched wrong, and tries
again — eight or ten executeCode runs, no answer, and a room watching a progress
spinner. Nothing in the transcript says "wrong system".
How to check in two seconds: every code row names the server it read, after the
portal's name — Company AI Portal / Nexus MCP. If that is not the app your prompt
named, the binding is wrong and no amount of waiting will fix it. Grant the right server from the
composer, where you pick it yourself.
Which is the real reason to grant before you start, rather than politeness about not interrupting the demo.
Expect to grant each server, and do not read it as friction to be worked around. It is a real difference between the two clients, and it is architectural rather than an oversight.
opencode gets one broad authorisation. It completes a single OAuth flow against the portal and can then call anything the portal offers that person — every server, every tool, for as long as the token lives. One login, total reach.
Cloudflare OS holds narrow capabilities instead. Each grant becomes a
binding in the agent's environment, scoped to one upstream server and named after it
(MCP_…_WORK, MCP_…_CRM), and the agent can reach only what it
has a binding for. Not what it is permitted to call — what it holds. Granting WorkBox
gets it no closer to Ledger, and no amount of rephrasing changes that, because there is nothing
there to call.
Which is worth saying out loud during the demo rather than apologising for beforehand. This whole walkthrough argues that the agent-to-tool path is where controls belong; a client that makes you name each system the agent may reach, and that shows the agent asking when it has not been given one, is that argument implemented. The Access policy on Ledger and the per-server grant are the same idea at two layers.
Practically: the walkthrough needs WorkBox and Pipeline, so two passes before you start. Leave Ledger alone — the portal will not list it for Alice anyway, because her Access session is not entitled to it, which is step 8's point and the cheapest control in the demo.
Skip that and the agent meets the portal for the first time mid-demo. It has no connection, so it
reasons about that at length, calls listConnectableResources, and asks you to set one up
— interesting once, and a poor way to open a walkthrough whose argument depends on the tools
being unremarkable.
Running without the protection layer first
If the suite was deployed with DEPLOY_PROTECTION=false there is no portal and no AI
Gateway yet. Point the client at the MCP servers individually and at the model provider directly:
{
"mcp": {
"servers": {
"workweek": { "type": "remote", "url": "https://hr-mcp./mcp" },
"pipeline": { "type": "remote", "url": "https://crm-mcp./mcp" },
"relay": { "type": "remote", "url": "https://work-mcp./mcp" },
"nexus": { "type": "remote", "url": "https://wiki-mcp./mcp" },
"ledger": { "type": "remote", "url": "https://fin-mcp./mcp" }
}
}
}
Each server runs its own OAuth 2.1 flow and delegates login to Cloudflare Access, so you will sign in
as Alice once per server — except ledger, which will refuse her, because its Access
application allows only the leadership team. Every demo script works in this mode
— that is the "before" half of each one.
The portal namespaces every tool with its server id, so list_employees becomes
hr_list_employees, get_pipeline_summary becomes
crm_get_pipeline_summary, and so on with work_, wiki_ and
fin_. The
demo scripts name the underlying tool; your transcript will show the prefixed one.
Check it works
Before running any script, ask the agent something harmless that proves both legs are live:
Who am I, and which tools do you have available?
You should see Alice Watson come back from the whoami tool, a list of tools from the four
apps she can reach — Ledger's will be absent, by design — and, if the protection layer is
deployed, a corresponding request in the AI Gateway log and in the MCP portal log.