AI Agent Tool Gateway: Who Holds the Keys When One Token Calls Thousands of Paid Tools
One token can reach thousands of paid tool endpoints. Register who holds each key, run the --local check, cap per-agent spend and reconcile the meter monthly.
Go deeper. Build your own.
treg run stripe -- get /v1/balance reads like a call that keeps your Stripe key on somebody’s server. By default it runs the Stripe CLI on your own machine with the organization’s credential injected. The server-side version needs one extra flag, and the README’s own example comment spells out the difference: with --server, “the key never reaches you.”
That flag is the whole custody question for an AI agent tool gateway in miniature. One token in front of thousands of paid endpoints is convenient: no provider signups, a price before every call, one bill. It also means one place holds the keys and one meter counts the spend, and the place depends on how each call is made.
This playbook gives you a tool-key custody and per-call spend register, one row per endpoint class, with the key holder, revoke path, cap, allowed agents, spend owner and evidence filled in. It adds the custody rule, a --local check you can run this afternoon, and a monthly reconciliation of the call log against the gateway’s metered total.
treg, the “OpenRouter for agent tools,” as of Oct 7
The hook is treg, a repository from Superdesign whose GitHub description reads “OpenRouter for agent tools.” The repo was created on Jul 15, 2026 and had about 4.8K stars by early Oct 8 (4,761 on the GitHub API at 03:37 UTC Oct 8). It is a tool-endpoint gateway, not a model proxy: the README promises “a curated catalog of thousands of endpoints across many providers,” from SEO and backlinks to enrichment, ads and scraping, “priced per call, from a cent,” with no provider signup, plus “your own team’s keys, skills and CLIs, callable by every teammate’s agent.”
Screenshot: GitHub, “superdesigndev/treg: OpenRouter for agent tools” (repository page, live counters), captured Oct 7, 2026.
It is self-hostable and also hosted at treg.to. The catalog size depends on who is counting. The live platform catalog summed to 3,401 endpoints across 110 providers when read at 03:37 UTC Oct 8, while treg’s own llms.txt says 3,800+ endpoints across 110 providers. Use the live number for planning and the marketing number for nothing.
The licence is unclear. GitHub reports it as NOASSERTION because the LICENSE file is Apache-2.0 text plus additional terms that take precedence and bar offering the code as a hosted service, managed registry or embedded component sold to third parties without written authorization; internal use inside your own organization is expressly permitted. That is not an OSI-approved licence as written, so do not file it as “Apache-2.0” in your vendor inventory.
Now the correction. The README’s one rule is that the proxy “injects auth server-side,” and for relayed /call requests, treg call, treg run --server and MCP that holds. But for vendor CLIs, “--local (default) runs on your machine; --server runs on the registry and streams output back.” The README also says treg run executes the CLI “with the org’s credential injected, so you never hold the key.”
Both sentences are true at once: you never typed or saw the key, and the key was present on your machine for the length of the run. “Credentials stay server-side” holds only for server-routed calls and MCP.
Screenshot: GitHub, “superdesigndev/treg” README, section “2. CLIs” (repository page), captured Oct 7, 2026.
Why per-call tools move the key closer to the agent
A model gateway holds one or two provider keys and a token budget. A tool gateway holds a key per provider, often dozens, and the agents decide at run time which ones to spend. Every call is a purchase made by software, against a key somebody else may hold, through an execution path the agent may choose.
That changes two questions from “set once” to “check per row.” Where does this key live at the moment of the call? And how much can this agent spend through it before a person notices? The register answers both.
Keep the tool-key custody and per-call spend register
Run the steps in order. The diagram below shows the path every call takes, and where it forks between server-side and local custody.
The fork in the middle is the custody decision. Default every lane to the server-side branch.
Step 1: Sort every endpoint class by whose key it uses
Gateways of this kind carry two kinds of key. The gateway’s own provider accounts, billed to you per call from a prepaid balance, and your keys, which you register and the gateway relays. treg’s README is explicit that “your own key always wins over treg’s, and those calls are never metered.” So the same endpoint class can be metered in one team and free in another, and the register has to say which.
A sales team that already pays for an enrichment provider registers its own key and the gateway relays it unmetered; a team without that subscription calls the same class on the gateway’s account and pays per call. Same endpoint, different custody, different meter.
Step 2: Apply the custody rule to every row
The custody rule: a key an agent can print is a key the agent holds. Three locations cover nearly every case:
- Gateway server. The key is injected on the gateway and never reaches the agent’s machine. Relayed calls, server-mode CLI runs and MCP land here.
- Local CLI. The gateway hands the credential to a CLI process on the agent’s machine for the run. treg’s default
treg runlands here. Whether the agent can read it depends on how the CLI receives it, which the README does not spell out, so assume yes until you test. - Agent environment. The key sits in an environment variable, dotfile or config the agent can read. Treat it as held by every process on that machine.
Revoke speed and blast radius differ by location, which is why custody is the first column after the endpoint class.
Step 3: Fill the register
The table is the artifact. The rows are endpoint classes a marketing-and-sales agent fleet commonly reaches through a tool gateway. Every price, cap and timing is illustrative, not any vendor’s price list; the column set is what matters.
| Endpoint class | Who holds the key | Revoke path and time | Per-call price · per-agent daily cap | Agents allowed | Spend owner | Evidence |
|---|---|---|---|---|---|---|
| SEO / SERP data | Gateway server (gateway’s account, metered) | Revoke the agent’s gateway token, 1 min | $0.01 · 300 calls ($3.00) | research-agent, content-agent | Marketing ops lead | Gateway call log: agent, endpoint, cost, cache hit |
| People and company enrichment | Gateway server (gateway’s account, metered) | Revoke the agent’s gateway token, 1 min | $0.05 · 100 calls ($5.00) | sales-agent | Revenue ops lead | Gateway call log plus monthly CSV export |
| Scraping | Gateway server (gateway’s account, metered) | Revoke the agent’s gateway token, 1 min | $0.002 per page · 2,000 pages ($4.00) | research-agent | Platform owner | Gateway call log |
| Ads (own ad account) | Gateway server (your key, relayed, not metered) | Rotate at the ad platform, 15 min; deleting the gateway connection does not revoke it | Platform spend, not per call · writes need approval | ads-agent (read), writes approved by a person | Paid media lead | Gateway log plus ad platform change history |
| Email sending (own provider CLI) | Local CLI today; target: gateway server via --server |
Rotate at the email provider, 10 min, then check the local machine | Provider’s per-send price · 200 sends | outreach-agent | Growth lead | treg runs plus provider send log |
| Payments (own provider CLI) | Agent environment today; target: gateway server, restricted read-only key | Roll the key in the provider dashboard, 5 min | No per-call price · reads only, zero writes | finance-agent (read) | Finance controller | treg runs plus provider request log |
Two cells carry most of the risk: the two rows still marked local CLI or agent environment. Each gets a target and a date, and the payments row gets a restricted read-only key before anything else changes. First-party model-provider keys, the ones your agents use to reach the models themselves, do not belong here; their migration and rotation drill is the fleet API key migration drill.
Step 4: Run the --local check
Find every place an agent or script invokes a gateway CLI without forcing server-side execution. A starting sweep, illustrative; adjust paths to where your agents keep scripts and skills:
grep -rnE "treg run " ./agents ./scripts ~/.claude 2>/dev/null | grep -v -- "--server"
grep -rn "treg shell start" ./agents ./scripts 2>/dev/null
The first line lists CLI runs that will take the local default. The second finds whole sessions where every registered CLI injects locally, which is the same custody decision made once for an entire shell. Neither catches a teammate who typed the command by hand an hour ago, so pair the sweep with the gateway’s own run log: treg runs records each CLI execution, and any entry for an agent token without a matching server-side flag is a row to fix.
Then test custody directly, from inside the agent’s own sandbox and as the agent: run one harmless local CLI command through the gateway and, in the same session, list the environment and the CLI’s config files. If any key-shaped string appears, the row’s custody is local CLI or agent environment, whatever the vendor page says. Rewrite each hit to the server-side form, treg run --server <cli> -- <args>, or move the call to a relayed endpoint. treg’s usage guide also notes that the background proxy service (treg serve start) writes its own access token to a file on disk, owner-readable only, where treg shell --proxy keeps it in the subshell’s environment; that token is the proxy’s, never a vendor key, but it belongs in the evidence column too.
Step 5: Set the caps where they actually stop spending
Write down which limit truly blocks, because not all of them do. treg’s llms.txt is blunt about its own budget caps: “These caps are advisory — your balance is the hard limit.” Its usage guide adds a per-day ceiling on spend against treg’s keys “so a runaway agent has a bounded blast radius,” an X-Treg-Route-Max-Cost header that caps a routed call (default ceiling $1), and two status codes agents must treat as stop signals:
- 402 when the balance runs out, carrying
balance_micro,estimated_cost_microand atopup_url. An agent should stop and report, never top up. - 503
provider_capacity_unavailablewhen the gateway’s own provider account is exhausted. Nothing is charged and the body namesresets_atwhen known, so allow one delayed retry and then stop.
For a per-agent cap, issue each agent its own gateway token and a tool allowlist; treg scopes tokens to one member of one organization, with roles (owner, admin, member, viewer) and per-member tool access via treg org access <member> --tools a,b. Keep the prepaid balance small enough that the hard limit is also an acceptable worst day.
Step 6: Write the revoke paths before you need them
There are two revocations per row, and they are not interchangeable. Revoking the agent’s gateway token stops that agent in about a minute. It does nothing to a provider key that leaked, because that key stays valid at the provider until the provider revokes it.
Treat deleting a connection inside the gateway as cleanup, not revocation. For every own-key row, the register names the provider console, the person with rights in it and the minutes it takes.
Step 7: Reconcile the call log against the metered total every month
The gateway’s metered total is a claim. Your call log, keyed by agent token, is the evidence. Once a month, sum the log by endpoint class and compare it with what the gateway charged; gateway chargeback and invoice reconciliation covers the mechanics of turning that into per-team charges.
Worked scenario: one month, three metered classes, one gap
Take the register above and a 30-day month, all numbers illustrative. A typical day spends $1.80 on SEO data (180 calls), $2.50 on enrichment (50 calls) and $2.00 on scraping (1,000 pages), each well under its cap. The three own-key rows cost the gateway nothing.
Illustrative: typical-day spend against the daily cap per endpoint class. Own-key rows never touch the meter.
Over the month the log sums to $54.00 for SEO, $75.00 for enrichment and $60.00 for scraping: $189.00. The gateway’s metered total says $198.40. The $9.40 gap is 5%, over a 2% tolerance, so somebody looks before the next top-up.
In this scenario two causes account for it: a batch of enrichment calls served through the gateway’s overflow relay, which treg’s usage guide says carries the relay’s real price and an X-Treg-Served-Via header, and a teammate’s local test token that never got a register row. The first gets a note and a decision on whether to turn overflow off; the second gets revoked.
Where tool-gateway key custody quietly fails
| What breaks | Signal you would see | First action |
|---|---|---|
| A “server-side” key is actually local | The --local sweep finds treg run without --server, or a key-shaped string in the agent’s environment |
Rewrite to --server or a relayed call, then rotate that key at the provider |
| The gateway itself is breached | A vendor bulletin, or calls in your log from tokens you never issued | Revoke every gateway token, then rotate every own-key row at its provider in register order |
| A runaway loop drains the balance | 402 responses clustered on one agent token; one class spikes against its cap | Revoke that agent’s token; lower the prepaid balance until the loop is fixed |
| An agent retries a 503 in a tight loop | Repeated provider_capacity_unavailable from one token within a minute |
Change the agent’s stop rule to one delayed retry, then report |
| Spend lands on an unregistered token | A token in the gateway’s call log with no row or owner in the register | Revoke it the same day; add a row only if someone claims it with a reason |
| Metered total drifts from the log | Monthly gap over tolerance; responses carrying an overflow header | Reconcile by token and class; decide whether overflow stays on |
| Telemetry leaves the machine unannounced | A CLI used with a hosted gateway sends usage events | treg documents an opt-out (TREG_TELEMETRY=0 or DO_NOT_TRACK=1); set it fleet-wide if policy requires |
Custody belongs on the same map as the agents
A tool gateway is one more layer between your agents and the world, and it should sit on the same control map as the rest. The agent gateway control plane describes that layer in general; this register is its key-and-meter page. The opposite design, tools that need no stored key at all, is the keyless MCP trust tier, and when you are the provider on the other side handing scopes to callers’ agents, the personal agent front door is the companion piece.
The register earns its keep on the bad day. When a bulletin like Composio’s lands, the question is which doors to re-lock and in what order, and the answer should take ten minutes because it is already written down. Seeing which agent spent what through which key, next to every session the fleet is running, is the job of a multi-agent command center.
FAQ
What is an AI agent tool gateway?
A service that gives agents one token and endpoint for many third-party tools, such as SEO data, enrichment, scraping or payments APIs. It injects the provider keys, meters each call and logs it. It differs from a model proxy: it fronts tool endpoints, not language models, and often holds dozens of provider keys.
Is treg open source?
Its licence is unclear. GitHub reports NOASSERTION because the LICENSE file combines Apache-2.0 text with additional terms that bar offering the code as a hosted service or managed registry to third parties without written authorization. Internal use inside your own organization is expressly permitted. Check the current file before relying on it.
How do you cap pay per call API spend for agents?
Give each agent its own gateway token and tool allowlist, set a per-agent daily cap per endpoint class, and keep the prepaid balance small enough to be an acceptable worst day. Treat 402 as stop, allow one delayed retry on 503, and reconcile the call log against the metered total monthly.
Sources
- GitHub: superdesigndev/treg repository
- treg README (credential ladder, –local default)
- treg USAGE guide (402, per-day ceiling, 503, overflow relay)
- treg LICENSE (“tools-registry License”)
- treg.to llms.txt (catalog claim, budgets advisory)
- treg.to live platform catalog
- Composio: May 2026 security incident bulletin (updated Sep 16, 2026)
- Arcade: Smithery joins Arcade (Aug 5, 2026)
- Business Wire (Arcade Series A release, Jun 15, 2026)
