Give your agent an API
Agents act on the world through tools. When the thing you need is behind an
HTTP API (your own backend, an AWS Lambda function URL, a Cloud Run service,
a vendor API), register it as an endpoint and your tool extensions get one call,
endpoint_invoke, that reaches it:
Register an endpoint
That is all an endpoint is: a name, a base URL and a key. The two
values are stored as agent secrets named after the endpoint,
ENDPOINT_BILLING_URL and ENDPOINT_BILLING_KEY, so the desktop's secrets
tab shows them like any other secret and you can set them there instead. The
agent picks up a new or changed endpoint at its next restart. --key-stdin
reads the key from standard input, keeping it out of your shell history;
--key none registers an API that needs no key. Replacing an endpoint
briefly unregisters it (the URL is removed, the key written, the URL
written), so a failure part-way leaves it unregistered rather than pointing
a key at the wrong host; rerun the command to finish.
What the agent can and cannot reach
Every call goes to the base URL plus the path the agent gives: /invoices,
/invoices/123?expand=lines. A path is checked to stay under the base (no
other host, no ..), redirects are not followed, and the key is sent only
there. There is no free URL anywhere in the tool, so a prompt that tries to
send the key elsewhere has nothing to hold on to. To expose only part of an
API, register a deeper base: https://api.example.com/v1/invoices makes
everything above it unreachable.
A call may add its own headers (headers: { "Accept": "text/csv", "X-Tenant": "acme" });
the ones that carry the key, the caller identity or the transport are set by
the platform and refused.
The key travels on both Authorization: Bearer <key> and x-api-key: <key>,
the two headers APIs commonly read; use whichever your API expects and
ignore the other. Who asked arrives as x-mutiro-agent, x-mutiro-user,
x-mutiro-role (owner or user) and x-mutiro-conversation; trust them only
on requests that carried your key.
A call must answer within 25 seconds and responses are cut at 1 MiB. An HTTP
error status is returned to the agent as success: false with your API's own
message (an error or message field in a JSON body), so answer errors that
way. Longer work is started by one call and collected by another.
Calls are not retried by the platform. Whether a failed call is safe to
repeat depends on the API and the operation, so the caller decides: a tool
extension in front of the endpoint can retry a read, and send its own
Idempotency-Key header (via headers) when it retries a create.
The model never calls it directly
endpoint_invoke is reserved for tool extensions
and hooks, always: it is not offered to the model and no setting opens it.
A raw HTTP client is a poor tool for a model (every API has its own shapes)
and the one place where a prompt could drive arbitrary operations on your
API with data from a conversation. So a developer puts a verb in front of
it:
The model sees create_invoice with validation and a stable result, never
the HTTP. A developer who wants the raw call available can write an
extension that passes its arguments straight through; that file in the repo
is the deliberate decision, reviewed and tested like any other.
Who may call the verbs. Extensions run with the tools of whoever
triggered them, so a user of the agent can reach an endpoint through any
extension the user may call. Mark an extension owner_only when the API
behind it should only act for you.
Serverless receivers
An AWS Lambda function URL with auth type NONE, or API Gateway with
an API key, works as is: read the key from x-api-key (API Gateway does this
for you) or Authorization. A Cloud Run or Cloud Functions service
with unauthenticated invocations works the same way with a shared key it
checks itself; an IAM-only service is not supported yet, since it needs OIDC
tokens rather than a key.