Keyless by Design: Authentication and Authorisation for an AI Gateway
Keyless by Design: Authentication and Authorisation for an AI Gateway
This is the second post in the series. The first one made the case for putting one door in front of every model an organisation runs, and closed on a design decision worth unpacking properly: there are no API keys anywhere in the request path, not on the backends and not even APIM subscription keys on the front door.
This post covers how that works. Every request the gateway handles has to answer three questions, in order:
- Who is calling? That is authentication.
- What are they allowed to do? That is authorisation.
- How does the gateway prove itself to the model? That is backend authentication.
The standard way to teach APIM security answers those three with three layers, and the first layer is almost always a subscription key. This setup diverges on that first point, so that is where we start.
Why Not Subscription Keys
A subscription key is APIM's standard credential. The caller puts it in an Ocp-Apim-Subscription-Key header. APIM checks it against a product subscription and lets the request through, attributing it to a consumer.
The problem is what that actually proves. A subscription key answers "who is calling?" only in the weakest sense: it proves the caller holds a copy of a shared secret. It does not prove which caller, because keys get copied, which is the failure mode described in the first post. It carries no tenant, no expiry, no signature, and no statement of what the holder is entitled to. It identifies, but it does not authorise.
For an AI gateway, where the resource behind the door is expensive and contended, "someone has a copy of the key" is not a strong enough answer. So the layer is removed entirely. There is no product, no subscription, and no front-door key. Every question about identity is answered by Entra ID instead.
Subscription keys remain a perfectly valid approach, and plenty of teams run them well: each consuming team gets its own key, product and backend. The decision here was to go against the APIM default, and the rest of this post is the reasoning.
Microsoft's own Well-Architected Framework for API Management recommends the same direction, as does the Azure security baseline for API Management.
Layer 1: Authentication — Entra ID at the Door
The first thing every request meets after the network boundary is a single policy, validate-azure-ad-token, and it does considerably more than a subscription key check ever could.
A calling application authenticates to Entra ID first, using the client-credentials flow: app to app, no user in the loop. It receives a JWT and presents it to the gateway as a normal Bearer credential. On every request the gateway checks that the token came from the expected tenant and was issued for the expected audience. It also checks the token has not expired and carries a valid signature.
The important change is what the client now holds. There is no shared secret in the request path. The client holds its own Entra credentials and mints its own short-lived tokens, so there is nothing for the gateway operator to distribute and nothing for the consumer to rotate on the gateway's behalf.
When the token validates, the gateway reads the azp claim and stores it in a variable. That claim is the verified client application ID. That variable is the single source of truth for who is calling. Every downstream policy reads it rather than re-parsing the header, which matters for a reason covered later in this post.
Two details are worth stating because they decide real behaviour. Tokens issued under the v1.0 endpoint carry appid rather than azp, so the gateway reads either, which means a managed identity fetching a token through IMDS still keys to its own identity. And a token carrying neither claim is refused with a 403 rather than allowed through. An empty identity would place every such caller into one shared bucket, where they would share each other's limits and could be served each other's cached completions while still receiving a 200. Refusing is the safer default.
Layer 2: Authorisation — Admission by Role, Limits by Config
Authentication establishes who. Authorisation is the separate question of what they are allowed to do, and this is where the keyless model earns its place, because the answer travels inside the token the caller has already presented.
The gateway's app registration defines exactly one application app role, AI.Gateway.Standard. It answers a single question: is this identity an approved workload permitted to reach the gateway at all? When a token arrives, the same validate-azure-ad-token policy inspects its roles claim and requires that role to be present. A token without it is rejected with a 401.
What the role deliberately does not carry is just as important: no tier, no limits, and no model rights. Those come from configuration, as a set of tier presets, with an onboarding registry for per-team differentiation. The limits themselves are keyed on the caller's verified application ID rather than on any role they hold.
That separation is the design decision I would most defend. Admission is a directory concern. It changes rarely and granting it requires directory privilege. Consumption is a repository concern. Limits change often and belong in a reviewed pull request. Encoding tiers as app roles welds the two together, and every rate-limit adjustment then becomes a directory change.
This layer also produced the most surprising finding of the build. APIM's traditional tiering mechanism is products, and it never fires for a keyless caller at all. Product-scope policies only execute when a subscription key is present. With the key gone, the entire tier apparatus has to live in API-scope policy instead. Anyone going keyless while still reaching for products will wonder why none of their tier limits apply.
Onboarding a team follows from the same split. The team brings its own managed identity or service principal, and it is assigned the admission role on the gateway's app registration. That single assignment is the onboarding: from then on the role appears in the roles claim of every token they present, and the gateway admits them. Which limits they receive is a separate decision recorded in the registry, so raising a team's ceiling never requires a directory change. Offboarding is deleting the assignment, and nothing else moves.
A later post will cover team onboarding in its own right.
Layer 3: Backend Authentication — Managed Identity
The caller is authenticated and admitted. The gateway now has to call the model, which is the third question, and this is the most straightforward of the three layers.
The answer is a managed identity, which is also Microsoft's recommended way to authenticate an AI gateway to its models. After the caller is admitted and their limits are applied, the gateway requests a short-lived token for its own system-assigned identity. That token is scoped to https://cognitiveservices.azure.com and replaces the caller's in the Authorization header.
From the model's point of view the request arrives as the gateway, proven by a token issued moments earlier. There is no secret anywhere in the exchange: nothing is stored and nothing rotates.
On the model side that identity is granted exactly the Azure RBAC role it needs and no more. For Azure AI Foundry inference it holds Cognitive Services OpenAI User, the least-privilege role for calling deployments. The other services behind the same door get Cognitive Services User instead. That covers Speech, Language, Document Intelligence and Content Safety. The backends themselves have key authentication disabled outright and no public network access. The only route to a model runs through the gateway over a private endpoint, as the gateway's own identity.
One distinction is worth drawing clearly, because the two role systems are easy to conflate. The app role from Layer 2 authorises the client calling the gateway and rides inside the token's roles claim.
The Azure RBAC role here authorises the gateway's managed identity calling the model and never appears in a token at all. The two never cross over. A client never receives Azure RBAC on the model, and the gateway's identity never receives an app role.
The Perimeter: Keyless Is Not the Same as Open
"No keys" is easily misread as "no boundary", so it is worth being precise. The three layers above are identity. Underneath them sits the network perimeter.
The gateway is VNet-injected, which supports two modes. External mode gives a public front door, gated by every policy check described above. Internal mode removes the public endpoint entirely and is fronted with something like Application Gateway or Azure Front Door.
In front of the JWT check, an IP allow-list turns away anything outside the expected egress ranges before the token is examined at all, which makes it the cheapest possible rejection and the right thing to run first. The backends sit behind private endpoints with private DNS and public network access switched off.
Order Matters: The Shape of the Policy Chain
Layers are a useful way to think about the design, but on the wire it is one ordered chain, and the order is load-bearing. Cheapest checks first, identity before spend, and two orderings that exist for reasons which only become obvious once you have got them wrong.
The caller's JWT is validated once, at the front, and the caller's application ID is stored in a variable at that moment rather than re-read from the header later. That is a correctness requirement rather than an optimisation. By the time the request reaches the backend, the managed-identity step has overwritten the Authorization header with the gateway's own token, which carries no azp claim. Any policy attempting to derive the caller's identity from the header downstream would be reading the gateway's identity instead of the client's.
The second ordering has the same character. The content-safety check and the model call both travel on the Authorization header, so the managed-identity swap has to happen in a specific relationship to both, and content safety has to run before the semantic cache is consulted. Otherwise a cached answer becomes a way to skip screening entirely. The sequence is itself the security posture. Shuffle it and the same policies end up enforcing a weaker guarantee.
The chain below is in the order it actually executes. Each stage carries the status code it returns when it refuses, so every red line is a request that never reached a model and never cost a token.

Stage four is the one most people expect to find at the end: the managed identity is injected partway through the chain rather than just before the backend call.
Putting the Three Layers Together
The diagram below follows a single request through every identity it touches. The numbers mark the order, and the caller never holds anything beyond its own Entra credentials.

In words: a production application mints a client-credentials token from Entra ID and calls the gateway. The IP filter confirms the request came from an expected range. validate-azure-ad-token checks the token came from the right tenant and was issued for the right audience, has not expired and carries a valid signature. It then reads the roles claim and confirms the caller holds the admission role. The caller's verified application ID is captured for everything that follows, and the configured rate and token limits are applied against that identity.
The gateway then exchanges the caller's token for its own managed-identity token, the prompt is screened by content safety, and the request goes to Azure AI Foundry over a private endpoint as the gateway itself. The model, which trusts nothing else, answers.
No key was involved at any step, and every log line names a cryptographically validated caller.
That is the architecture worth having in place before production rather than as a hardening pass promised afterwards: authentication that proves identity instead of possession, authorisation that lives where access reviews already look, and a backend that trusts exactly one caller by construction.
What's Next in This Series
With identity settled, the next post takes apart the problem the gateway was really built to solve: carving one contended pool of model capacity fairly across many teams, through rate limits, tokens-per-minute ceilings and longer-period quotas, all keyed on the validated identity established here.
- Keyless by design (you are here) — Entra ID end to end: JWT validation, a single admission role, managed identity to the backend
- Carving up the token pool — rate limits, tokens-per-minute, and quotas across many teams
- Guardrails at the gate — content safety and Prompt Shield on prompts and completions
- Answering twice, paying once — semantic caching with Azure Managed Redis
- Staying up — backend pools, circuit breakers, and zone-redundant Premium
- Who spent what — token metrics, chargeback, and the observability stack
- Moving in — landing-zone adoption with bring-your-own everything
The Takeaway
The default way to secure a model endpoint is a key, and a key answers the wrong question. It proves possession rather than identity, and possession is precisely the thing that leaks.
Trading it for Entra ID collapses three separate problems into one primitive. Authentication becomes a signed token from your own tenant. Authorisation splits cleanly in two. Admission is governed by the access reviews you already run. Consumption limits live in reviewed configuration. Backend trust becomes a managed identity holding a least-privilege role with no secret to lose.
Keyless is not less security. It is the same three layers with the weakest link, the shared secret, removed from all of them. The whole setup is open source and deploys with one apply: okaneconnor/terraform-azurerm-ai-gateway.