Industry Analysis

Inference Hooks in Claude Enterprise: What Changes for MCP and AI Agent Security

INS Security Team
6 min read

Until recently, inline control of AI agents meant putting something in the path: a network proxy, a browser extension, or an agent on every device. Since August 2026, Claude Enterprise can do part of that job itself. Anthropic's new Inference hooks make the platform call your policy server before a request reaches the model, and that changes where the real work in AI security sits.

What Inference hooks are

Inference hooks are a beta feature of Claude Enterprise. Before each governed request, Anthropic sends the conversation transcript to an HTTPS server run by your organization or your security vendor and waits for a verdict: allow or deny. A denied request never reaches the model, and the user sees a block message that combines the reason your server returned with a standing message set by administrators.

The check runs on Anthropic's side, so nothing is installed on devices. One hook covers claude.ai, Cowork, Claude Code and Claude Tag, on web, desktop, mobile, CLI and Slack. There are two events: prompt, which fires before each request, and tool_call, which fires before tools run if the “Validate tool calls” option is on. Tool results are not checked separately; they arrive in the next prompt event, before the model sees them.

What they cover out of the box

The transcript is passed in full: message text, tool calls with their arguments, tool results, and text extracted from attachments. That is enough for the controls most teams want first:

  • DLP on prompts, pasted content and attachment text.
  • Detection of prompt injection that enters through tool results, caught at the next prompt event.
  • Policy by MCP server: for third-party tools, the tool_call event carries the server's origin and a verified_origin flag.
  • One policy across every Claude surface, including mobile, with shadow mode, percentage rollout and role exclusions for a gradual start.
  • Every denial is recorded in the Activity Feed of the Compliance API.

What they do not do

This is the part to read twice before calling the problem solved. Per the documentation:

  • No tool descriptions or system prompts. Neither is sent to the hook. Tool poisoning, where malicious instructions sit in a tool's description, and rug pulls, where a description changes silently after approval, live in exactly the data the hook never receives.
  • Allow or deny only. There is no redaction or masking. A blocked message stays in the conversation and is resent with every later message, so the user has to edit or rewind it away. That is worse than transparent masking.
  • One verdict per response. A single verdict covers all tool calls in a response, so one bad call cannot be blocked on its own.
  • Partial content coverage. Raw file and image bytes are never sent, so image-only content is not inspected. Voice mode is not covered.
  • Claude Enterprise only. Hooks are not available on Amazon Bedrock or Google Cloud, and API access through the Claude Platform is out of scope. Other assistants and in-house agents are out of scope.
  • Limited origin signal. verified_origin is true only when Anthropic's servers connect to the MCP server themselves, and a missing origin means unknown, not safe. Hooks say who provides a tool, not who runs it.
  • Operational limits. The default verdict timeout is 5 seconds (configurable up to 10 seconds). If your server fails, a setting decides between blocking and allowing uninspected, and sustained failures trip a circuit breaker.

The feature is in beta, and Anthropic states that field names, request shapes and headers may change. Build against it with that in mind.

What this means for the market

Hooks do not replace third-party security tooling. They cover part of the scenarios and change the balance of power.

Being in the traffic path used to be a differentiator, and many products competed on where they could intercept: proxy, endpoint agent, browser. For Claude Enterprise, the platform now provides that position to anyone who can host an endpoint. The point of interception becomes a commodity, and the barrier for DLP and AI-security vendors drops sharply.

What follows is a forecast, not something Anthropic has said. Competition will likely move to two areas. The first is detection quality: every vendor receives the same transcript, so the difference is accuracy, false-positive rate and speed within a short timeout. The second is coverage: many enterprises use more than one assistant and run MCP servers that no single vendor's hook sees, so value shifts to consistent policy across assistants and to data hooks omit, such as tool definitions and their changes.

Questions for CISOs

  1. Which of our AI risks live in data the hook does not see: tool descriptions, system prompts, file contents, voice? Who covers them?
  2. What share of our AI usage is Claude Enterprise, and what is the policy for everything else?
  3. Is fail-open or fail-closed right for each user group, given a short timeout and a circuit breaker?
  4. If a vendor runs our endpoint, can it mask instead of only block, track MCP tool definitions over time, and show the evidence behind each verdict?

If the answers point to gaps outside Claude's own surface, that is where a gateway such as INS, which sits in front of MCP servers across assistants, is designed to help.

Documentation: Inference hooks, Anthropic.

Secure Your AI Agents Across Every Assistant

Join the waitlist to get early access.

Join the Waitlist

Related Posts