AI Server provides an OpenAI-compatible API for applications on infrastructure the organisation controls. An agent integration needs a more specific acceptance test than a successful chat response: the chosen model must propose the intended operation, and the host application must handle that proposal correctly. Our AI Server product page describes the deployment role; the following procedure describes how we recommend checking an individual integration.
MLCommons' July 2026 edge agentic benchmark documentation illustrates the distinction. Its examples send tool definitions to a Chat Completions endpoint and inspect structured calls returned by the model.[1] An endpoint name alone does not establish that every model, runtime and framework combination will complete the same workflow.
Start with a request that cannot change business data
Choose a disposable test workspace and a model intended for the task. Record the server build, runtime, model identifier and agent framework version. Include the effective endpoint and transport settings, but keep credentials out of the report. The record should distinguish the tested installation from a similarly named model or an older server still running elsewhere.
Begin with one ordinary text request. Confirm which endpoint received it, which model handled it and whether the caller received a complete response. Repeat through the framework that will be used in the pilot. A direct HTTP success and a framework success are separate observations; keep both results.
Use a short input with a predictable answer so that connection failures are easy to separate from task difficulty. A generated greeting establishes very little about a purchasing workflow, but it is a useful first check of the route. Do not add business credentials merely to make this initial test more realistic.
Check a harmless tool proposal before execution
Define a narrow test operation, such as looking up an item in a fabricated inventory. Give it a small schema with required fields and an explicit set of allowed values. At this stage, capture the model's proposed operation without executing it. Check the operation name, the parsed arguments and the framework's treatment of missing or unexpected fields.
Keep the fixture simple enough that a person can inspect every result. A model that returns plausible prose about an item has not necessarily produced a usable tool call. Conversely, a well-formed call may still name the wrong item. Record format acceptance and task correctness separately.
If the production application streams responses, run the same check through streaming. Verify that the consumer waits until the relevant structured value is complete before interpreting it. Treat any difference between streamed and non-streamed handling as an integration finding that needs resolution, rather than assuming the successful path covers both.
Complete the round trip with a controlled result
After the proposal passes validation, allow the host to call the fabricated inventory and return its result to the model. Inspect the subsequent answer. It should use the supplied result and preserve the requested item identity. Include a lookup that returns no match so the application has to represent absence honestly.
Then exercise a short sequence that requires a second lookup. Verify how the framework associates each result with its originating call and how the next request carries the conversation forward. Save a sanitised trace that shows the sequence without retaining real customer documents.
This is the point at which a single-turn demonstration becomes a workflow test. The trace should expose a wrong association, an unnecessary repeat or an abandoned task. The October benchmark article explains why workload boundaries matter when evaluating these longer interactions.
Keep permission checks outside the model's judgement
OWASP identifies excessive functionality, permissions and autonomy as causes of excessive agency. Its guidance places authorisation checks in the systems that execute actions and recommends human approval for high-impact operations.[2] For this integration, test those controls independently of whether the model usually asks for sensible actions.
Extend the fabricated inventory fixture with one operation the caller is not allowed to perform. Verify that the host or downstream service refuses it even when the proposed arguments look valid. The refusal should remain understandable to the user and should not cause the framework to substitute a more privileged credential.
Include a retrieved test document containing an irrelevant instruction. OWASP describes indirect prompt injection through external content and notes that retrieval-augmented generation does not fully remove the vulnerability.[3] The acceptance criterion here is concrete: retrieved prose must not grant the application additional permissions. This test is a useful check, not proof that every injection attempt will be blocked.
Finish with interruption and a recorded scope
Interrupt the test tool, cancel a request and make a dependency temporarily unavailable. Observe what the user sees and what the framework retries. Before introducing a tool that changes records, agree how a retry will avoid duplicating an already completed action. That decision belongs in the application design and the downstream service contract.
Our proposed acceptance record lists the exact operations tested, the permitted identities and any unresolved behaviour. A successful inventory lookup should authorise only the agreed pilot scope. A new connector, broader permission or different model should prompt a review of the affected checks.
For a deployment discussion, bring that record to a Software Tailor integration review, together with a redacted failing example if one exists. The useful next step is a reproducible gap with an owner. The operational handover should preserve those owners after the initial integration work ends.
References
[1] MLCommons. Call for Submission: Edge Agentic Inference Benchmark for MLPerf Inference v6.1. July 2026. Accessed 2026-09-26.
[2] OWASP Gen AI Security Project. LLM06:2025 Excessive Agency. 2025 edition. Accessed 2026-09-26.
[3] OWASP Gen AI Security Project. LLM01:2025 Prompt Injection. 2025 edition. Accessed 2026-09-26.
Related articles
- MLPerf Inference 6.1 and the move to workflow benchmarks
- The handover record for a private AI service
- Before installing a local model