7 April 2026 is the date NIST gives for its concept note on an AI Risk Management Framework profile for critical infrastructure.[1] The distinction matters: a concept note describes work towards guidance. It is not evidence that a particular product has passed an assessment. For private AI buyers, the practical response is to make the proposed deployment specific enough to examine.
Our position is that the strongest pilot record follows a real task through its operating boundary. The document should name what the system may do, what it may not do, and who takes over when the output cannot be trusted. A label such as “on-premises” cannot supply those answers by itself.
Separate the source from the interpretation
NIST describes the AI RMF as voluntary and identifies the original framework's release in January 2023. Its current overview also describes revision work and the critical-infrastructure profile initiative.[1] Those are useful facts to preserve with their dates. They should not be rewritten as a new legal obligation or a guarantee that an existing checklist is complete.
An internal briefing can keep this distinction visible with two short paragraphs. The first reports what the source actually says and links to it. The second states what the organisation proposes to do in response. That makes later updates manageable: a changed source does not require guessing which parts of the briefing were facts and which were local decisions.
This article proposes a way to assemble technical evidence. It does not determine which sector-specific obligations apply to an organisation or whether a deployment satisfies them.
Draw the actual operating boundary
A maintenance-document assistant and a system that changes equipment settings are different proposals. Describe the first allowed task in ordinary terms before discussing models. State the input, the person using the result, and the action that person is authorised to take. Include the consequence of an incorrect answer in the same description.
Then trace the data path. A proposed AI Server deployment belongs in a diagram with its clients, identity systems and storage. Model downloads, diagnostics and optional provider routes deserve their own entries. The question is where each activity runs and who operates it, not whether everything can fit under one reassuring product label.
For a document-only pilot, the team might prohibit generated text from directly triggering an operational action. Record that restriction as an actual integration decision. Merely adding a sentence to a training slide does not establish what the software can call.
The enterprise platform overview provides a starting point for discussing topology. A buying record still needs the configuration chosen for the particular site.
Preserve the identity of what was tested
Hugging Face documents model cards as a place for model information including intended use, limitations and evaluation.[2] Save the candidate's card with the revision reference used during review. Record the runtime and deployment settings alongside it. A later reviewer should not have to infer which model file was behind an old result.
Keep test inputs under the organisation's own access controls. The review record can refer to an approved test set without copying confidential source documents into a procurement slide deck. State who may retrieve the inputs and how the test can be repeated.
Record the surrounding workflow too. An answer generated from a different document collection is a different test, even if the model file did not change. So is an answer evaluated against a relaxed acceptance rule. Versioning only the model leaves those changes invisible.
Test the failure and the handover
Choose examples that expose the boundaries of the task. For a document assistant, include a question with no answer in the permitted documents, two passages that conflict, and a scanned page whose reading order is awkward. Agree in advance how the reviewer will judge each response. These are suggested test cases, not a certified assessment suite.
Performance evidence also needs its conditions. MLCommons describes inference benchmarks using defined workloads, quality targets and measurement scenarios.[3] A result gathered under those conditions is useful in its proper context. It is not an observed service level for an untested installation.
For the pilot, measure the actual review workflow and document what happens when the service stops. Who receives the failure? Can the user return to the original document? What record survives a cancelled request? Exercise those paths while the system is still small enough for the team to understand.
Recovery deserves a named owner. Keep the previous known configuration available under the organisation's change process, and define who can approve returning to it. A proposed fallback that nobody can execute is an unfinished part of the design.
Make the next decision narrow
The review should end with a bounded decision: continue this task under these conditions, repeat these tests after a change, or stop until a named deficiency is addressed. Avoid turning a successful document pilot into an endorsement of unrelated operational uses.
The cost record should use the same boundary. Our article on cost per accepted task explains why review and rejected outputs belong in that calculation. Technical and financial reviewers can then discuss the same unit of work.
Use the AI Server information to identify the deployment questions, then bring the proposed task and acceptance record into the evaluation. Evidence becomes useful when another person can repeat the test and understand the decision.
References
- NIST. AI Risk Management Framework. Accessed 2026-09-12.
- Hugging Face. Model Cards. Accessed 2026-09-12.
- MLCommons. MLPerf Inference: Datacenter. Accessed 2026-09-12.
Related articles
- Private AI costs: measure the accepted task
- Keeping an honest record of AI-assisted articles