AMR Site Acceptance Testing From Safety Claims to a Release-Ready Evidence Matrix

August 31, 2026

A Successful Demonstration Is Not a Release Decision

An autonomous mobile robot can complete a polished demonstration while still being unready for production. A clear aisle, a light payload, a fresh wireless connection and a vendor engineer standing nearby create favorable conditions; they do not establish that the deployed application will control risk across the operating envelope. The release question is harder: can the owner show, with traceable evidence, that each claimed function works under the site conditions that matter, fails in a defined way and remains governed after handover?

That is the real purpose of AMR site acceptance testing. It is not a ceremonial drive-through at the end of commissioning. It is a structured decision process that connects the risk assessment, application design, software and map baseline, physical environment, test method, measured result, deviation record and accountable approval. The output is not merely a signed checklist. It is a defensible argument explaining why the configured application may enter service, under which limits, and what changes would invalidate that decision.

This article provides a site-specific matrix method for that argument. It intentionally goes deeper than our broader AMR and AGV acceptance testing guide, which covers the overall FAT-to-handover journey. Here, the unit of control is one auditable evidence row: one requirement, one defined condition, one expected response, one result and one release consequence.

Place SAT Correctly in the Assurance Lifecycle

AMR assurance lifecycle from factory acceptance and commissioning through site acceptance and operational qualification.

Acceptance becomes confused when factory testing, commissioning, site testing and operational proving are treated as interchangeable. They answer different questions and should produce different records.

Factory acceptance asks whether the supplied product baseline is ready to ship

Factory acceptance can verify product configuration, component identity, basic safety functions, interface simulations, declared performance and document completeness in a controlled environment. It is valuable because defects are cheaper to correct before equipment reaches the site. However, it cannot fully represent the customer’s floor, traffic mix, wireless behavior, lighting, dust, slopes, door timing, load geometry, human work patterns or production congestion.

Commissioning asks whether the application has been installed and configured

A disciplined AMR commissioning checklist confirms installation details such as charging infrastructure, map and route deployment, scanner mounting, field sets, network identities, interface endpoints, user roles, time synchronization and parameter values. Commissioning makes the system testable. It does not, by itself, prove that the resulting configuration satisfies the risk-reduction claims.

Site acceptance asks whether the configured application earns release

An autonomous mobile robot acceptance test should challenge the exact installed configuration against approved requirements at the real operating site. The test team observes behavior, measures critical quantities, captures records and assigns each result a disposition. SAT therefore sits between “installed” and “authorized for use.” Starting production before this decision reverses the logic: operations become an uncontrolled experiment and evidence is collected after exposure has begun.

Operational qualification asks whether performance remains acceptable over time

Soak tests, capacity trials and early-life monitoring can expose intermittent faults, traffic bottlenecks, battery limitations and maintenance weaknesses. These activities extend confidence, but they should not be used to postpone unresolved safety acceptance. A production KPI may be observed for days; a required safe stop response must be demonstrated before personnel are exposed to the released application.

The SAT Evidence Cube: Claim × Condition × Evidence

SAT evidence cube mapping AMR safety claims to operating conditions and auditable test evidence.

A flat checklist encourages shallow questions such as “Does the emergency stop work?” A stronger model treats acceptance as a three-dimensional evidence problem. We call it the SAT Evidence Cube. The first axis is the claim being made, the second is the condition under which the claim must remain true, and the third is the evidence needed to make a release decision.

Axis one: the claim

A claim may concern collision avoidance, protective stopping, speed limitation, stability, load retention, docking, fleet coordination, access control, manual operation, fault response or recovery. Each claim should be expressed in observable terms. “Robot is safe” is not testable. “When the specified test object enters the active protective field during forward travel at the authorized speed and payload, the vehicle commands the defined stop state without contact” is testable.

Axis two: the condition

Conditions define the boundaries of validity: direction, speed, payload, load overhang, floor state, route geometry, scanner field, lighting, network quality, traffic density, approach angle, human activity, software version and operating mode. A function that passes once under nominal conditions has not been tested at its boundaries. The matrix must sample the combinations capable of changing the outcome.

Axis three: the evidence

Evidence identifies what will prove the outcome: calibrated measurement, controller log, safety-device diagnostic, video synchronized to event time, network trace, configuration export, photograph, witness record or physical inspection. An AMR safety validation matrix becomes credible only when the result can be reconstructed by someone who was not present at the test.

The Evidence Cube changes the review conversation. Instead of asking whether a long list of tests was completed, reviewers ask whether every material claim is covered across its meaningful conditions and whether each pass has adequate evidence. Missing coverage becomes visible before the test campaign begins.

Freeze the Test Article Before Collecting Results

Signed AMR test baseline manifest recording vehicle configuration, site state and environmental conditions.

Results are not portable across an uncontrolled configuration. Before the first formal run, create a signed baseline manifest for the exact test article and site state. At minimum, record vehicle model and serial number, safety controller and scanner hardware, firmware, application software, fleet manager version, map version and checksum, route package, parameter set, safety field configuration, load fixture, battery condition, network configuration, interface versions, user-role model and open deviations.

The manifest should also identify the environment: floor zones and known defects, rack geometry, crossings, doors, lifts, conveyors, chargers, pedestrian boundaries, fire routes, lighting bands, temperature limits and wireless design. Where a test depends on a temporary setup, record it explicitly. A movable pallet, open door or borrowed access point can change the result.

Use immutable identifiers, not screenshots alone

A screenshot proves what a screen displayed, not necessarily which executable, map or parameter file produced it. Store export files and cryptographic hashes where possible. Link each result to the baseline identifier. If a safety parameter changes during troubleshooting, invalidate affected results, create a new baseline revision and rerun the defined regression set.

Separate safety-relevant and performance-only parameters

Classification prevents a productivity setting from silently altering a safety assumption. Maximum speed, deceleration behavior, load envelope, scanner field switching, localization confidence responses and recovery permissions may participate in the safety argument even when configured outside a safety controller. The risk assessment should identify these dependencies and the matrix should make their verification visible.

Anatomy of One Auditable Matrix Row

A useful AMR SAT checklist is not a collection of vague tick boxes. Each row should function as a small, independently reviewable test specification. The following fields create sufficient structure without turning the matrix into an unreadable test script.

Field Purpose Example
Requirement ID Connects the row to an approved source SAF-PS-014
Claim and hazard Explains the risk-reduction purpose Protective stop for pedestrian intrusion
Source Provides traceability Risk assessment, clause, drawing or specification
Test object Identifies the exact vehicle and configuration Vehicle 03, baseline B17
Preconditions Defines the starting state Automatic mode, full charge, no active fault
Condition vector Captures speed, load, direction and environment 1.6 m/s, rated load, forward, dry floor
Stimulus Defines the initiating event Test piece enters field at 90 degrees
Expected response States observable success before execution Protective stop; no contact; fault logged
Measurement method Makes the result reproducible Position encoder plus synchronized video
Acceptance limit Prevents post-test interpretation Stop before boundary with defined reserve
Sample count Addresses repeatability Ten valid runs per condition
Raw evidence link Preserves original data Repository path and file hash
Observed result Records fact rather than opinion Maximum travel 1.42 m
Result status Supports controlled disposition Pass, fail, blocked or not applicable
Deviation ID Links anomalies to corrective action DEV-027
Retest scope Prevents narrow, unapproved retesting Row plus related reverse and turning cases
Witness and date Establishes accountability Integrator, owner and safety representative
Release effect Connects the result to operation Blocks Route Group C

Expected results and acceptance limits must be approved before execution. If the team decides what “good enough” means after seeing the data, the test has become a negotiation rather than verification.

Build the Campaign Around Five Kinds of Challenge

Autonomous mobile robot undergoing operational qualification while engineers monitor its route in a test facility.

Coverage should not be generated by multiplying every variable mechanically. That creates thousands of low-value combinations while missing interactions that matter. Use the risk assessment, design assumptions, field data and failure analysis to select representative challenges in five categories.

Nominal cases establish the intended path

Nominal tests prove that a correctly configured vehicle can complete ordinary missions, interface handshakes and safety responses. They are necessary controls and useful for validating instrumentation, but they are the beginning of the campaign, not the conclusion.

Boundary cases test the edge of authorization

Run at maximum authorized speed, payload and overhang; minimum clearance; weakest permitted wireless condition; tightest turn; steepest allowed slope; worst approved floor zone; and shortest interface timeout. Boundary selection must match the released operating envelope, not the vendor’s theoretical capability.

Negative cases prove that invalid requests are rejected

Send an unauthorized mission, stale command, incompatible load identifier, invalid map revision or restart request while a protective device is active. A system can appear correct because it accepts valid inputs while remaining unsafe because it also accepts invalid ones.

Compound cases expose dependency failures

Many incidents require more than one condition: a blocked aisle during network degradation, a scanner contamination warning while turning with an overhanging load, or a charger fault while traffic queues behind the vehicle. Select combinations that challenge shared resources, timing assumptions and common-cause dependencies.

Recovery cases verify controlled restoration

Every injected failure should have a defined route back to service. Test who may authorize the transition, what must be inspected, whether state is fresh, how the robot localizes, what happens to the old mission and whether nearby personnel receive a clear warning. Recovery is part of the safety lifecycle, not housekeeping after the “real” test.

Measure Stopping Performance as a Distribution, Not a Demonstration

Loaded AMR stopping distance test with synchronized cameras, detection point and protected-distance measurements.

The AMR stopping distance test is among the most consequential SAT activities because site geometry, payload, tires, floor friction, speed, controller latency and field switching all affect the available separation. A single successful stop is weak evidence. The campaign should characterize response and braking across relevant conditions and compare the upper observed behavior, plus justified uncertainty and reserve, with the protected distance.

Separate the timing chain

Where instrumentation permits, measure detection time, safety-device processing, communication delay, controller reaction, brake actuation and mechanical travel. This decomposition helps distinguish a perception delay from a brake-performance issue and prevents one favorable total result from hiding a fragile component.

Test the released motion envelope

Include forward and reverse travel, turning, loaded and unloaded states, speed transitions, representative battery states and site floor categories. Test routes where load swing or vehicle yaw changes the swept envelope. If environmental variation such as dust, condensation or floor contamination is within the authorized use case, include controlled representative conditions or explicitly exclude them from the release.

Use conservative decision statistics

Record every valid run, not only the average. Review maximum observed stopping travel, dispersion, measurement uncertainty and anomalous traces. Define how many repetitions are required and how invalid runs are handled before testing. A pass should retain an engineering margin; matching a boundary within measurement tolerance is not robust acceptance.

The stopping campaign should align with the separation-distance reasoning in the site risk assessment and the protective-field configuration. Our guide to safety laser scanners for AGV and AMR navigation explains why navigation range and safety-certified protective coverage must not be treated as the same quantity.

Validate Protective Fields Against Real Geometry

Forklift AMR protective field validation covering load contour, field switching and scanner gaps.

AMR protective field validation must connect the configured field set to the vehicle’s true swept envelope and stopping behavior. The test is not complete when a scanner diagnostic shows a field number changing. Review the scanner mounting, occlusion, contour, tolerance stack, load overhang, towing or attachment geometry and transition logic at the physical site.

Approach from more than one direction

Introduce specified test pieces at frontal, lateral and oblique approach angles, including near field boundaries and during turns. Consider low-profile objects, legs emerging from behind racks, personnel stepping from doors, and areas obscured by the vehicle body or payload. The selected objects and approach speeds should follow the device instructions, application standard and risk assessment.

Challenge field switching

Verify each transition between speed- or direction-dependent field sets. Look for gaps caused by timing, localization, direction reversal or delayed state information. Confirm that an invalid field-selection input produces the designed safe response rather than retaining an optimistic field indefinitely.

Inspect the complete hazardous contour

The protected geometry must cover the body, wheels, forks, conveyors, lifted load, carried load and anything that can sweep during rotation. A scanner may detect a pedestrian correctly while an unmodeled load corner reaches the person first. Record contour measurements and the approved dimensional allowances in the evidence pack.

Challenge Localization Without Mistaking Confidence for Safety Integrity

Localization supports route following and may influence speed zones, field selection, docking and traffic permissions. Yet a localization score produced by navigation software is not automatically a safety-rated signal. SAT should test the behavior that depends on localization and confirm the architecture assumed by the risk assessment.

Test map and environment disagreement

Change representative landmarks, open and close large doors, place reflective or repetitive surfaces, alter rack occupancy and introduce approved environmental variation. Confirm how the system detects map mismatch, ambiguity or loss of localization. The expected response may be reduced speed, controlled stop, mission abort or operator intervention, but it must be defined and observable.

Verify spatial boundaries at their tightest points

Measure lateral, longitudinal and yaw errors near minimum clearances, crossings and docking points. Include vehicle and load envelope tolerance, control error, map tolerance and required reserve. A high average localization accuracy does not compensate for an occasional large jump where clearance is small.

Control map lifecycle

Attempt deployment of an unapproved or stale map and verify rejection or a governed update path. Record map identifiers in mission logs and test records. A map editor should not be able to change a safety-relevant route boundary without authorization, review and defined regression testing.

Make Fleet Behavior Part of the Site Test Article

Individual vehicle acceptance cannot establish system behavior when a fleet manager assigns missions, reserves zones, controls intersections and coordinates shared resources. AMR fleet acceptance testing should therefore verify both traffic performance and the failure responses of the control architecture.

Test contention, deadlock and queue spillback

Create opposing traffic, blocked destinations, unavailable chargers and occupied narrow aisles. Verify priority rules, deadlock detection, reservation release and queue limits. A vehicle stopped safely can still create a secondary hazard by blocking an emergency route, pedestrian crossing or lift exit.

Disturb messages and clocks

Inject delayed, duplicated, reordered and stale messages at approved test interfaces. Interrupt the fleet service and restore it. Verify that vehicles handle lost authority conservatively, that leases or reservations expire correctly, and that old commands cannot regain validity after reconnection. Time synchronization should be tested because evidence reconstruction and timeout logic both depend on it.

Verify interoperability at the boundary of responsibility

Where VDA 5050 or another interface is used, test the exact implemented version, optional features, state semantics, order updates, instant actions, error handling and connection recovery. VDA 5050 can standardize communication, but its official scope does not provide commissioning, validation or acceptance procedures. An interface-conformant message is not proof that the application is safe.

For the broader traffic architecture, see our analysis of AMR fleet traffic management and safety and the separate guide to multi-vendor mobile robot interoperability.

Test Every Physical and Digital Handshake

Doors, elevators, conveyors, machines, chargers and docking stations form distributed control loops with the vehicle. A green signal on one side may have a different meaning, lifetime or failure response on the other. Each handshake needs an end-to-end requirement and test.

Prove the state sequence, not only the happy path

For each interface, document request, acknowledgement, permission, action, completion, timeout, cancellation and recovery. Test loss of power and communications at each meaningful state. Confirm that a stale “clear” or “ready” state cannot authorize motion after physical conditions have changed.

Include docking and charging tolerances

Test approach accuracy, contact alignment, charger identification, electrical interlocks, foreign-object conditions, blocked egress and failed undocking. Battery or charger faults should not create an uncontrolled manual-recovery task. Detailed docking considerations are covered in our AMR precision docking and autonomous charging guide.

Assign ownership for both ends

The matrix should name the owner of the vehicle logic, infrastructure controller, network path, physical guarding and recovery procedure. “Vendor interface” is not an accountable role. Acceptance gaps often survive because every supplier verifies its own component while nobody verifies the composed behavior.

Degrade Wireless and Cyber Controls Deliberately

Wireless service is an operating condition, not a background assumption. Map coverage and roaming performance under loaded traffic; then test what the application does when latency rises, packets are lost, an access point fails or the connection changes. Vehicle behavior during communication loss must match the architecture: some local safety functions should remain independent, while mission and traffic authority may need to expire or restrict motion.

Cybersecurity tests should include role enforcement, unauthorized command rejection, credential expiry, certificate replacement, remote-access control, configuration audit logs, secure update behavior and restoration from approved backups. Do not use penetration testing as a substitute for verifying safety consequences. The key SAT question is how a digital failure or misuse can alter motion authority, safety-related configuration, evidence integrity or recovery.

Use the site’s approved change window and test environment for intrusive exercises. Our industrial wireless network guide for AMRs and AGVs covers coverage and roaming design, while the mobile robot cybersecurity guide addresses identity, segmentation and lifecycle controls.

Verify Human Tasks, Modes and Recovery Authority

Technicians verify AMR operating modes, human interaction and fault recovery scenarios on a factory floor.

The application includes people who load, unload, dispatch, clean, maintain, rescue and supervise the vehicle. SAT should observe realistic tasks rather than treating humans as test objects who only step into scanner fields. Verify visibility, audible and visual indications, emergency-stop access, manual controls, mode display, restricted-area entry, load placement, fault messages and escalation instructions.

Test mode boundaries

Confirm who can select automatic, manual, maintenance or recovery modes; what motion is permitted in each; which protective functions remain active; and how the current mode is indicated. Attempt prohibited transitions. Removing a key, closing an application or releasing an emergency stop must not automatically restore motion authority.

Run recovery from representative faults

An AMR failure recovery test should cover blocked paths, localization loss, safety-device trips, communication faults, load exceptions, charger failures and stranded vehicles. The operator must receive enough information to make a correct decision without bypassing safeguards. Test limited recovery motion, speed and hold-to-run behavior where provided.

Challenge stale state

After a pause, power cycle or network interruption, verify the freshness of localization, route clearance, load identity, mission authorization, surrounding occupancy and infrastructure state. The vehicle should not resume because it remembers that conditions were safe before the interruption. Our related guide to AMR failure recovery and operational resilience provides a wider operational perspective.

Turn Raw Test Data into a Release-Ready Evidence Pack

The AMR test evidence matrix is the index of the acceptance case, not the entire case. Every row should point to controlled raw records and to the requirement it verifies. Keep original logs, configuration exports, calibration certificates, synchronized video, measurement files, network traces, photographs and witness records in a repository with retention, access and revision controls.

Preserve provenance

Record the creator, device clock, collection method, original filename, hash and any processing applied. A graph is useful for review, but the underlying measurements must remain available. If a video is trimmed or annotated, retain the original. If log times are converted, document the transformation and clock offset.

Make failures as traceable as passes

Do not delete failed runs after correction. Link the deviation, root cause, configuration change, impact analysis, regression scope and successful retest. A clean report containing only final passes hides the engineering history that explains why confidence is justified.

Write a decision summary for non-testers

The release record should summarize coverage, unresolved deviations, operational restrictions, training dependencies, maintenance conditions, cybersecurity assumptions and revalidation triggers. It should identify the named technical owner and business authority who accepted residual risk. A signature without a readable decision basis is weak governance.

Classify Deviations by Release Effect

Not every failed row has the same consequence, but no team should improvise disposition at the closing meeting. Define categories before testing and tie them to permitted decisions.

Decision Meaning Minimum control
Pass Requirement met with valid evidence Result approved and baseline retained
Pass with bounded condition Release is acceptable only inside an explicitly reduced envelope Restriction engineered, documented, communicated and monitored
Hold Evidence is missing, invalid or an unresolved issue could affect the claim Affected operation remains disabled pending correction and retest
Reject The design cannot meet the approved requirement or residual risk is unacceptable Redesign and renewed assessment before another acceptance attempt

A temporary restriction is not a verbal promise to “be careful.” It must be technically enforceable where practical, reflected in routes and permissions, included in training, shown in the release record and assigned an expiry or review date. If the restriction can be removed casually, it is not a reliable control.

Control retest scope through impact analysis

When a defect is corrected, identify all requirements influenced by the changed component, parameter, interface or assumption. A brake adjustment may require stopping tests across several loads and speeds; a map change may affect traffic, speed zones, localization and field switching. Retesting only the row that failed can preserve an unknown regression elsewhere.

Use a Requirements-to-Evidence Coverage Map

Before the final decision, review the matrix from both directions. Requirement-to-test review asks whether every approved requirement has adequate coverage. Test-to-requirement review asks whether every executed test has a purpose and approved acceptance limit. Orphan requirements indicate missing evidence; orphan tests often indicate inherited scripts that do not match the application.

Group coverage by claim family: protective stopping, emergency stopping, speed control, stability, load handling, localization-dependent behavior, traffic coordination, infrastructure interfaces, wireless failure, cybersecurity, operating modes, manual intervention, recovery, maintenance and documentation. Mark the condition variants represented for each family. This makes a concentration of easy nominal tests immediately visible.

Use risk to prioritize, not to excuse omissions

Risk-based sampling can reduce redundant combinations, but it cannot remove a material safety claim from verification. Document why selected cases represent omitted variants. Where modeling or supplier certification supports the rationale, define its validity limits and confirm that the installed configuration stays within them.

Write Acceptance Requirements into Procurement

A rigorous test matrix is difficult to build after equipment arrives if the contract does not provide data, tools, access or supplier support. Procurement specifications should require a requirement traceability matrix, proposed SAT plan, configuration exports, diagnostic access, raw log retention, interface simulators, calibration information, fault-injection support, evidence ownership and correction response times.

Define which party supplies test objects, loads, instruments and witnesses. State the required production-representative vehicle count and whether a change to one unit must be propagated and verified across the fleet. Specify the format of the final evidence pack and the conditions for provisional acceptance, payment milestones and final release.

Do not purchase a performance number without its conditions

Throughput, docking accuracy, stopping distance and wireless availability claims should identify payload, speed, environment, confidence level, sample method and excluded conditions. A headline number without an operating envelope cannot become an enforceable acceptance limit.

Align the Matrix with the Current Standards Landscape

AMR safety standards alignment matrix connecting ISO and ANSI requirements to site acceptance evidence.

As of August 2026, ISO 3691-4:2023 is the published international standard addressing safety requirements and means for verification for driverless industrial trucks and their systems, including AGVs and AMRs. Its application-level perspective means the operating zone, hazards and installed configuration matter. ISO also lists a draft replacement under development, so organizations should record the exact edition used for the assessment rather than writing “latest standard” into controlled documents.

In the United States, ANSI/A3 R15.08-2-2023 addresses integration, configuration and customization of industrial mobile robot systems and applications. ANSI/A3 R15.08-3-2026 adds responsibilities for users and maintaining acceptable risk during use, including attention to risk assessment, the current operating environment and management of change. Part 3 makes the handover from project acceptance to lifecycle governance especially important.

ISO 12100 supports the machinery risk-assessment method. ISO 13849-1:2023 supports design and integration of safety-related control-system parts, while ISO 13849-2:2012 addresses validation by analysis and testing; a replacement for Part 2 is in development. ISO 13850 and IEC 60204-1 may also be relevant to emergency-stop and electrical-equipment aspects. Applicability depends on jurisdiction, machine boundaries and the risk assessment, so the matrix should cite specific clauses and project interpretations approved by qualified professionals.

Standards provide requirements and methods; they do not supply one universal field size, test count or acceptance script for every site. The test matrix must translate applicable requirements into the actual vehicle, load, environment, interfaces and operating concept.

Focused FAQ

What is the difference between commissioning and site acceptance?

Commissioning installs and configures the system so it can operate as designed. Site acceptance verifies, with approved methods and evidence, that the configured application meets requirements in its real environment. A completed installation is a prerequisite for acceptance, not proof of acceptance.

How many stopping tests are enough?

There is no universal number. The test plan should justify repetition based on variability, risk, measurement uncertainty and the range of speed, load, direction, floor and motion conditions. Define sample counts and statistical decision rules before viewing the results.

Can vendor certificates replace on-site tests?

Certificates and product validation can support the evidence case within their declared boundaries. They normally cannot prove site-specific field geometry, stopping behavior, interfaces, traffic rules, environmental conditions or recovery procedures. Use them to reduce justified duplication, not to erase application verification.

Should every AMR in a fleet be tested?

Each vehicle should receive identity, configuration and essential functional checks. More extensive type or boundary testing may use a justified representative sample when units are demonstrably equivalent. The sampling rationale, propagation controls and fleet-wide regression requirements should be documented.

Does VDA 5050 compliance mean the fleet has passed acceptance?

No. VDA 5050 defines communication between mobile robots and a master control, and its published scope excludes validation and acceptance procedures. The implemented interface still requires site tests for semantics, failure handling, timing and safe system behavior.

Can a failed test be accepted with an operating restriction?

Only when a renewed risk evaluation shows the bounded operating envelope is acceptable and the restriction is enforceable, documented, communicated, monitored and approved. A warning in meeting minutes is not an adequate substitute for a control.

What changes trigger revalidation?

Typical triggers include software or firmware updates, scanner or brake replacement, map and route changes, new payloads or attachments, speed increases, altered floor or rack geometry, wireless redesign, interface changes, new traffic participants, changed recovery procedures and recurring incident trends. Each trigger requires impact analysis and a defined regression scope.

Who signs the final release?

The organization should name technical reviewers for safety, controls, operations, IT or cybersecurity and maintenance as applicable, plus the accountable business authority accepting residual risk. Supplier sign-off alone cannot replace the user organization’s responsibility for its operating environment.

The Release Question Is Whether the Evidence Survives Scrutiny

A credible SAT does not attempt to prove that failure is impossible. It proves that the organization has identified material claims, tested them under representative and boundary conditions, preserved reproducible evidence, controlled deviations and defined the limits of release. The matrix turns a complex mobile-robot application into reviewable decisions without pretending that a generic checklist can replace engineering judgment.

The strongest acceptance package can answer five questions quickly: what exactly was released, which claims were verified, under what conditions, with what evidence, and what change requires reconsideration. If any answer depends on the memory of the commissioning team, the evidence chain is incomplete. Freeze the baseline, test the boundaries, retain failed as well as successful data, and make the final authorization explicit. That is how site acceptance becomes a durable control rather than a one-day ceremony.