What Evidence Should an AMR Simulation Provide Before Fleet Purchase
Start With a Purchase Decision the Model Can Test
AMR simulation validation asks whether a model is credible enough to support a specific purchasing decision. Before approving a fleet, buyers need evidence that the model represents the intended operating conditions, reproduces relevant measured behavior and remains useful when uncertain assumptions change. An attractive factory animation cannot establish those properties by itself.
A useful study begins with a decision statement: choose between defined layouts, vehicle configurations or infrastructure options while meeting an agreed material-service requirement. It then connects each recommendation to input records, model checks, controlled experiments and the conditions under which the recommendation remains valid. The outcome should identify what the buyer can approve, what still requires measurement and what must be demonstrated after installation.
Consider a supplier proposing eight robots and another proposing ten. Comparing their headline throughput figures reveals little if one assumes continuous station availability while the other includes interruptions. The buying question is whether both proposals were tested against the same operating problem with comparable evidence.
This article provides an engineering review method for that problem. The example values below are explicitly hypothetical. They illustrate a review process; they are not results from a customer installation or a simulation executed for this article.
The existing preliminary fleet quantity calculation provides an initial range of configurations. The next task is to evaluate how much confidence the buyer should place in the model used to compare those configurations.
Select the Model According to the Claim
An AGV simulation model may represent vehicles as moving agents, transport events, detailed physical systems or a combination of these. The choice should follow the evidence needed. A model intended to estimate station queues does not automatically predict localization performance. A detailed sensor simulator does not automatically include an entire production schedule.
| Model approach | Useful purchasing question | Evidence still needed |
|---|---|---|
| Discrete-event material-flow model | Can transport and stations serve the defined request pattern? | Measured timing distributions, resource constraints and verified event logic |
| Agent-based movement model | How do individual vehicles interact under the proposed operating rules? | Representative vehicle behavior and the actual limits of the traffic implementation |
| Physics and sensor simulation | How might motion or perception respond to defined physical conditions? | Validated physical parameters, sensor assumptions and corresponding physical tests |
| Controller-connected emulation | Does the intended control software behave correctly against a simulated plant? | Software-version equivalence, interface fidelity and realistic timing |
These approaches can be combined. However, each connection adds assumptions about timing, data exchange and model boundaries. A simulator may advance one hour of logistics events in seconds while a controller expects events at real-time intervals. The integration must preserve the ordering and latency behavior relevant to the test.
NVIDIA describes a robot fleet simulation workflow that connects fleet management, robot policies, a world simulator, sensor simulation and scheduling. It demonstrates the range of components a detailed virtual environment can contain. The existence of such components is not, by itself, evidence that a particular factory model has been validated.
Define what the digital twin actually represents
For AMR digital twin validation, ask which physical system the model represents, how its data are maintained, which operating conditions have been checked and how changes reach the model. A one-time planning model may be entirely suitable for a capital decision. Calling it a digital twin does not expand the decisions it can reliably support.
NIST's discussion of digital twin credibility places verification, validation and uncertainty assessment within the model lifecycle. For a buyer, the practical implication is to request an evidence trail tied to the intended use, rather than accepting a software category as a credibility certificate.
Audit the Inputs Before Adjusting the Model
The most productive first review in an intralogistics simulation study is often a data review. Even a correctly implemented model can produce misleading recommendations when the demand stream, station behavior or initial conditions describe a factory that will never exist.
Every influential input needs a value or distribution, unit, source, observation period, operating context, owner and uncertainty statement. Distinguish measured inputs from engineering estimates and contractual assumptions. A spreadsheet with three decimal places should not hide that the underlying transfer time came from one quiet demonstration.
| Input group | Evidence to request | Common source of bias |
|---|---|---|
| Transport requests | Timestamped releases, request classes, destinations and deadlines | Smoothing shift changes or production batches into a constant average |
| Travel behavior | Representative segment times by load, turn and speed zone | Applying rated maximum speed to every meter of movement |
| Station service | Arrival, readiness, transfer and release timestamps | Omitting occupied approaches or counting overlapping delays twice |
| Energy behavior | Consumption and charging observations under relevant duty conditions | Assuming all vehicles begin every shift fully charged |
| Disruptions | Observed interruption patterns and defined stress assumptions | Removing delays as outliers without explaining their cause |
| Operating context | Shift calendars, staffing, product mix, layout revision and traffic conditions | Combining observations from incompatible operating periods |
Preserve timing relationships
A production machine may release several transport requests together. The same shift change can increase pedestrian crossings and reduce station staffing. If these events are sampled independently, the model can underestimate the combined disruption even when each individual distribution looks reasonable.
Use recorded time sequences where appropriate, or explicitly model the common driver. When constructing synthetic demand, document which relationships were retained and which were simplified. Replay answers how a design handles a particular recorded period; generated scenarios explore a broader set of possible periods. Neither approach automatically replaces the other.
Make the measurement boundary explicit
A station observation of ninety seconds might include thirty seconds waiting for readiness and sixty seconds transferring material. If the model samples ninety seconds as transfer time and separately adds readiness waiting, it duplicates the delay. If it samples sixty seconds and never represents readiness, it removes a real constraint.
A practical timestamp check
For each measured cycle, identify when the request became executable, when the robot reached the approach, when transfer became permissible, when material movement ended and when the resource became available again. Resolve clock offsets and missing events before calculating durations. Store the definitions beside the input data so another reviewer can reproduce the calculation.
Greenfield projects need a different evidence strategy because complete site histories do not exist. Combine supplier component tests, observations from comparable processes and explicit ranges for unmeasured behavior. Record the differences between the reference process and the proposed installation. The remaining uncertainty belongs in the experiment plan.
Verify the Logic, Then Validate the Behavior

Verification asks whether the implementation behaves according to its specification. Validation asks whether that behavior represents the real application adequately for the intended decision. ASME's overview of verification, validation and uncertainty distinguishes these activities. Both are necessary when a model will influence equipment expenditure.
Begin with cases whose answers are known
Test a single vehicle on an unobstructed route with fixed travel and transfer times. Its cycle duration should match a manual calculation within the model's declared timing resolution. Then remove all demand and verify that no transport jobs appear. Make one station unavailable and check that requests cannot complete through it.
Add conservation checks. For example, released requests should reconcile to completed, open, cancelled or explicitly rejected requests under the defined accounting boundary. A load cannot occupy two exclusive locations simultaneously. A finite buffer cannot silently hold unlimited vehicles. An unavailable vehicle must not continue contributing service capacity.
These checks are particularly valuable in robot fleet simulation testing because a visually plausible movement sequence can conceal an impossible transaction. Keep an event trace for failed checks and identify the model revision in which each issue was resolved.
Calibrate parameters without disguising structural errors
AMR model calibration adjusts uncertain parameters using observed behavior. It should not compensate for missing mechanisms by changing unrelated variables. Reducing travel speed until a model matches observed throughput may hide omitted station waiting. That model can then predict the wrong response to an additional station or a different route.
Compare component behavior before aggregate totals. Review segment travel times, readiness waiting, transfer duration, resource occupancy and request completion by class. Agreement in total daily output is insufficient when the model serves one mission family too quickly and another too slowly.
Reserve observations for an independent check
Use one set of observations for adjustment and another appropriate set to evaluate the adjusted model. The second set should cover the conditions relevant to the purchase, such as a busy shift or a different product mix. If it comes from a materially different process, explain that difference before interpreting the error.
Report signed errors as well as absolute errors. A model that consistently predicts shorter delays is commercially different from one with small, balanced deviations. Plot residuals against demand, location and time. A pattern that appears only near high utilization may expose a missing interaction that average statistics conceal.
There is no universal percentage that makes mobile robot simulation accuracy acceptable. The required accuracy depends on the margin separating the candidate designs from the purchasing threshold. A model can be useful for rejecting an obviously inadequate layout while remaining too uncertain to distinguish between two nearly equivalent proposals.
Write the Experiment Protocol Before Looking at Winners

Well-designed warehouse simulation experiments separate the configuration being evaluated from the conditions under which it operates. Configuration variables might include station count or charger location. Operating variables might include request timing, load mix or a defined interruption. Record both in a run register rather than changing settings informally between demonstrations.
Use the correct start and finish conditions
A terminating experiment represents a defined period, such as a production shift. Its opening backlog, vehicle positions, battery states and station occupancy should reflect that period. Starting every case with empty queues and fully charged vehicles creates a favorable initial state that may not describe normal operations.
A steady-state experiment addresses long-run behavior after initialization effects have diminished. It may require a justified warm-up period. Automatically deleting the first hour from a shift experiment is different: it can remove the very startup problem the project needs to understand.
Define how requests near the end of a run are handled. If only completed missions enter a response-time report, the slowest unfinished missions can disappear from the metric. Report outstanding work separately and, where the analysis requires it, follow the defined request cohort until its outcomes are known.
Make random runs reproducible
Save the model version, input revision, scenario, random seeds, software configuration and output definitions. Independent replications expose variation between possible operating periods. A single seed can provide a reproducible demonstration, but it does not establish that the chosen outcome is representative.
AnyLogic's parameter variation documentation describes repeated runs and confidence-based evaluation for stochastic models. The purchasing principle is broader than any particular tool: select replication effort according to the decision's required precision and report the rule used to stop experimentation.
When comparing designs, deliberately coordinated random streams can expose both designs to equivalent external demand and disruption sequences. Simply entering the same seed may not achieve this if model changes alter the order of random-number consumption. Verify the pairing and keep the experimental unit clear.
Keep averages, percentiles and uncertainty separate
A 95th-percentile response time describes a tail of a response-time distribution. A 95% confidence interval around a mean addresses uncertainty in estimating that mean. These are different quantities. Neither establishes that every future shift will satisfy a service limit.
For a percentile-based requirement, specify how the percentile is estimated, which missions enter it, how unfinished work is treated and how estimation uncertainty is evaluated. Depending on the question, the study may also need the proportion of shifts meeting the requirement. More random replications reduce sampling uncertainty; they do not correct biased inputs or omitted behavior.
Maintain a register of AGV simulation scenarios covering the expected operating envelope, credible disturbances and combinations capable of changing the decision. Label deliberately severe stress cases separately from expected operating cases. A stress scenario reveals vulnerability without claiming that its assumed event frequency has been measured.
Worked Review: Three Proposals and One Uncertain Station
Suppose a factory expects forty-five eligible deliveries per hour during a sustained busy period. Proposal A uses eight robots and one receiving position. Proposal B increases the fleet to ten robots while retaining that position. Proposal C keeps eight robots and provides two independently usable receiving positions.
For this illustrative review, assume each position is available for service during 90% of the period. Its nominal service time is sixty seconds per load, but an unverified planning range extends to eighty seconds. Availability here excludes the time already included in service duration, avoiding duplicate loss accounting.
Check a necessary capacity condition first
The approximate long-run service ceiling for one continuously supplied position is its available service time divided by its mean service time. At sixty seconds, the ceiling is 0.90 × 3,600 / 60 = 54 loads/hour. At eighty seconds, it falls to 0.90 × 3,600 / 80 = 40.5 loads/hour.
| Proposal | Robots | Independent receiving positions | Station ceiling at 60 seconds | Station ceiling at 80 seconds |
|---|---|---|---|---|
| A | 8 | 1 | 54 loads/hour | 40.5 loads/hour |
| B | 10 | 1 | 54 loads/hour | 40.5 loads/hour |
| C | 8 | 2 | 108 loads/hour | 81 loads/hour |
These are arithmetic screening bounds, not simulated throughput predictions. They assume sufficient incoming work, the stated availability and, for Proposal C, genuinely independent service capacity. Shared operators, a common downstream conveyor or a single transfer permission could invalidate that independence.
Under the eighty-second assumption, Proposals A and B have a station ceiling below the sustained demand of forty-five loads per hour. The long-run demand-capacity gap is at least 4.5 loads per hour before considering other constraints. Extra robots cannot remove that station limit.
Proposal C clears this necessary station-capacity screen. It has not yet proved that eight robots can transport the required loads, that two approaches can operate together or that delays meet the service requirement. The screen identifies which experiment is worth running and which assumption can invalidate a proposal.
Ask the simulation to resolve the remaining comparison
Run the candidate layouts across the agreed station-time range, preserving demand bursts, vehicle eligibility, approach geometry and the receiving process. Include the loss of one position in the two-position design if that disturbance belongs to the project requirement.
Inspect service by destination and time window. A fleet can complete enough total deliveries while repeatedly missing the deadline of one production-critical destination. Also check whether a second position relocates the constraint to a shared aisle, a downstream process or vehicle travel.
The relevant traffic-resource design principles should be represented in the model. Record the implemented policy and its version so that a favorable result cannot depend on behavior absent from the purchased software.
Identify the measurement that changes the decision
For one receiving position to exceed forty-five loads per hour under the assumed availability, its mean service time must be below 0.90 × 3,600 / 45 = 72 seconds. At exactly seventy-two seconds, the arithmetic leaves no capacity margin for clearing accumulated work. Meeting a response-time target generally requires additional headroom and depends on variability.
The sixty-to-eighty-second uncertainty range crosses this boundary. The appropriate next action is a better station measurement or a redesigned receiving process. Running hundreds of replications at an arbitrary sixty-second assumption would refine the wrong conditional answer.
This is the commercial contribution of AMR simulation validation: identifying which uncertainty deserves resolution before the purchase becomes difficult to change.
Challenge the Assumptions That Could Reverse the Recommendation

Sensitivity analysis should reveal whether a recommendation survives plausible input changes. Begin with factors that combine uncertainty and operational influence, such as a poorly measured transfer time at a heavily used station. A highly uncertain parameter that barely affects the decision may deserve less attention.
Screen individual factors to understand direction, then test important combinations. A longer transfer time and a burstier request stream may create a failure that neither factor produces alone. Likewise, a lower starting battery state may matter only when a busy window overlaps the charging schedule.
Distinguish variation within a known process from uncertainty about the process itself. Different request sequences are operating variation. Not knowing whether a shared operator will support both receiving positions is a structural assumption. Repeating the same structure with new random seeds cannot resolve that assumption.
Test the recommendation on cases that did not select it
If an optimizer evaluates many configurations, the apparent winner may partly reflect favorable random outcomes. Re-evaluate shortlisted candidates using an appropriate independent experiment set. Report the size of the decision margin and whether uncertainty could change the ranking.
For close alternatives, the useful conclusion may be that the current evidence cannot reliably distinguish their operational performance. The buyer can then request a targeted measurement, compare other documented requirements or preserve flexibility in the installation. An unjustified ranking is less useful than a clearly identified evidence gap.
Check whether faster approximations preserve the decision
A simplified or surrogate model can reduce computation. Before using it to select equipment, compare its relevant outputs with the more detailed reference under conditions near the decision boundary. Good agreement in routine cases does not establish agreement during bottlenecks or disturbances.
Keep the approximation's validated range visible. If a proposed layout, vehicle behavior or operating condition falls outside it, review the applicability before accepting an extrapolated result. Model speed is valuable when it enables better experiments; its value depends on preserving the distinctions the buyer needs to make.
Convert the Study Into Purchase Conditions and Field Evidence
AMR simulation acceptance criteria should separate acceptance of the model from acceptance of the proposed system. A well-validated model may show that a design is inadequate. A favorable design result from an unvalidated model may still be unsuitable for purchasing approval.
| Review question | Model evidence | Purchase or field follow-up |
|---|---|---|
| Were the relevant inputs represented? | Traceable input register, context and assumption ranges | Resolve material assumptions and identify the party providing each missing measurement |
| Does the implementation follow its specification? | Verification cases, event checks and defect disposition | Retain the reviewed model and software revisions |
| Does it reproduce relevant observations? | Calibration records and independent validation comparisons | Declare the operating range and remaining prediction limitations |
| Does the proposed design meet the requirement? | Scenario results, uncertainty assessment and outstanding-work accounting | Translate the claim into a measurable performance condition |
| Can another reviewer reproduce the recommendation? | Run register, seed policy, output definitions and accessible result files | Agree access rights, executable environment and retention responsibility |
| Will installed behavior match the modeled assumptions? | Explicit link from each claim to a physical verification activity | Use a corresponding site test and investigate material prediction differences |
One possible project requirement might specify response time for a named mission class over a defined request stream, with a stated initial condition and disturbance case. The test method must also define sample size, timing boundaries, exclusions and how unfinished or failed requests affect the result. A bare promise of an average hourly throughput omits too much information.
The site's site acceptance test matrix provides a related framework for field evidence. Use simulation to explain the expected response; use commissioning measurements to establish what the installed application actually does.
If field results differ, first reconcile demand, configuration, timing definitions and operating conditions. Then examine model structure and parameter values. Quietly adjusting the model after installation until it matches the observed total does not explain why the original purchase prediction was wrong.
Preserve the original approved prediction and record later revisions separately. Changes to station layout, dispatch software, vehicle behavior or operating schedules should trigger a review of the conclusions they affect. A useful model remains attached to its assumptions and evidence as the installation evolves.
Physical safety performance, wireless coverage and actual transfer behavior require their relevant site assessments and tests. A simulation result can inform those activities, but a planning model's successful run does not certify the installed application.
Focused FAQ
What should a buyer receive from an AMR simulation supplier?
Request the decision statement, model scope, input register, assumption ranges, verification evidence, validation comparisons, experiment protocol, scenario results and limitations. Include enough model and run information to reproduce the recommendation through an agreed process. A presentation containing only an animation and a fleet count is insufficient for independent technical review.
How many simulation runs are enough?
There is no universal count. The answer depends on the variability of the selected output, the required estimation precision and the decision margin. Specify the experimental unit and stopping rule. A percentile or shift-level reliability requirement may require a different analysis from a mean-throughput comparison.
Can a greenfield model be validated before robots are installed?
Some components and behaviors can be checked against supplier tests, comparable operations and controlled experiments. Whole-system predictions will retain uncertainty where the proposed conditions have no direct observations. State that limitation, test relevant ranges and plan the field measurements required to close the remaining evidence gaps.
Why can a simulation match throughput but predict delays incorrectly?
Different combinations of travel, waiting and service times can produce a similar total output. Errors may also cancel across mission classes. Review component timings, response-time distributions and performance by destination. Aggregate agreement alone does not establish that the modeled mechanisms will respond correctly to a design change.
Does a digital twin need more detail than a planning model?
Its necessary detail depends on its intended use and the physical relationships it must represent. A lifecycle-connected model also needs a defined method for maintaining relevant data and configuration. Additional visual or physical detail is useful only when it improves the decisions or assessments the model is expected to support.
What is the difference between model uncertainty and random variation?
Random variation concerns different outcomes within the represented process, such as different request sequences. Model uncertainty includes incomplete knowledge of parameters and whether the represented mechanisms are correct. Independent replications help measure stochastic variation; measurement, validation and alternative structural assumptions address other uncertainty sources.
When should a favorable simulation result delay a purchase?
Delay the affected commitment when a material unverified assumption can reverse the conclusion, when results cannot be reproduced or when the model omits behavior essential to the performance requirement. Record the specific evidence needed to resolve the decision. Further analysis should target that gap rather than generate more favorable demonstrations.
The purchasing deliverable is a traceable recommendation with a defined range of validity. Approve the proposed system when its evidence supports the required service under the agreed conditions, and preserve the model, assumptions and tests that make that approval explainable.