Answer
Most DR tests fail quietly at the same three points: startup order, access, and time. Startup order breaks when a runbook lists which subsystems and applications to bring up but nobody has validated the dependencies between them, so a job scheduler or interface starts before the data it needs is ready. Access breaks when the team restores the system successfully but discovers that 5250 terminal emulation, VPN profiles, or exit point security rules were never configured for the recovery environment, so users technically have a working system they cannot reach. Time breaks when nobody has ever run the clock on the full sequence, so the organization is working from an estimate instead of a measured result.
A realistic test also has to include the people, not just the technology. Whoever declares the disaster, whoever executes each step, and whoever validates the applications should participate under conditions that resemble the real event, including limited communication, a defined time window, and no quiet correction from someone who already knows the answer. Buyers should ask how the DR software supports timed, evidenced test runs, including automatic capture of start and completion times for each recovery step, because that record is what turns a tabletop exercise into a defensible, auditable proof of readiness.