Answer
Proof requires production data, not vendor samples. Buyers should bring an actual batch of spool files, including whatever oddball forms and edge cases exist in real output, along with a stack of genuinely messy scanned paper, and watch the software index all of it in the same session. It is worth timing the process and asking what percentage of documents come through with correct index values on the first pass, since that number, not the vendor's marketing claim, is what determines how much manual cleanup staff will actually do every week.
Throughput matters as much as accuracy. A capture engine that indexes cleanly one document at a time may behave very differently against a month-end spool file run of several thousand pages, so buyers should ask to see or simulate a realistic batch volume rather than a handful of samples. Exception handling deserves equal scrutiny: when a document cannot be indexed automatically, does it route to a queue where a person can fix it quickly, or does it get lost in a way nobody notices until someone goes looking for a document that was never properly filed.