ISSUE
When importing .eml files with PDF attachments, customer has noticed that the OCR appears to be crashing - as a result, the Extraction fails with the following error:
Error from ExtractionProcess.exe:
BeforeExtraction: BeforeExtract - Document:*****.xdc - Extraction class: ***** - executeDocumentAfterProcessEvent: True: The execution of a locator method failed. Class = "PH_0001_INV_AVIS", Locator = "AEL_Advice", Original error message:
Web service failure: error code=0x803d0013
The server returned a fault: External component has thrown an exception.
CAUSE
It seems like the PDF attachment in the email has poor text representation and, when the PDF text extraction in the process is also configured to use “All text” in conjunction with the poor text representation, the Extraction gets suspended with the error above.
SOLUTION
The following bug has been submitted to address this issue:
Bug 2265724:Failed Extraction: "Web service failure: error code=0x803d0013 The server returned a fault: External component has thrown an exception."
There are two workarounds that can be applied until the bug is fixed:
1) Change the PDF text extraction from "All text" to "Ignore all text layer”.
Sidenote: The process level setting will apply to all incoming PDF files, even the ones with a "good" text representation, so the customer might not want to ignore the text layer for all PDF files.
2) Alternatively:
1. Enable "Reject document on exception" on the Extraction activity
2. If the Folder.HasRejections property is true, do the following:
- Unreject the document using the UnrejectDocuments() method.
- Pass the document through an OCR activity using OmniPage with the "Force OCR" property enabled.
Any documents rejected with this error can then be routed back to Extraction.
REFERENCES
| Product | Version | Build | Environment | Hardware |
|---|---|---|---|---|
| TotalAgility |