GetOCRDiagnosticsJSON
OCR, diagnostics
Description
Returns version-1 JSON describing the latest OCR phase and layer action, with reports from the currently assigned instance-owned engine when available
Syntax
Delphi
Function TPDFlib.GetOCRDiagnosticsJSON: WideString;
ActiveX
Function PDFlib::GetOCRDiagnosticsJSON() As String
DLL
const wchar_t* DLGetOCRDiagnosticsJSON(int InstanceID);
const char* DLGetOCRDiagnosticsJSONA(int InstanceID);
Return values
A JSON object containing version, phase, status, complete, errorCode, and error, with page and word counts when applicable
Layer status
| notStarted | No OCR operation has started on this instance |
|---|---|
| configured | An engine configuration was accepted; this does not prove that a trained-data model can recognize a page |
| skipped | The page has non-owned existing text; reason is existingText |
| empty | No word met the threshold; old OCR layers remain intact |
| appended | A new owned OCR layer was added |
| replaced | At least one owned old OCR layer was replaced |
| error | The operation failed; inspect the error fields |
| cancelled | The provider cancelled the current operation |
Engine reports
The optional provider object contains its latest status, complete, exitCode, error, and engine log
Provider status codes are 0 success, 1 unavailable, 2 invalid input, 3 invalid output, 4 failed, 5 timed out, and 6 output limit
The optional orientation object includes orientationDegrees, rotateClockwise, confidence, and its own completion and error fields
Orientation status codes are 0 success, 1 unavailable, 2 invalid input, 3 invalid output, 4 failed, 5 timed out, 6 output limit, and 7 low confidence
Always check the completion flag: a newly configured engine has not completed recognition even if its default numeric status is zero
Remarks
The report does not rerun OCR or switch the selected page
The configured engine's reports describe its latest run, so a later skipped page may retain the prior engine report; the top-level page action identifies the latest page decision
OCRDocumentEx leaves the report for its last attempted page; its integer return value provides the number of processed pages
errorCode reflects LastErrorCode; read diagnostics immediately after the OCR call because other API calls can change the library's last error code
The DLL wide pointer follows the instance's wide return-buffer lifetime; the A variant returns UTF-8 and follows the corresponding ANSI return-buffer lifetime
See also
OCRPageEx, OCRDocumentEx, ConfigureTesseractOCR, LastOCRError