ApplyOCRTextLayerEx
OCR, text, document editing
Description
Validates provider-neutral OCR JSON, prepares invisible text in a separate buffer, and commits an owned searchable text layer to one page
Syntax
Delphi
Function TPDFlib.ApplyOCRTextLayerEx(Page: Integer; Const ResultJSON,
FontName: WideString; MinConfidence: Double; Options: Integer): Integer;
ActiveX
Function PDFlib::ApplyOCRTextLayerEx(Page As Long, ResultJSON As String,
FontName As String, MinConfidence As Double, Options As Long) As Long
DLL
int DLApplyOCRTextLayerEx(int InstanceID, int Page, const wchar_t* ResultJSON,
const wchar_t* FontName, double MinConfidence, int Options);
int DLApplyOCRTextLayerExA(int InstanceID, int Page, const char* ResultJSON,
const char* FontName, double MinConfidence, int Options);
Parameters
| Page | The one-based target page |
|---|---|
| ResultJSON | The version-1 schema described by ApplyOCRTextLayer, optionally including verified baseline endpoints and hierarchy IDs |
| FontName | An embeddable TrueType font or an empty string for installed fallback selection |
| MinConfidence | The inclusive word-confidence threshold from 0 through 1 |
| Options | A bitwise combination of the layer policies below; zero appends a new owned layer |
Options
| 1 — PDF_OCR_SKIP_EXISTING_TEXT | Return success without applying OCR when a retained, non-owned content stream or invoked Form XObject contains text, including invisible text |
|---|---|
| 2 — PDF_OCR_REPLACE_OWNED_LAYER | Replace every existing layer with the exact library OCR marker; append if no owned OCR layer exists |
| 3 | Skip pages with non-owned text, while allowing owned-only OCR pages to be replaced |
Return values
| 1 | A layer was committed, the existing-text policy skipped the page, or no word met the confidence threshold |
|---|---|
| 0 | The input, geometry, options, font encoding, or layer write failed |
Remarks
Unknown option bits are rejected
Source coordinates use the rendered image's top-left origin and the selected render box; page rotation and nonzero box offsets are mapped back to raw PDF user space
An optional baseline array contains [x1, y1, x2, y2] in source units and defines the text direction and advance; a complete baseline must not also carry a nonzero textAngle
Every word boundary and baseline is checked before confidence filtering; an invalid low-confidence word still causes the operation to fail
Font fitting and missing-glyph handling finish before the new content layer is committed; no partially drawn OCR words replace the old layer
Only streams with all three exact fields /PDFlibOCROwner /PDFlibPas, /PDFlibOCRType /InvisibleText, and /PDFlibOCRVersion 1 are owned OCR layers
Replacement creates a new contents array and new stream, preserving non-owned content and the streams used by sibling pages
The OCR layer uses rendering mode 3 and an isolated graphics state; placement cancels the retained content's current transformation matrix before applying the word matrix
An empty accepted result leaves old layers intact; use GetOCRDiagnosticsJSON to distinguish empty, skipped, appended, and replaced
Searchability and unchanged visible pixels are verified after saving and reloading the document
See also
OCRPageEx, OCRDocumentEx, ConfigureTesseractOCR, GetOCRDiagnosticsJSON, LastOCRError