ApplyOCRTextLayerEx

OCR, text, document editing

Description

Validates provider-neutral OCR JSON, prepares invisible text in a separate buffer, and commits an owned searchable text layer to one page

Syntax

Delphi

Function TPDFlib.ApplyOCRTextLayerEx(Page: Integer; Const ResultJSON,
  FontName: WideString; MinConfidence: Double; Options: Integer): Integer;

ActiveX

Function PDFlib::ApplyOCRTextLayerEx(Page As Long, ResultJSON As String,
  FontName As String, MinConfidence As Double, Options As Long) As Long

DLL

int DLApplyOCRTextLayerEx(int InstanceID, int Page, const wchar_t* ResultJSON,
  const wchar_t* FontName, double MinConfidence, int Options);
int DLApplyOCRTextLayerExA(int InstanceID, int Page, const char* ResultJSON,
  const char* FontName, double MinConfidence, int Options);

Parameters

PageThe one-based target page
ResultJSONThe version-1 schema described by ApplyOCRTextLayer, optionally including verified baseline endpoints and hierarchy IDs
FontNameAn embeddable TrueType font or an empty string for installed fallback selection
MinConfidenceThe inclusive word-confidence threshold from 0 through 1
OptionsA bitwise combination of the layer policies below; zero appends a new owned layer

Options

1 — PDF_OCR_SKIP_EXISTING_TEXTReturn success without applying OCR when a retained, non-owned content stream or invoked Form XObject contains text, including invisible text
2 — PDF_OCR_REPLACE_OWNED_LAYERReplace every existing layer with the exact library OCR marker; append if no owned OCR layer exists
3Skip pages with non-owned text, while allowing owned-only OCR pages to be replaced

Return values

1A layer was committed, the existing-text policy skipped the page, or no word met the confidence threshold
0The input, geometry, options, font encoding, or layer write failed

Remarks

Unknown option bits are rejected

Source coordinates use the rendered image's top-left origin and the selected render box; page rotation and nonzero box offsets are mapped back to raw PDF user space

An optional baseline array contains [x1, y1, x2, y2] in source units and defines the text direction and advance; a complete baseline must not also carry a nonzero textAngle

Every word boundary and baseline is checked before confidence filtering; an invalid low-confidence word still causes the operation to fail

Font fitting and missing-glyph handling finish before the new content layer is committed; no partially drawn OCR words replace the old layer

Only streams with all three exact fields /PDFlibOCROwner /PDFlibPas, /PDFlibOCRType /InvisibleText, and /PDFlibOCRVersion 1 are owned OCR layers

Replacement creates a new contents array and new stream, preserving non-owned content and the streams used by sibling pages

The OCR layer uses rendering mode 3 and an isolated graphics state; placement cancels the retained content's current transformation matrix before applying the word matrix

An empty accepted result leaves old layers intact; use GetOCRDiagnosticsJSON to distinguish empty, skipped, appended, and replaced

Searchability and unchanged visible pixels are verified after saving and reloading the document

See also

OCRPageEx, OCRDocumentEx, ConfigureTesseractOCR, GetOCRDiagnosticsJSON, LastOCRError