ConfigureTesseractOCR
OCR, configuration
Description
Creates an instance-owned Tesseract provider and assigns it to OnOCRPage, making the page and document OCR methods usable without an application callback
Syntax
Delphi
Function TPDFlib.ConfigureTesseractOCR(Const Executable,
TessDataDirectory: WideString; DetectOrientation: Boolean= False;
MinOrientationConfidence: Double= 15; TimeoutMS: Integer= 120000): Boolean;
ActiveX
Function PDFlib::ConfigureTesseractOCR(Executable As String,
TessDataDirectory As String, DetectOrientation As Long,
MinOrientationConfidence As Double, TimeoutMS As Long) As Long
DLL
int DLConfigureTesseractOCR(int InstanceID, const wchar_t* Executable,
const wchar_t* TessDataDirectory, int DetectOrientation,
double MinOrientationConfidence, int TimeoutMS);
int DLConfigureTesseractOCRA(int InstanceID, const char* Executable,
const char* TessDataDirectory, int DetectOrientation,
double MinOrientationConfidence, int TimeoutMS);
Parameters
| Executable | The path to an existing Tesseract executable |
|---|---|
| TessDataDirectory | The trained-data directory, or an empty string to let the engine use its normal data lookup |
| DetectOrientation | False by default; true runs OSD before recognition and explicitly corrects the reported 0, 90, 180, or 270 degree clockwise rotation |
| MinOrientationConfidence | A finite non-negative Tesseract OSD confidence score, default 15; this score is distinct from the 0-to-1 word-confidence threshold |
| TimeoutMS | The positive timeout in milliseconds for each engine process, default 120000; an oriented page can run separate OSD and recognition processes |
Return values
Delphi returns true when the configuration is accepted; DLL and ActiveX return 1 for success and 0 for an invalid configuration
Remarks
The engine is optional and is supplied by the application; configuration does not install an executable, download models, or modify the user's Tesseract configuration
The executable path and optional data-directory path are checked during configuration; model availability and actual engine execution are checked when a page is recognized
The requested language needs its trained-data file, such as eng.traineddata; orientation detection also needs osd.traineddata
The Windows adapter launches the executable directly without a shell or visible console, uses private temporary files, enforces time and output limits, and parses hOCR hierarchy and verified line baselines
Orientation detection is disabled by default; when enabled, a missing or low-confidence OSD result fails explicitly rather than guessing a direction
A verified correction rotates supported noninterlaced 8-bit PNG pixels by a quarter turn without interpolation, recognizes the upright image, then maps word boxes and baseline endpoints back to the original image coordinates
The mapped baseline already carries the text direction and does not receive the orientation angle a second time
Invalid reconfiguration preserves the existing provider; accepted reconfiguration replaces the previous owned provider
An application callback assigned afterward remains application-owned and is preserved by ClearTesseractOCR and instance destruction
Example
if not PDF.ConfigureTesseractOCR('C:\Tools\Tesseract\tesseract.exe',
'C:\OCR\tessdata', True, 15, 120000) then
raise Exception.Create(PDF.LastOCRError);
Processed := PDF.OCRDocumentEx('', 300, 'eng', 'Arial', 0.5,
PDF_OCR_SKIP_EXISTING_TEXT or PDF_OCR_REPLACE_OWNED_LAYER);
ReportJSON := PDF.GetOCRDiagnosticsJSON;
See also
ClearTesseractOCR, OCRPageEx, OCRDocumentEx, GetOCRDiagnosticsJSON, OnOCRPage