ExtractStructuredDocument

Text, extraction, structured data

Description

Extracts a page range into one JSON document, one XHTML document, or a sequence of detected CSV tables

Syntax

Delphi

Function TPDFlib.ExtractStructuredDocument(Const PageRange: WideString;
  Format, Options: Integer): WideString;

ActiveX

Function PDFlib::ExtractStructuredDocument(PageRange As String,
  Format As Long, Options As Long) As String

DLL

const wchar_t* DLExtractStructuredDocument(int InstanceID,
  const wchar_t* PageRange, int Format, int Options);
const char* DLExtractStructuredDocumentA(int InstanceID,
  const char* PageRange, int Format, int Options);

Parameters

PageRangeA one-based page range such as 1-3,7, or an empty string for every page
FormatPDF_STRUCTURED_TEXT_JSON (0), PDF_STRUCTURED_TEXT_XHTML (1), or PDF_STRUCTURED_TEXT_CSV (2)
OptionsZero or PDF_STRUCTURED_TEXT_INCLUDE_STYLES (1)

Return values

JSON contains a versioned pages array and XHTML contains one semantic section per selected page

CSV contains only detected tables and is empty when no selected page has one

An invalid range, format, or option and any extraction failure return an empty string

Remarks

Only the current page text model is retained while the output is assembled, so extraction memory for geometry and hierarchy follows the most complex selected page rather than the document page count

The selected page is restored before the call returns

Progress and cooperative cancellation use the normal document operation lifecycle

LastErrorCode is 111 for invalid input and 515 when extraction fails

See also

ExtractStructuredPage, ExtractPageRangeText