ExtractStructuredPage

Text, extraction, structured data

Description

Renders one page into a structured text model and serializes it as JSON, XHTML, or CSV

The model groups text into paragraphs and detected tables, with lines and styled spans under paragraphs and rows, cells, and spans under tables

Syntax

Delphi

Function TPDFlib.ExtractStructuredPage(Page, Format,
  Options: Integer): WideString;

ActiveX

Function PDFlib::ExtractStructuredPage(Page As Long, Format As Long,
  Options As Long) As String

DLL

const wchar_t* DLExtractStructuredPage(int InstanceID, int Page,
  int Format, int Options);
const char* DLExtractStructuredPageA(int InstanceID, int Page,
  int Format, int Options);

Parameters

PageThe one-based page number
FormatPDF_STRUCTURED_TEXT_JSON (0), PDF_STRUCTURED_TEXT_XHTML (1), or PDF_STRUCTURED_TEXT_CSV (2)
OptionsZero or PDF_STRUCTURED_TEXT_INCLUDE_STYLES (1) to include font, size, colour, and vertical alignment in JSON and XHTML

JSON hierarchy

{
  "version": 1,
  "page": 1,
  "width": 612,
  "height": 792,
  "blocks": [
    {"type": "paragraph", "bbox": [40, 50, 300, 70],
     "lines": [{"bbox": [40, 50, 300, 70], "spans": []}]},
    {"type": "table", "columns": 3, "rows": [[]]}
  ]
}

Return values

JSON and XHTML return complete standalone documents

CSV returns detected tables only and is empty when the page has no table

An invalid request or extraction failure returns an empty string

Remarks

Text is rendered once, then every format is produced from the same intermediate hierarchy

Table detection merges normal word gaps inside a cell and requires aligned multi-cell rows, reducing paragraph false positives

Bounding boxes use top-left page coordinates and are expressed in PDF points

DLL return pointers remain valid until the next string-returning call on the same instance

LastErrorCode is 111 for invalid page, format, or options and 515 when extraction fails

See also

ExtractStructuredDocument, ExtractPageTextBlocks, GetPageText