ExtractStructuredPage
Text, extraction, structured data
Description
Renders one page into a structured text model and serializes it as JSON, XHTML, or CSV
The model groups text into paragraphs and detected tables, with lines and styled spans under paragraphs and rows, cells, and spans under tables
Syntax
Delphi
Function TPDFlib.ExtractStructuredPage(Page, Format,
Options: Integer): WideString;ActiveX
Function PDFlib::ExtractStructuredPage(Page As Long, Format As Long,
Options As Long) As StringDLL
const wchar_t* DLExtractStructuredPage(int InstanceID, int Page,
int Format, int Options);
const char* DLExtractStructuredPageA(int InstanceID, int Page,
int Format, int Options);Parameters
| Page | The one-based page number |
|---|---|
| Format | PDF_STRUCTURED_TEXT_JSON (0), PDF_STRUCTURED_TEXT_XHTML (1), or PDF_STRUCTURED_TEXT_CSV (2) |
| Options | Zero or PDF_STRUCTURED_TEXT_INCLUDE_STYLES (1) to include font, size, colour, and vertical alignment in JSON and XHTML |
JSON hierarchy
{
"version": 1,
"page": 1,
"width": 612,
"height": 792,
"blocks": [
{"type": "paragraph", "bbox": [40, 50, 300, 70],
"lines": [{"bbox": [40, 50, 300, 70], "spans": []}]},
{"type": "table", "columns": 3, "rows": [[]]}
]
}Return values
JSON and XHTML return complete standalone documents
CSV returns detected tables only and is empty when the page has no table
An invalid request or extraction failure returns an empty string
Remarks
Text is rendered once, then every format is produced from the same intermediate hierarchy
Table detection merges normal word gaps inside a cell and requires aligned multi-cell rows, reducing paragraph false positives
Bounding boxes use top-left page coordinates and are expressed in PDF points
DLL return pointers remain valid until the next string-returning call on the same instance
LastErrorCode is 111 for invalid page, format, or options and 515 when extraction fails
See also
ExtractStructuredDocument, ExtractPageTextBlocks, GetPageText