Typed Table Extraction

THotPDF.ExtractLoadedTypedTables converts semantic table rows into a public canonical column grid without modifying the loaded document

Table continuity

Compatible adjacent fragments can add canonical columns, turning missing intermediate starts into explicit ColumnSpan values instead of discarding merged headings

Compatible tables on consecutive pages become one table, whilst matching leading header rows remain present and carry IsRepeatedHeader

Typed values and confidence

Each cell retains original Unicode text, page-space bounds, row and column spans, normalised value, typed storage, currency code and confidence

Inference recognises empty, string, integer, number, boolean, date, date-time, currency and percentage values under caller-selected decimal, thousands and date-order rules

Bounded export

ExportLoadedTypedTables writes expanded RFC 4180-style CSV rows or metadata-rich JSON through a bounded staging stream and rollback-safe final publication

Page, glyph, source-table, output-table, row, cell, text, output-byte, confidence and cancellation limits are checked before publishing results

XLSX workbooks

The HPDFXLSXExport unit exports THPDFTypedTables as an Office Open XML workbook without installing Office or loading an external spreadsheet library

HPDFExportTypedTablesXLSX writes to a readable, writable, seekable stream at its current position, staging the complete workbook before publication and restoring overwritten bytes when a recoverable write failure occurs

HPDFExportTypedTablesXLSXFile uses an exclusively created temporary file beside the target, flushes it, and atomically replaces the target only after successful export

HPDFExportLoadedTypedTablesXLSX combines extraction with stream export and rejects a destination that aliases the loaded source

Each extracted table becomes a worksheet named Table N, with numeric and boolean cells, date and date-time formats, currency codes in number formats, percentage formats, and validated merged ranges

Integers with more than 15 decimal digits are stored as text to preserve their exact value because Excel numeric cells have limited precision

Strings remain literal text, including values beginning with =, and XML control characters and literal Office Open XML escape sequences are preserved through SpreadsheetML escaping

By default, a hidden Source metadata worksheet records the workbook cell address, zero-based source page and table indexes, original and normalised text, value kind, currency code, page-space bounds, and confidence for every source cell

THPDFXLSXExportOptions.Default permits up to 4,096 tables, 1,000,000 rows, 1,000,000 cells, and 256 MiB of output, with an optional THPDFCancellationToken and an option to omit the metadata worksheet

Export rejects overlapping merges, unordered or duplicate column indexes, non-finite numbers, invalid currency codes, dates outside 1900 through 9999, text longer than 32,767 UTF-16 units, and worksheets outside Excel's 1,048,576-row or 16,384-column limits

The metadata sheet has one header row, so enabling it additionally limits source cells to 1,048,575 across the workbook; an empty extraction still produces one visible sheet

ZIP members use the stored compression method, so output size can be larger than compressed workbooks; the exporter retains neither PDF visual styling nor charts, formulas, macros, or embedded images

JSON jobs and the C ABI expose the same workbook pipeline through tables.export with format: "xlsx" and the shared extraction and output budgets

Related APIs