Typed Table Extraction
THotPDF.ExtractLoadedTypedTables converts semantic table rows into a public canonical column grid without modifying the loaded document
Table continuity
Compatible adjacent fragments can add canonical columns, turning missing intermediate starts into explicit ColumnSpan values instead of discarding merged headings
Compatible tables on consecutive pages become one table, whilst matching leading header rows remain present and carry IsRepeatedHeader
Typed values and confidence
Each cell retains original Unicode text, page-space bounds, row and column spans, normalised value, typed storage, currency code and confidence
Inference recognises empty, string, integer, number, boolean, date, date-time, currency and percentage values under caller-selected decimal, thousands and date-order rules
Bounded export
ExportLoadedTypedTables writes expanded RFC 4180-style CSV rows or metadata-rich JSON through a bounded staging stream and rollback-safe final publication
Page, glyph, source-table, output-table, row, cell, text, output-byte, confidence and cancellation limits are checked before publishing results
XLSX workbooks
The HPDFXLSXExport unit exports THPDFTypedTables as an Office Open XML workbook without installing Office or loading an external spreadsheet library
HPDFExportTypedTablesXLSX writes to a readable, writable, seekable stream at its current position, staging the complete workbook before publication and restoring overwritten bytes when a recoverable write failure occurs
HPDFExportTypedTablesXLSXFile uses an exclusively created temporary file beside the target, flushes it, and atomically replaces the target only after successful export
HPDFExportLoadedTypedTablesXLSX combines extraction with stream export and rejects a destination that aliases the loaded source
Each extracted table becomes a worksheet named Table N, with numeric and boolean cells, date and date-time formats, currency codes in number formats, percentage formats, and validated merged ranges
Integers with more than 15 decimal digits are stored as text to preserve their exact value because Excel numeric cells have limited precision
Strings remain literal text, including values beginning with =, and XML control characters and literal Office Open XML escape sequences are preserved through SpreadsheetML escaping
By default, a hidden Source metadata worksheet records the workbook cell address, zero-based source page and table indexes, original and normalised text, value kind, currency code, page-space bounds, and confidence for every source cell
THPDFXLSXExportOptions.Default permits up to 4,096 tables, 1,000,000 rows, 1,000,000 cells, and 256 MiB of output, with an optional THPDFCancellationToken and an option to omit the metadata worksheet
Export rejects overlapping merges, unordered or duplicate column indexes, non-finite numbers, invalid currency codes, dates outside 1900 through 9999, text longer than 32,767 UTF-16 units, and worksheets outside Excel's 1,048,576-row or 16,384-column limits
The metadata sheet has one header row, so enabling it additionally limits source cells to 1,048,575 across the workbook; an empty extraction still produces one visible sheet
ZIP members use the stored compression method, so output size can be larger than compressed workbooks; the exporter retains neither PDF visual styling nor charts, formulas, macros, or embedded images
JSON jobs and the C ABI expose the same workbook pipeline through tables.export with format: "xlsx" and the shared extraction and output budgets