Structured Loaded-Page Export
THotPDF.ExportLoadedPagesStructured converts selected loaded pages into one UTF-8 HTML, XHTML, XML, or JSON document without mutating the PDF
One semantic model
Every format consumes one accessible reading model so structure order, document language and title, headings, lists, tables, figures, marked-content replacements, object references, navigation, page labels, annotations, and form controls remain aligned across exporters
Tagged structure is preferred when it provides a usable projection for a page, while geometric paragraph, heading, list, caption, column, and table analysis is used only as the configured fallback
Generation completes in a bounded staging stream, and publication snapshots any destination bytes that would be overwritten so a failed final write can restore the original stream state
Each span carries safely escaped Unicode text plus optional page-space bounds, source content-stream object identity, source font resource, and point size
Accessible structure and interactions
HTML and XHTML use semantic elements and stable page anchors, while XML and JSON retain equivalent begin, end, content, object-reference, interaction, outline, and page-label records
Inline and named marked-content properties are resolved across page streams and recursive Form XObjects, ActualText replaces visual glyph text, artifacts stay outside the accessible model, and figure alternative text is associated through tagged content rather than visual order
Safe http, https, mailto, and local destinations can be exported, while executable or control-character-bearing URIs are omitted
Form names, labels, options, states, and ordinary values can be retained, but password, file-select, NoExport, and signature payloads are never placed in the shared export model
Images
Image XObject occurrences retain their resource name, object number, intrinsic pixel size, and transformed page-space bounds
simEmbeddedPNG decodes each supported image through the normal bounded image pipeline and emits a PNG data URI, while simMetadataOnly avoids pixel materialization and simOmit excludes images
Bounded atomic output
Page, glyph, image-count, structure-node, outline, per-page interaction, per-field option, image-byte, aggregate-image, output-byte, and cancellation limits are enforced before the completed buffer is copied to the caller stream
Extraction and budget failures leave the destination unchanged, while a destination write failure attempts to restore its original position and size