Scalable and Random-Access PDF Loading
HotPDF keeps large loaded files responsive by delaying stream copies and compressed-object parsing until the data is required
Windowed local-file mapping
LoadFromFile uses a shared read-only memory mapping when the source reaches LocalFileMappingThreshold, which defaults to 8 MiB
The parser maps fixed-size regions on demand, reuses them through a bounded least-recently-used cache, and copies across region boundaries without mapping the complete file into the process address space
LocalFileMappingWindowSize defaults to 16 MiB and LocalFileMappingCacheWindows defaults to four windows; a zero threshold disables mapping, while mapping failures automatically fall back to TFileStream
Unencrypted stream bodies share the same mapping as immutable source ranges, avoiding a separate file open for each lazy read; mappings allow atomic rename or deletion while retaining a stable view of the opened bytes
LastLoadUsedMemoryMapping reports the selected path and GetLastMappedFileStatistics returns mapped size, window geometry, reads, bytes, mappings, cache hits, and any positional-read fallback activity
File-backed stream windows
Unencrypted documents opened with LoadFromFile retain immutable stream bodies as source-file ranges
OpenLoadedStreamSourceSlice creates an independent read-only view over an unfiltered source range without copying or materializing the lazy stream
The first write materializes the stream in memory when it is at or below LoadedStreamMemoryThreshold, or in an owned THPDFCompressedSpillStream when it is larger
Spill streams compress independent 256 KiB blocks, keep incompressible blocks raw, reuse suitable block slots after random overwrites, and drain a 4 MiB bounded queue in 1 MiB asynchronous batches
Lazy object streams
Each decoded /Type /ObjStm container is indexed by authoritative xref object number without parsing every member during open
Catalog, page-tree, and later link resolution materialize individual members once and reuse their parsed objects
A full rewrite automatically materializes every remaining member, while an unchanged file-backed save can preserve the original bytes directly
ReleaseDecodedStreamsAfterSave releases an object-stream container only after every indexed member is materialized, and ReleaseLoadedDecodedStreamCaches exposes the same safe release pass on demand
Lazy page trees
Ordinary page trees with at least LazyPageTreeThreshold declared pages retain compact page placeholders instead of recursively expanding every /Kids branch during open
Random page access follows subtree /Count values through an indirect-object hash index and caches the resolved page, while structural mutation and save paths expand the complete tree in document order
GetLoadedPageTreeStatistics exposes materialization, cache, traversal, depth, storage, and full-expansion counters
Custom random access
Derive from THPDFRandomAccessSource and implement GetSize plus ReadAt to load from HTTP ranges, databases, archives, or virtual storage; override IsRangeAvailable when fetched ranges can arrive independently
LoadFromRandomAccessSource adapts the source to HotPDF's seekable parser, can optionally own it for the document lifetime, and automatically coalesces adjacent reads into cached 256 KiB source ranges
The 2 MiB range cache combines least-recently-used order with bounded adaptive frequency admission, parser reads and background transport access are serialized for compatibility with non-thread-safe sources, and upcoming ranges are prefetched asynchronously between parser reads
A foreground seek outside the active prefetch range cancels that work before waiting for the source lock; override ReadAtCancellable to propagate cancellation into HTTP, database, or storage transport operations
Adaptive read-ahead grows from 1 to 2, 4, and 8 blocks during sustained forward parser access, is capped by cache capacity, and is suppressed immediately after random or backward seeks before recovering gradually
Wrap a source explicitly in THPDFCoalescingRandomAccessSource to select block and cache sizes, tune sequential tolerance and maximum read-ahead, schedule or cancel ranges, disable adaptive or asynchronous prefetch, clear cached data, and read THPDFRangeCacheStatistics
Direct stream transfer
HPDFTransferStream dispatches memory, mapped-file, lazy-file, adaptive-cache, and random-access sources to storage-aware transfer paths that write retained bytes directly without an intermediate copy
Unencrypted PDF stream serialization uses the same path automatically, avoiding a complete duplicate TMemoryStream for source-backed stream bodies; encryption continues to stage transformed bytes because the on-disk length changes
Parser scratch arena
Each THotPDF instance retains 64 KiB blocks for short-lived parser tokens and growing scratch buffers, resets their bump positions at the start of every load, and releases token lifetimes with constant-time marks
Persistent PDF objects retain normal ownership and never point into arena storage, while dictionary keys bypass the former temporary THPDFNameObject allocation entirely
Span-based simple-object parsing
HPDFParserParseSimpleDictionary and HPDFParserParseSimpleArray scan immutable source spans and share one cursor across nested containers instead of copying token and container substrings
Names and strings materialize only when their persistent PDF objects are created, numeric values convert directly from source bytes, and dictionary insertion uses an operation-local open-addressed index while retaining case-sensitive replacement semantics
The WithStatistics overloads return THPDFParserViewStatistics with source, token-view, materialized-string, object, container, depth, and error counters for workload diagnostics
Document string interning
Each load creates a fresh document-scoped pool for repeated PDF names, content-stream operators, and immutable literal or hexadecimal strings up to 64 bytes
Matching is case-sensitive and byte-exact, pooled AnsiString values remain safe under copy-on-write assignment, and persistent PDF objects keep their existing ownership
The pool admits at most 65,536 distinct values and 4 MiB of string bytes; empty, longer, or over-budget values bypass interning without changing their content
Independent cache admission
THPDFCacheAdmissionPolicy gives raw compressed ranges, decoded images, compiled display lists, and rendered raster pages separate enabled, entry-count, total-byte, and per-entry limits
The default policies retain at most 8 raw ranges in 2 MiB, 256 decoded images in 32 MiB, 64 display lists in 64 MiB, and 8 raster pages in 128 MiB, with an additional per-entry ceiling for every layer
SetCacheAdmissionPolicy validates and applies one layer without changing the other three, rejects oversized entries without affecting the current operation, and trims inactive least-recently-used entries when a bound shrinks
Adaptive admission compares candidates with required least-recently-used victims through a fixed-memory frequency sketch, bypasses colder candidates, favors recency on ties, and periodically ages history as the working set changes
GetCacheAdmissionStatistics reports active limits, occupancy, hits, misses, admissions, rejections, evictions, frequency decisions, and aging cycles, while ResetCacheAdmissionPolicies restores all defaults
Progressive linearized availability
GetProgressiveLinearizedLoadInfo reuses the strict range planner for a bounded first-object probe, validates /L, /H, /O, /E, /N, /T, and the containing indirect object, and reports the self-contained first-page section independently of the main cross-reference tail
ReadProgressiveLinearizedFirstPageSection copies [0, /E) in 64 KiB requests as soon as that section is available, accepts legal short reads, and checks operation cancellation between reads
PlanProgressiveLinearizedPageRanges plans any zero-based page from raw or Flate primary and optional overflow hint streams, accepts interoperable byte-aligned rows with a combined shared-object sequence and the continuous MSB-first split sequence described by Annex F, and rejects conflicting dual interpretations
The plan contains the complete [0, /E) bootstrap, hint objects, selected page objects, deduplicated shared groups, and exact traditional cross-reference fragments or the complete cross-reference stream object, accepts LF and CRLF entry-zero prefixes, then sorts and merges overlapping or caller-approved nearby ranges
THPDFLinearizedPageRangePlan.DependencyRanges preserves selected shared-object intervals independently of merged transfer ranges so a caller can distinguish first-paint data from dependencies even when coalescing produces an lrkMixed range
Every planned interval is a THPDFLinearizedByteRange carrying Offset, Length, Kind, and Available, and both transfer and dependency lists are exposed as THPDFLinearizedByteRangeArray values
THPDFLinearizedRangeKind classifies each interval as bootstrap, primary or overflow hint, page objects, shared objects, cross-reference, or mixed data
THPDFLinearizedPageRedrawScheduler.RefreshAvailableRanges reports the first-paint transition, newly arrived dependencies after first paint, and final page completion without repeating redraw requests when availability is unchanged
Each refresh returns a THPDFLinearizedRedrawUpdate whose FirstPaintScheduled, RedrawScheduled, and PageCompleted flags plus pending and arrived dependency counts tell the caller which repaint pass to run
THPDFLinearizedPageRangeOptions bounds decoded hints, page and shared-group counts, shared references, planned ranges, total bytes, and planner memory; AllowRowAlignedHints defaults to True for interoperable input and can be disabled for Annex F-only validation, while the default planner-memory ceiling is 256 MiB and an encoding that cannot be disambiguated within that ceiling fails closed
THPDFLinearizedPageRangePlan reports typed status, hint encoding, page and object geometry, transfer bytes and ratio, range availability, cross-reference format, explicit dependency ranges, and diagnostics without taking ownership of the source
THPDFLinearizedHintEncoding names the decoded hint-table interpretation, separating interoperable row-aligned combined data from the Annex F continuous split sequence and indistinguishable-compatible input accepted for both readings
THPDFProgressiveLoadStatus distinguishes an unavailable header, a non-linearized or invalid source, missing first-page ranges, first-page readiness, and complete-document readiness
Parallel image and high-ratio stream optimization
OptimizeLoadedStreams can select frmHighRatio to establish a deterministic four-strategy zlib baseline, run bounded Zopfli block splitting, and retain Zopfli only when it is strictly smaller
OptimizeLoadedImagesParallel decodes eligible image streams into bounded adaptive storage, compresses up to the caller's worker limit, and commits results serially in source-object order only when the requested byte saving is achieved; its telemetry aggregates Zopfli attempts, selections, budget rejections, bytes saved over zlib, and elapsed time
Diagnostics
GetLoadedStreamCacheInforeports source-backed, memory-materialized, and disk-materialized stream countsGetLoadedObjectStreamCacheInforeports indexed and materialized members, cache hits and misses, retained and peak decoded bytes, and per-container evictions and reloadsGetLoadedPageObjectGraphReleaseStatisticsreports page release requests, released pages and objects, reloads, retention reasons, and the last released pageGetLoadedObjectLifecycleStatisticsseparates clean, dirty, released, and must-write objectsGetLastMappedFileStatisticsreports local mapping activity and window-cache reuse for the last file loadGetLastParserArenaStatisticsreports retained blocks and bytes, peak scratch use, allocation and reuse counts, requested bytes, parsed tokens, and eliminated temporary objects for the most recent loadGetLastParserBudgetStatisticsreports configured token and container limits, token traffic, malformed-token counts, observed peaks, and the terminal limit breach for the most recent loadGetDocumentStringInternStatisticsreports category requests, hits, misses, bypasses, retained entries and bytes, and bytes reused during the current loadGetCacheAdmissionStatisticsreports independent occupancy, traffic, admissions, rejections, evictions, frequency decisions, and aging cycles for each cache layerTHPDFRangeCacheStatisticsreports source requests and bytes, cache hits and misses, prefetch requests, completions, cancellations, access-pattern classification, current and peak read-ahead windows, direct transfer calls and bytes, and current cache occupancyTHPDFLoadedStreamCacheInforeports compressed spill logical and physical bytes, current block count, asynchronous batch and block counts, producer waits, and peak queued bytes
See also: Span-Based Simple-Object Parsing, Bounded PDF Parser Budgets, Lazy Page-Tree Loading, Released Page Object Graphs, Bounded Object-Stream Cache, Windowed Local File Mapping, LoadFromRandomAccessSource, THPDFCoalescingRandomAccessSource, THPDFCompressedSpillStream, Cache Admission Policies, HPDFTransferStream, GetLastParserArenaStatistics, GetDocumentStringInternStatistics, GetProgressiveLinearizedLoadInfo, ReadProgressiveLinearizedFirstPageSection, OptimizeLoadedImagesParallel, OpenLoadedStreamSourceSlice, SourceBacked, LoadedStreamMemoryThreshold, ReleaseLoadedDecodedStreamCaches, DecompressAllObjectStreams