Scalable and Random-Access PDF Loading

HotPDF keeps large loaded files responsive by delaying stream copies and compressed-object parsing until the data is required

Windowed local-file mapping

LoadFromFile uses a shared read-only memory mapping when the source reaches LocalFileMappingThreshold, which defaults to 8 MiB

The parser maps fixed-size regions on demand, reuses them through a bounded least-recently-used cache, and copies across region boundaries without mapping the complete file into the process address space

LocalFileMappingWindowSize defaults to 16 MiB and LocalFileMappingCacheWindows defaults to four windows; a zero threshold disables mapping, while mapping failures automatically fall back to TFileStream

Unencrypted stream bodies share the same mapping as immutable source ranges, avoiding a separate file open for each lazy read; mappings allow atomic rename or deletion while retaining a stable view of the opened bytes

LastLoadUsedMemoryMapping reports the selected path and GetLastMappedFileStatistics returns mapped size, window geometry, reads, bytes, mappings, cache hits, and any positional-read fallback activity

File-backed stream windows

Unencrypted documents opened with LoadFromFile retain immutable stream bodies as source-file ranges

OpenLoadedStreamSourceSlice creates an independent read-only view over an unfiltered source range without copying or materializing the lazy stream

The first write materializes the stream in memory when it is at or below LoadedStreamMemoryThreshold, or in an owned THPDFCompressedSpillStream when it is larger

Spill streams compress independent 256 KiB blocks, keep incompressible blocks raw, reuse suitable block slots after random overwrites, and drain a 4 MiB bounded queue in 1 MiB asynchronous batches

Lazy object streams

Each decoded /Type /ObjStm container is indexed by authoritative xref object number without parsing every member during open

Catalog, page-tree, and later link resolution materialize individual members once and reuse their parsed objects

A full rewrite automatically materializes every remaining member, while an unchanged file-backed save can preserve the original bytes directly

ReleaseDecodedStreamsAfterSave releases an object-stream container only after every indexed member is materialized, and ReleaseLoadedDecodedStreamCaches exposes the same safe release pass on demand

Lazy page trees

Ordinary page trees with at least LazyPageTreeThreshold declared pages retain compact page placeholders instead of recursively expanding every /Kids branch during open

Random page access follows subtree /Count values through an indirect-object hash index and caches the resolved page, while structural mutation and save paths expand the complete tree in document order

GetLoadedPageTreeStatistics exposes materialization, cache, traversal, depth, storage, and full-expansion counters

Custom random access

Derive from THPDFRandomAccessSource and implement GetSize plus ReadAt to load from HTTP ranges, databases, archives, or virtual storage; override IsRangeAvailable when fetched ranges can arrive independently

LoadFromRandomAccessSource adapts the source to HotPDF's seekable parser, can optionally own it for the document lifetime, and automatically coalesces adjacent reads into cached 256 KiB source ranges

The 2 MiB range cache combines least-recently-used order with bounded adaptive frequency admission, parser reads and background transport access are serialized for compatibility with non-thread-safe sources, and upcoming ranges are prefetched asynchronously between parser reads

A foreground seek outside the active prefetch range cancels that work before waiting for the source lock; override ReadAtCancellable to propagate cancellation into HTTP, database, or storage transport operations

Adaptive read-ahead grows from 1 to 2, 4, and 8 blocks during sustained forward parser access, is capped by cache capacity, and is suppressed immediately after random or backward seeks before recovering gradually

Wrap a source explicitly in THPDFCoalescingRandomAccessSource to select block and cache sizes, tune sequential tolerance and maximum read-ahead, schedule or cancel ranges, disable adaptive or asynchronous prefetch, clear cached data, and read THPDFRangeCacheStatistics

Direct stream transfer

HPDFTransferStream dispatches memory, mapped-file, lazy-file, adaptive-cache, and random-access sources to storage-aware transfer paths that write retained bytes directly without an intermediate copy

Unencrypted PDF stream serialization uses the same path automatically, avoiding a complete duplicate TMemoryStream for source-backed stream bodies; encryption continues to stage transformed bytes because the on-disk length changes

Parser scratch arena

Each THotPDF instance retains 64 KiB blocks for short-lived parser tokens and growing scratch buffers, resets their bump positions at the start of every load, and releases token lifetimes with constant-time marks

Persistent PDF objects retain normal ownership and never point into arena storage, while dictionary keys bypass the former temporary THPDFNameObject allocation entirely

Span-based simple-object parsing

HPDFParserParseSimpleDictionary and HPDFParserParseSimpleArray scan immutable source spans and share one cursor across nested containers instead of copying token and container substrings

Names and strings materialize only when their persistent PDF objects are created, numeric values convert directly from source bytes, and dictionary insertion uses an operation-local open-addressed index while retaining case-sensitive replacement semantics

The WithStatistics overloads return THPDFParserViewStatistics with source, token-view, materialized-string, object, container, depth, and error counters for workload diagnostics

Document string interning

Each load creates a fresh document-scoped pool for repeated PDF names, content-stream operators, and immutable literal or hexadecimal strings up to 64 bytes

Matching is case-sensitive and byte-exact, pooled AnsiString values remain safe under copy-on-write assignment, and persistent PDF objects keep their existing ownership

The pool admits at most 65,536 distinct values and 4 MiB of string bytes; empty, longer, or over-budget values bypass interning without changing their content

Independent cache admission

THPDFCacheAdmissionPolicy gives raw compressed ranges, decoded images, compiled display lists, and rendered raster pages separate enabled, entry-count, total-byte, and per-entry limits

The default policies retain at most 8 raw ranges in 2 MiB, 256 decoded images in 32 MiB, 64 display lists in 64 MiB, and 8 raster pages in 128 MiB, with an additional per-entry ceiling for every layer

SetCacheAdmissionPolicy validates and applies one layer without changing the other three, rejects oversized entries without affecting the current operation, and trims inactive least-recently-used entries when a bound shrinks

Adaptive admission compares candidates with required least-recently-used victims through a fixed-memory frequency sketch, bypasses colder candidates, favors recency on ties, and periodically ages history as the working set changes

GetCacheAdmissionStatistics reports active limits, occupancy, hits, misses, admissions, rejections, evictions, frequency decisions, and aging cycles, while ResetCacheAdmissionPolicies restores all defaults

Progressive linearized availability

GetProgressiveLinearizedLoadInfo reuses the strict range planner for a bounded first-object probe, validates /L, /H, /O, /E, /N, /T, and the containing indirect object, and reports the self-contained first-page section independently of the main cross-reference tail

ReadProgressiveLinearizedFirstPageSection copies [0, /E) in 64 KiB requests as soon as that section is available, accepts legal short reads, and checks operation cancellation between reads

PlanProgressiveLinearizedPageRanges plans any zero-based page from raw or Flate primary and optional overflow hint streams, accepts interoperable byte-aligned rows with a combined shared-object sequence and the continuous MSB-first split sequence described by Annex F, and rejects conflicting dual interpretations

The plan contains the complete [0, /E) bootstrap, hint objects, selected page objects, deduplicated shared groups, and exact traditional cross-reference fragments or the complete cross-reference stream object, accepts LF and CRLF entry-zero prefixes, then sorts and merges overlapping or caller-approved nearby ranges

THPDFLinearizedPageRangePlan.DependencyRanges preserves selected shared-object intervals independently of merged transfer ranges so a caller can distinguish first-paint data from dependencies even when coalescing produces an lrkMixed range

Every planned interval is a THPDFLinearizedByteRange carrying Offset, Length, Kind, and Available, and both transfer and dependency lists are exposed as THPDFLinearizedByteRangeArray values

THPDFLinearizedRangeKind classifies each interval as bootstrap, primary or overflow hint, page objects, shared objects, cross-reference, or mixed data

THPDFLinearizedPageRedrawScheduler.RefreshAvailableRanges reports the first-paint transition, newly arrived dependencies after first paint, and final page completion without repeating redraw requests when availability is unchanged

Each refresh returns a THPDFLinearizedRedrawUpdate whose FirstPaintScheduled, RedrawScheduled, and PageCompleted flags plus pending and arrived dependency counts tell the caller which repaint pass to run

THPDFLinearizedPageRangeOptions bounds decoded hints, page and shared-group counts, shared references, planned ranges, total bytes, and planner memory; AllowRowAlignedHints defaults to True for interoperable input and can be disabled for Annex F-only validation, while the default planner-memory ceiling is 256 MiB and an encoding that cannot be disambiguated within that ceiling fails closed

THPDFLinearizedPageRangePlan reports typed status, hint encoding, page and object geometry, transfer bytes and ratio, range availability, cross-reference format, explicit dependency ranges, and diagnostics without taking ownership of the source

THPDFLinearizedHintEncoding names the decoded hint-table interpretation, separating interoperable row-aligned combined data from the Annex F continuous split sequence and indistinguishable-compatible input accepted for both readings

THPDFProgressiveLoadStatus distinguishes an unavailable header, a non-linearized or invalid source, missing first-page ranges, first-page readiness, and complete-document readiness

Parallel image and high-ratio stream optimization

OptimizeLoadedStreams can select frmHighRatio to establish a deterministic four-strategy zlib baseline, run bounded Zopfli block splitting, and retain Zopfli only when it is strictly smaller

OptimizeLoadedImagesParallel decodes eligible image streams into bounded adaptive storage, compresses up to the caller's worker limit, and commits results serially in source-object order only when the requested byte saving is achieved; its telemetry aggregates Zopfli attempts, selections, budget rejections, bytes saved over zlib, and elapsed time

Diagnostics

See also: Span-Based Simple-Object Parsing, Bounded PDF Parser Budgets, Lazy Page-Tree Loading, Released Page Object Graphs, Bounded Object-Stream Cache, Windowed Local File Mapping, LoadFromRandomAccessSource, THPDFCoalescingRandomAccessSource, THPDFCompressedSpillStream, Cache Admission Policies, HPDFTransferStream, GetLastParserArenaStatistics, GetDocumentStringInternStatistics, GetProgressiveLinearizedLoadInfo, ReadProgressiveLinearizedFirstPageSection, OptimizeLoadedImagesParallel, OpenLoadedStreamSourceSlice, SourceBacked, LoadedStreamMemoryThreshold, ReleaseLoadedDecodedStreamCaches, DecompressAllObjectStreams