Span-Based Simple-Object Parsing
The HPDFParser unit parses standalone PDF dictionaries and arrays without copying token or nested-container substrings
Entry points
function HPDFParserParseSimpleDictionary(
const Data: AnsiString;
RecursionLevel: Integer = 0): THPDFDictionaryObject;
function HPDFParserParseSimpleDictionaryWithStatistics(
const Data: AnsiString;
out Statistics: THPDFParserViewStatistics;
RecursionLevel: Integer = 0): THPDFDictionaryObject;
function HPDFParserParseSimpleArray(
const Data: AnsiString): THPDFArrayObject;
function HPDFParserParseSimpleArrayWithStatistics(
const Data: AnsiString;
out Statistics: THPDFParserViewStatistics): THPDFArrayObject;
The returned dictionary or array is owned by the caller and contains normal persistent HotPDF objects with no dependency on the source string lifetime
View statistics
THPDFParserViewStatistics exposes SourceBytes, TokenViewCount, MaterializedStringCount, MaterializedBytes, ObjectCount, ContainerCount, MaxDepth, ErrorCount, and Valid
Token views count bounded source ranges, while materialised-string counters cover bytes copied into durable PDF names, keys, and strings
Supported objects
The parser recognises names, integers, real numbers, booleans, nulls, indirect references, literal strings, hexadecimal strings, arrays, and dictionaries
Literal strings preserve nested parentheses, escaped delimiters, octal escapes, and escaped line continuations, while PDF whitespace and comments are ignored
Nested containers share one cursor over the immutable input, recursion is bounded, and duplicate dictionary keys replace earlier values with case-sensitive matching
Performance characteristics
Names and strings materialise once at their persistent object boundary, numeric conversion reads source bytes directly, and an operation-local open-addressed index prevents quadratic duplicate scans while constructing large dictionaries
A 30,000-pair Win32 regression fixture reduced five-pass allocation bytes from 50,866,510 to 17,936,800 and elapsed time from 3,656 ms to 46 ms
See also: Scalable and Random-Access PDF Loading, Bounded PDF Parser Budgets