Span-Based Simple-Object Parsing

The HPDFParser unit parses standalone PDF dictionaries and arrays without copying token or nested-container substrings

Entry points

function HPDFParserParseSimpleDictionary(
  const Data: AnsiString;
  RecursionLevel: Integer = 0): THPDFDictionaryObject;

function HPDFParserParseSimpleDictionaryWithStatistics(
  const Data: AnsiString;
  out Statistics: THPDFParserViewStatistics;
  RecursionLevel: Integer = 0): THPDFDictionaryObject;

function HPDFParserParseSimpleArray(
  const Data: AnsiString): THPDFArrayObject;

function HPDFParserParseSimpleArrayWithStatistics(
  const Data: AnsiString;
  out Statistics: THPDFParserViewStatistics): THPDFArrayObject;

The returned dictionary or array is owned by the caller and contains normal persistent HotPDF objects with no dependency on the source string lifetime

View statistics

THPDFParserViewStatistics exposes SourceBytes, TokenViewCount, MaterializedStringCount, MaterializedBytes, ObjectCount, ContainerCount, MaxDepth, ErrorCount, and Valid

Token views count bounded source ranges, while materialised-string counters cover bytes copied into durable PDF names, keys, and strings

Supported objects

The parser recognises names, integers, real numbers, booleans, nulls, indirect references, literal strings, hexadecimal strings, arrays, and dictionaries

Literal strings preserve nested parentheses, escaped delimiters, octal escapes, and escaped line continuations, while PDF whitespace and comments are ignored

Nested containers share one cursor over the immutable input, recursion is bounded, and duplicate dictionary keys replace earlier values with case-sensitive matching

Performance characteristics

Names and strings materialise once at their persistent object boundary, numeric conversion reads source bytes directly, and an operation-local open-addressed index prevents quadratic duplicate scans while constructing large dictionaries

A 30,000-pair Win32 regression fixture reduced five-pass allocation bytes from 50,866,510 to 17,936,800 and elapsed time from 3,656 ms to 46 ms

See also: Scalable and Random-Access PDF Loading, Bounded PDF Parser Budgets