Token-Preserving Content-Stream AST

HPDFContentStream provides a bounded lossless syntax model for targeted edits to decoded PDF content streams

Parse and inspect

HPDFParseContentStreamAST creates a THPDFContentStreamAST with default limits of 4,194,304 lexemes and 256 nested containers; callers that need tighter limits can call THPDFContentStreamAST.Create directly

THPDFContentLexemeKind distinguishes whitespace, comments, numbers, names, literal and hexadecimal strings, booleans, nulls, words, container delimiters, inline-image data, and malformed bytes; every input byte belongs to exactly one THPDFContentLexeme

THPDFContentSyntaxNodeKind distinguishes the root, operations, operands, arrays, dictionaries, procedures, inline images, and malformed regions; each THPDFContentSyntaxNode uses parent, first-child, last-child, and next-sibling indexes so traversal does not allocate child arrays

Use GetLexeme, GetLexemeBytes, GetLexemeText, and GetNode with LexemeCount and NodeCount; GetOriginalBytes returns an independent source copy

Valid, ErrorOffset, and ErrorMessage report malformed input while retaining byte-exact serialization, and GetStatistics fills THPDFContentASTStatistics with source size, counts, maximum depth, inline-image count, and malformed-node count

Edit overlays

ReplaceRange uses half-open byte offsets, while ReplaceLexeme, ReplaceNode, DeleteNode, InsertBeforeNode, and InsertAfterNode derive ranges from syntax records

Each THPDFContentEdit is copied into a sorted overlay; conflicting edits are rejected, adjacent ranges remain valid, EditCount reports accepted edits, and ClearEdits restores no-op output

Operator visitors

HPDFVisitContentStreamOperators layers reference-counted visitor and Delphi event callbacks over this AST, provides a read-only operand view, enforces source, operator, lexeme, depth, and replacement budgets, and reparses transformed output before success

Serialize

Serialize creates one output byte array after an overflow-checked size calculation, while WriteToStream writes untouched source spans and replacement spans directly to a caller-owned TStream

With no edits, either path reproduces the source byte for byte without changing numeric spelling, string escapes, comments, line endings, delimiters, or inline-image payloads

var
  AST: THPDFContentStreamAST;
  Output: TBytes;
begin
  AST:= HPDFParseContentStreamAST(ContentBytes);
  try
    if AST.Valid and AST.ReplaceRange(StartOfs, EndOfs, ReplacementBytes) then
      AST.Serialize(Output);
  finally
    AST.Free;
  end;
end;