Token-Preserving Content-Stream AST
HPDFContentStream provides a bounded lossless syntax model for targeted edits to decoded PDF content streams
Parse and inspect
HPDFParseContentStreamAST creates a THPDFContentStreamAST with default limits of 4,194,304 lexemes and 256 nested containers; callers that need tighter limits can call THPDFContentStreamAST.Create directly
THPDFContentLexemeKind distinguishes whitespace, comments, numbers, names, literal and hexadecimal strings, booleans, nulls, words, container delimiters, inline-image data, and malformed bytes; every input byte belongs to exactly one THPDFContentLexeme
THPDFContentSyntaxNodeKind distinguishes the root, operations, operands, arrays, dictionaries, procedures, inline images, and malformed regions; each THPDFContentSyntaxNode uses parent, first-child, last-child, and next-sibling indexes so traversal does not allocate child arrays
Use GetLexeme, GetLexemeBytes, GetLexemeText, and GetNode with LexemeCount and NodeCount; GetOriginalBytes returns an independent source copy
Valid, ErrorOffset, and ErrorMessage report malformed input while retaining byte-exact serialization, and GetStatistics fills THPDFContentASTStatistics with source size, counts, maximum depth, inline-image count, and malformed-node count
Edit overlays
ReplaceRange uses half-open byte offsets, while ReplaceLexeme, ReplaceNode, DeleteNode, InsertBeforeNode, and InsertAfterNode derive ranges from syntax records
Each THPDFContentEdit is copied into a sorted overlay; conflicting edits are rejected, adjacent ranges remain valid, EditCount reports accepted edits, and ClearEdits restores no-op output
Operator visitors
HPDFVisitContentStreamOperators layers reference-counted visitor and Delphi event callbacks over this AST, provides a read-only operand view, enforces source, operator, lexeme, depth, and replacement budgets, and reparses transformed output before success
Serialize
Serialize creates one output byte array after an overflow-checked size calculation, while WriteToStream writes untouched source spans and replacement spans directly to a caller-owned TStream
With no edits, either path reproduces the source byte for byte without changing numeric spelling, string escapes, comments, line endings, delimiters, or inline-image payloads
var
AST: THPDFContentStreamAST;
Output: TBytes;
begin
AST:= HPDFParseContentStreamAST(ContentBytes);
try
if AST.Valid and AST.ReplaceRange(StartOfs, EndOfs, ReplacementBytes) then
AST.Serialize(Output);
finally
AST.Free;
end;
end;