GetTextBlockCharUnicodeFlags
Text, Extraction, Diagnostics, Unicode
Description
Returns Unicode mapping and fallback flags for one character in a selected-document text block
Syntax
Delphi
Function TPDFlib.GetTextBlockCharUnicodeFlags(TextBlockListID, Index, CharIndex: Integer): Integer;Parameters
| TextBlockListID | Handle returned by ExtractPageTextBlocks |
|---|---|
| Index | 1-based text block index |
| CharIndex | 1-based UTF-16 code-unit position inside the block text |
Return values
| -1 | The handle, block index or character index is invalid |
|---|---|
| 0 or flags | No source glyph mapping applies or the result contains one or more Unicode diagnostic flags |
Flags
PDF_TEXT_CHAR_UNICODE_EXPLICIT | The character was produced by an explicit valid ToUnicode mapping |
|---|---|
PDF_TEXT_CHAR_UNICODE_FALLBACK | The character was produced by a font encoding, glyph, CID or raw-code fallback |
PDF_TEXT_CHAR_UNICODE_MISSING_TOUNICODE | The font has no ToUnicode entry |
PDF_TEXT_CHAR_UNICODE_UNMAPPED | A valid ToUnicode CMap has no mapping for the source code |
PDF_TEXT_CHAR_UNICODE_INVALID | The ToUnicode CMap is unusable or the selected mapping is invalid UTF-16 |
PDF_TEXT_CHAR_UNICODE_REPLACEMENT | The returned text contains U+FFFD because no usable fallback was available |
Remarks
Use bitwise tests because a fallback result also carries the missing, unmapped or invalid reason
Flags are retained during extraction and queried in constant time without reparsing page content
A supplementary Unicode scalar occupies two UTF-16 positions and both positions carry the same flags