Binary lexer, ISO 32000-2 §7.2 — bytes → token stream.
Module pdfTokenizer | Source packages/front/office/pdf/src/syntax/tokenizer.js | Deps pdfErrors, pdfShared | Worker-safe yes
Emits { kind, value?, offset, end } tokens for the PDF lexical conventions.
The tokenizer is stream-aware at the keyword level: when it emits
{ kind: 'kw', value: 'stream' }, the consumer skips the binary payload itself
(using the preceding dictionary's /Length) and then resumes at endstream.
Resolve
const tokMod = runtime.resolve('pdfTokenizer');
// Returns: { tokenize, lastIndexOfBytes }
API
| Method |
Signature |
Returns |
tokenize |
(bytes: Uint8Array, opts?: { keepWhitespace?, start?, end? }) => Tokenizer |
Iterator (next, peek, pos, seek, bytes). |
lastIndexOfBytes |
(bytes: Uint8Array, needle: Uint8Array, from?: number) => number |
Offset (or -1) of the last match. |
Token kinds
kind |
Payload |
ws |
whitespace run (only when keepWhitespace) |
comment |
comment body (bytes between % and EOL) |
name |
value: string (#xx escapes decoded) |
int / real |
value: number, raw: string |
string |
value: Uint8Array — literal (...) form only |
hex |
value: Uint8Array — hex <...> form |
open_arr / close_arr |
[ / ] |
open_dict / close_dict |
<< / >> |
kw |
value: string (obj, endobj, R, stream, endstream, xref, trailer, startxref, true, false, null, n, f) |
eof_marker |
%%EOF |
Tokenizer interface
| Member |
Description |
next() |
Consumes and returns the next token (or null at EOF). |
peek() |
Looks ahead without consuming. |
pos() |
Current offset. |
seek(n) |
Forces the offset (used after skipping a stream body). |
bytes |
Reference to the source Uint8Array. |
Examples
Iterating a stream
const tok = tokMod.tokenize(bytes);
let t;
while ((t = tok.next())) {
if (t.kind === 'kw' && t.value === 'xref') break;
}
Locating startxref by backward scan
const needle = new Uint8Array([0x73,0x74,0x61,0x72,0x74,0x78,0x72,0x65,0x66]);
const at = tokMod.lastIndexOfBytes(bytes, needle);
Starting from a known offset
const tok = tokMod.tokenize(bytes, { start: 12345 });
tok.next(); // first token at or after 12345
Errors
| Code |
Class |
When |
pdf/tokenizer/bad-input |
ParseError |
bytes is not a Uint8Array. |
pdf/tokenizer/bad-seek |
ParseError |
seek(n) out of range. |
pdf/tokenizer/bad-name-escape |
ParseError |
#xx truncated or non-hex. |
pdf/tokenizer/bad-number |
ParseError |
Empty or unparsable numeric token. |
pdf/tokenizer/bad-string |
ParseError |
Trailing backslash. |
pdf/tokenizer/unterminated-string |
ParseError |
(...) never closed. |
pdf/tokenizer/bad-hex |
ParseError |
Non-hex character inside <...>. |
pdf/tokenizer/unexpected-rangle |
ParseError |
Lone > outside a hex string. |
pdf/tokenizer/empty-keyword |
ParseError |
Empty keyword. |
See also