Chainable constructive DSL for assembling a PDF document from scratch.

Module pdfBuilder | Source packages/front/office/pdf/src/document/builder.js | Deps pdfErrors, pdfParserObj, pdfWriter | Worker-safe yes

An ergonomic layer over pdfWriter.writeDocument: instead of hand-assembling a typed indirect-object graph, builder() returns a chainable object (addPage → addContent/addFont/addImage → … → build()) that tracks pages, content streams, fonts (referenced or embedded), image XObjects and the Info dict internally, then emits Uint8Array bytes on .build(). The returned chain is closure-based (no class, no this), so it stays worker-transportable.

Resolve

const b = runtime.resolve('pdfBuilder');
// Returns: { builder }
const doc = b.builder();
// doc: { addPage, addContent, addFont, addImage, addMetadata, setVersion, setId, build }

API

Method Signature Returns
builder () => Builder Creates a fresh chainable builder (own closure state — call once per document).

Builder chain

Method Signature Notes
addPage (opts?: { mediaBox?: number[4], cropBox?: number[4], rotate?: number, resources?: DictObj }) => Builder Starts a new page (default mediaBox is US Letter [0,0,612,792]); becomes the target of subsequent addContent/addFont calls.
addContent (data: string | Uint8Array) => Builder Appends one content stream to the current page.
addFont (spec: { name: string, baseFont: string, subtype?: string, encoding?: 'WinAnsiEncoding' | 'MacRomanEncoding' | 'StandardEncoding' }) => Builder Legacy shape — registers a non-embedded font reference (subtype defaults to Type1) into the current page's /Resources /Font. The font dictionary is allocated once per distinct (baseFont, subtype, encoding) and shared by every page that registers it. Without encoding the emitted bytes are frozen.
addFont (spec: { name: string, embedded: SimpleEmbed | CidEmbed }) => Builder Embedded shape — takes a pdfFontEmbed result and allocates every indirect it needs (see Embedded fonts). Exactly one of baseFont / embedded must be given.
addImage (spec: { name, width, height, colorSpace, bitsPerComponent, data, filter?, decodeParms?, sMask? }) => Builder Allocates an image XObject as an indirect stream and registers it into the current page's /Resources /XObject (see Images). Decodes nothing; emits no content operator.
addMetadata (meta: object) => Builder Merges Info-dict fields. Title/Author/Subject/Keywords/Creator/Producer/CreationDate/ModDate are always emitted as PDF strings; other keys are kept only if their value is already a string.
setVersion (v: string) => Builder Overrides the PDF header version (default '2.0'); must match /^\d\.\d$/.
setId (a: string | Uint8Array, b?: string | Uint8Array) => Builder Sets the /ID pair — hex string or raw bytes; b defaults to a.
build () => Uint8Array Assembles Catalog + Pages tree + page objects + Info (if any), then calls pdfWriter.writeDocument.

Every chain method except build returns the same Builder instance.

addFont spec — non-embedded shape

Key Type Notes
name string Resource name registered under the page's /Resources /Font (e.g. F1). Required.
baseFont string /BaseFont name (e.g. Helvetica). Exactly one of baseFont / embedded.
subtype string /Subtype; defaults to Type1.
encoding 'WinAnsiEncoding' | 'MacRomanEncoding' | 'StandardEncoding' Optional. One of the three predefined simple-font encodings of ISO 32000-1 § 9.6.6 / Annex D, emitted as /Encoding /<name> after the existing keys: << /Type /Font /Subtype /Type1 /BaseFont /Helvetica /Encoding /WinAnsiEncoding >>. Absent (or undefined) emits no /Encoding key and the bytes are identical to a spec without the option — the default is unchanged. Any other value, or encoding combined with embedded, throws pdf/builder/bad-font with { name, encoding } in its context. No /Differences or custom encoding dictionaries.

Shared dictionary. A non-embedded font dictionary is allocated once per distinct (baseFont, subtype, encoding) and shared by every page that registers it: subtype is the defaulted value, so { baseFont: 'Helvetica' } and { baseFont: 'Helvetica', subtype: 'Type1' } share one object. Each page (and each resource name) still gets its own /Resources /Font entry; only the referenced object is shared. A 3-page document that registers the same four faces on every page therefore carries 4 font dictionaries, not 12. A document that never repeats a (baseFont, subtype, encoding) triple is byte-identical to one built without sharing.

Examples

Minimal one-page document

const { builder } = runtime.resolve('pdfBuilder');
const bytes = builder()
    .addPage({ mediaBox: [0, 0, 612, 792] })
    .build();

Page with a font and content stream

const bytes = builder()
    .addPage({ mediaBox: [0, 0, 612, 792] })
    .addFont({ name: 'F1', baseFont: 'Helvetica', subtype: 'Type1' })
    .addContent('BT /F1 12 Tf 100 700 Td (Hello) Tj ET')
    .build();

Metadata + explicit /ID

const bytes = builder()
    .addPage()
    .addMetadata({ Title: 'Demo', Author: 'awa' })
    .setId('0102030405060708090a0b0c0d0e0f10')
    .build();

Embedded fonts

addFont({ name, embedded }) accepts a pdfFontEmbed result — embedSimple (a SimpleEmbed: /TrueType + /WinAnsiEncoding) or embedCid (a CidEmbed: /Type0 + /Identity-H + /CIDFontType2). The adapter deliberately returns inline dicts and leaves every indirect to its consumer; the builder is that consumer.

An embedded value must carry fontFile (Uint8Array), descriptor, toUnicodeStream and either fontDict or type0Dict + cidFontDict — never both — else pdf/builder/bad-font.

What gets allocated

Per embed result object, not per addFont call:

# Indirect Content Route
1 font-program stream << /Length1 <fontFile.length> >> (plus /Subtype /OpenType when fontFileKey is FontFile3) over embedded.fontFile both
2 FontDescriptor embedded.descriptor cloned, with [fontFileKey] → 1 0 R added both
3 /ToUnicode embedded.toUnicodeStream (the serializer refuses an inline stream) both
4 CIDFont embedded.cidFontDict cloned, /FontDescriptor → 2 0 R CID only
5 font object simple: embedded.fontDict cloned, /FontDescriptor → 2 0 R, /ToUnicode → 3 0 R · CID: embedded.type0Dict cloned, /DescendantFonts [4 0 R], /ToUnicode → 3 0 R both

So a simple embedding costs 4 indirects and a composite one 5 — i.e. +3 and +4 over the single object a legacy addFont allocates. Only the last one lands in the page's /Resources /Font.

Identity cache. The builder keeps a WeakMap keyed by the embedded object, so registering the same result on several pages (under any resource names) allocates the graph once and every page references the same font object number. Two distinct results — even from the same face — are two graphs.

No mutation. Every patched dict is shallow-cloned; the caller's descriptor / fontDict / type0Dict / cidFontDict come back exactly as pdfFontEmbed returned them, so one result can be reused across builders.

Example — pdfFontEmbed → addFont → addContent

const fontsMod = runtime.resolve('fonts');
const embed    = runtime.resolve('pdfFontEmbed');
const { builder } = runtime.resolve('pdfBuilder');

const face = fontsMod.read(ttfBytes);
const text = 'Unicode fi ⁄ ⁴';
const cps  = [...new Set([...text].map(c => c.codePointAt(0)))];

const e = embed.embedCid(face, cps);           // or embedSimple for WinAnsi
const hex = [...e.encode(text)]
    .map(b => b.toString(16).padStart(2, '0')).join('');

const bytes = builder()
    .addPage({ mediaBox: [0, 0, 612, 792] })
    .addFont({ name: 'F1', embedded: e })
    .addContent(`BT /F1 12 Tf 72 700 Td <${hex}> Tj ET`)
    .build();

e.encode(text) yields the character codes the written font expects — WinAnsi bytes for the simple route, big-endian 2-byte CIDs for Identity-H — and e.widthOf(codePoint) gives the 1000/em advance for layout.

Images

addImage({ name, … }) is the image counterpart of addFont's embedded route: it allocates the image as an indirect stream object — the serializer refuses an inline one — and maps name to it in the current page's /Resources /XObject.

The seam decodes nothing. The caller supplies data that is already encoded plus the parameters that describe it. This module does not parse a PNG IHDR, a JPEG SOF, or anything else; it does not transcode, resample or colour-manage, and it never inspects data — the bytes reach the file verbatim. Choosing /Filter, /ColorSpace and /BitsPerComponent consistently with those bytes is the caller's job.

And it places no ink. addImage makes the resource reachable; drawing it is a content-stream matter, so the caller emits the operators itself (see the snippet below).

Spec

Key Type /Key Notes
name string — Resource name without the leading slash (e.g. 'Im0'); non-empty.
width number /Width Positive safe integer.
height number /Height Positive safe integer.
colorSpace string /ColorSpace Emitted as a name (e.g. 'DeviceRGB', 'DeviceGray'); non-empty.
bitsPerComponent number /BitsPerComponent Positive safe integer.
data Uint8Array stream body The encoded bytes, verbatim; /Length is added by the serializer.
filter string? /Filter Emitted as a name (e.g. 'DCTDecode', 'FlateDecode'). Omit for unfiltered data — then no /Filter is written.
decodeParms DictObj? /DecodeParms Passthrough — must already be a typed obj.dict, exactly like addPage's resources.
sMask object? /SMask A soft mask: the same spec shape minus name, and with no nested sMask.

The emitted dict is << /Type /XObject /Subtype /Image /Width … /Height … /ColorSpace … /BitsPerComponent … >>, plus /Filter, /DecodeParms and /SMask when those were given.

Soft masks

An sMask is allocated first, as its own image XObject, and the parent's /SMask holds a reference to it. It is referenced, never named: it does not appear in the page's /XObject dict, so no content operator can draw it directly. A masked image therefore costs 2 indirects instead of 1.

builder().addPage()
    .addImage({
        name: 'Im0', width: w, height: h,
        colorSpace: 'DeviceRGB', bitsPerComponent: 8,
        filter: 'FlateDecode', data: rgbDeflated,
        sMask: {                       // no `name` here
            width: w, height: h,
            colorSpace: 'DeviceGray', bitsPerComponent: 8,
            filter: 'FlateDecode', data: alphaDeflated
        }
    })
    .build();

Resource precedence

/XObject follows the pre-existing /Font rule exactly: addPage's resources passthrough is merged after the built resource classes and only fills keys the builder did not produce. So on a page that called addImage, a caller-supplied /XObject is dropped; on a page that did not, it passes through untouched. Nothing about that rule changed.

Example — place a JPEG on the page

const bytes = builder()
    .addPage({ mediaBox: [0, 0, 612, 792] })
    .addImage({
        name: 'Im0',
        width: 800, height: 600,          // the image's own pixel size
        colorSpace: 'DeviceRGB',
        bitsPerComponent: 8,
        filter: 'DCTDecode',              // the bytes ARE a JPEG already
        data: jpegBytes
    })
    // The seam placed the resource; the caller places the ink.
    // `cm` is width height 0 0 x y in USER SPACE units, not pixels:
    // 400x300 pt with its lower-left corner at (100, 400).
    .addContent('q 400 0 0 300 100 400 cm /Im0 Do Q')
    .build();

q … Q brackets the transform so the CTM is restored afterwards; the image XObject's own space is the unit square, which is why the cm matrix carries the on-page size directly.

Errors

Code Class When
pdf/builder/no-page RenderError addContent, addFont or addImage called before any addPage.
pdf/builder/bad-bytes RenderError addContent data is neither string nor Uint8Array.
pdf/builder/bad-font RenderError addFont spec has no name; or neither / both of baseFont and embedded; or an embedded value that is not a well-formed pdfFontEmbed result; or an encoding that is not one of the three predefined names, or is combined with embedded.
pdf/builder/bad-image RenderError addImage spec is not an object, or one of name / width / height / colorSpace / bitsPerComponent / data / filter / decodeParms is missing or ill-typed, or an sMask carries a nested sMask. err.context.keys lists the offending keys; err.context.sMask is true when the failure is on the mask leg.
pdf/builder/bad-metadata RenderError addMetadata argument is not an object.
pdf/builder/bad-version RenderError setVersion value doesn't match /^\d\.\d$/.
pdf/builder/bad-id RenderError setId part is neither hex string nor Uint8Array.
pdf/builder/bad-id-hex RenderError setId hex string has odd length.
pdf/builder/bad-box RenderError mediaBox/cropBox is not a 4-element array.
pdf/builder/no-pages RenderError build() called with zero pages added.

It also propagates every code from pdfWriter (raised inside writeDocument).

See also