Audience: on-prem/air-gap RAG integrators, sovereign-cloud teams, defence and regulated-sector engineers who ingest documents into a chunked, auditable index inside their own perimeter. This page is the wire contract: every key, why it exists for that audience, and what it does and does not promise. It is a reference for anything consuming
@awacloud/oconv's output, not just for@awacloud/oconvitself.Prerequisites. Producing a profile-v1 document takes
oconv.toMdfrom the@awacloud/oconvmain entry (registered on an@awacloud/fwModuleRuntime— see the package README's Quick Start), in a browser or on the runtimes the package'senginesfield names (Bun ≥ 1.0, Node ≥ 20);oconv.fromMdaccepts one back as input. Consuming one needs no oconv code at all: any YAML reader plus any CommonMark/GFM parser reads it.
A profile-v1 document is a literal YAML front-matter block followed by a CommonMark/GFM body:
---
profile: v1
ir: oconv-ir/v1
sourceFormat: docx
sourceName: report.docx
sourceBytes: 1556
sourceSha256: 7be541047cdb2255188fdf68ed6ddc98c47ed22332093df494dcb3f5c7f60b7c
convertedAt: 2026-07-20T00:00:00Z
converter: oconv
converterVersion: 1.0.0
engine: bun
blocks: 6
anchors:
- { level: 1, anchor: sovereign-rag-ingestion }
- { level: 2, anchor: why-air-gap-matters }
- { level: 3, anchor: chunking }
lossy: true
losses:
- { code: block/dropped, detail: footnote }
assets:
- { kind: image, name: logo.png }
---
Every line through the anchors: list is what toMd emits for the package's own
docx-smoke.docx corpus fixture, called with that convertedAt and
engine: 'bun'. The losses: and assets: entries are added to show every
key: that fixture converts with an empty ledger and no image, so its real
front matter ends at lossy: false.
Every front-matter key below is emitted exactly in this order by
src/write/ir-to-md.js (oconvIrToMd); anchors:, losses: and
assets: are each omitted entirely when empty (never an empty []).
Key-by-key semantics
| Key | Meaning | Why this audience needs it |
|---|---|---|
profile |
Wire-profile version (v1), independent of ir |
The RAG ingestion pipeline can pin its parser to a profile version even after the pivot evolves. |
ir |
The pivot IR version that produced this document (oconv-ir/v1) |
Traceability to the exact internal model — useful when auditing a conversion discrepancy. |
sourceFormat |
docx | odt | xlsx | ods | pptx | odp | pdf |
Which reader ran. |
sourceName |
Original file name | Human/audit trail back to the source document. |
sourceBytes |
Source byte length | A cheap sanity check alongside the hash. |
sourceSha256 |
Lowercase hex SHA-256 of the source bytes | Auditable: pins the exact input a reviewer can re-fetch and re-hash. Computed by the facade via crypto.subtle.digest — no external hashing tool, no network. |
convertedAt |
ISO-8601 timestamp | Always caller-injected, never Date.now() — see Determinism below. |
converter / converterVersion |
Producer identity | Pins which version of @awacloud/oconv produced the file, for reproducing a conversion later. |
engine |
Host engine label (bun, browser, …) |
Omitted entirely when the caller doesn't supply one — never invented. |
blocks |
Top-level IR block count | A structural checksum independent of exact text — catches gross truncation. |
anchors |
{level, anchor}[], document order |
Chunkable: a chunker splits on headings without re-parsing the markdown; citations can address #anchor stably. |
lossy |
losses.length > 0 |
A single boolean gate a pipeline can branch on ("flag lossy documents for human review"). |
losses |
{code, detail}[], reader losses then writer losses, document order |
Honest: losses the pipeline detects are recorded; the loss matrix lists the known silent cases — see Loss ledger below. Present only when lossy: true. |
assets |
{kind, name, bytes?}[], deduplicated by name, document order |
Image references extracted from the source; bytes travels by reference when the reader captured the image bytes: the same Uint8Array, never copied and never written into the front matter. The writer (ir-to-md.js) reads it from escapes.docx.bytes, the field the docx reader fills — the only reader that captures image bytes — so bytes is present for a docx image whose bytes resolved and absent for every other source. Measured on poi-with-gif.docx: one asset (Grafik 1) whose bytes is a 6554-byte Uint8Array, lossy: false (tests/fidelity.integration.test.js). Present only when the document has images. |
Determinism and auditability
Two properties matter for a sovereign/air-gap offer, and both are structural (not best-effort):
sourceSha256pins the input. A reviewer with the original file can recompute the hash and confirm the front matter describes THAT exact byte sequence.convertedAtis always caller-injected — never defaulted toDate.now().src/oconv.js'stoMdthrowsoconv: convertedAt is requiredwhen omitted. Consequence, verified by the fidelity harness: converting the same bytes twice with the sameconvertedAtproduces byte-identical markdown, and changing onlyconvertedAtchanges exactly that one front-matter line — nothing else in the document shifts. This is the property an audit trail needs: a re-run is a diff, not a fresh unknown.
This determinism claim is scoped to the to-md direction, with two
from-md exceptions: the .docx and .pdf targets. The .docx target of
fromMd is byte-reproducible: @awacloud/ooxml stamps every zip entry with a
fixed 1980-01-01 00:00 timestamp, so two identical fromMd(..., 'docx') calls
on the same markdown produce byte-identical output —
tests/roundtrip.integration.test.js § "reproducibility — md -> docx is
byte-reproducible". The .odt target is not: the ODF package writer stamps
the current time, so two identical calls give equal document models but may
give different bytes; see docs/loss-matrix.md
"From-md reproducibility" for the full statement. The .pdf target IS
byte-reproducible: irToPdf never stamps a date or an /ID, so two
identical fromMd(..., 'pdf') calls on the same markdown produce
byte-identical output — tests/roundtrip.integration.test.js §
"reproducibility — md -> pdf is byte-reproducible".
Chunking usage
anchors: is the section index a RAG ingestion pipeline chunks against:
split the body on #/##/… headings, and address each chunk by its
#anchor slug. Slugs are deterministic and ASCII-only by construction
(NFKD → strip combining marks → lowercase → [^a-z0-9]+ → -, with a
-1, -2 … collision suffix for duplicate headings) — safe for any
downstream index, URL, or filename. anchors.length equals the number of
#-heading lines the writer emits for real headings (headings inside a
GFM table cell are excluded on both sides — a cell renders inline-only, so
such a heading never becomes a #-line and is correctly un-indexed), but
this is not an absolute count of #-prefixed lines in the body: a
fenced code block can carry #-lines of its own. A source paragraph line
that itself begins with a literal # (e.g. # Ligne FACT) is escaped by
@awacloud/md's Markdown renderer (\# Ligne FACT; likewise a line-leading 1., 1), -, +, * or >),
so it re-parses as paragraph text, never as a heading. Measured on
facturx-minimum-sample.pdf: its body used to carry 5 #-prefixed lines
against 0 headings; it now carries 0 against 0, and
tests/fidelity.integration.test.js includes that fixture in the
heading-count/anchor-count equality assertion again.
Profile v1 as INPUT (fromMd)
fromMd accepts a profile-v1 document — or plain
CommonMark/GFM with no front matter at all — and writes a .docx, .odt
or .pdf container. The relationship with this profile is read-only and
one-directional: a leading front-matter fence (---/+++/;;;,
detected the same way @awacloud/md's mdFrontmatter.stripFrontmatter
detects it on the toMd side) is stripped before parsing
(src/read/md-to-ir.js, composing mdFrontmatter.stripFrontmatter from
@awacloud/md/extra/frontmatter.js — never an @awacloud/md internal), and:
- it is never parsed — no YAML/TOML/JSON decoding of its content, by
design (same "emit, never parse" stance the reader takes for the
stripped block on the
toMdwrite side, generalized to the read direction too); - it is recorded as a loss, not a silent strip — one
{code: 'frontmatter/stripped', detail: <lang>}entry per document (langis'yaml'|'toml'|'json', matching the fence syntax), verified bysrc/read/md-to-ir.test.jsfor all three fence kinds and bytests/roundtrip.integration.test.jsfor a realtoMd-produced profile document fed back in (sourceSha256/converterVersionand all) — a deliberate design decision, not a gap; - it is never written into the target container — no
.docx/.odtcore-properties or custom-metadata part is populated from it. The target container's own fixed provenance is unrelated to it: none fordocx,meta:generator: '@awacloud/odf'forodt,/Producer+/Creator'@awacloud/oconv'forpdf(see convert.md).tests/roundtrip.integration. test.js's profile-v1 input leg confirms the front-matter text is absent from the produced docx body viadocxApi.toText(docxApi.read(bytes) .document).
In short: a profile-v1 document's front matter describes a prior
conversion (the toMd call that produced it) — it is provenance about
the PAST, not an instruction for the container fromMd is about to build,
and it does not round-trip. A document produced by fromMd carries no
sourceSha256/convertedAt/losses front matter of its own; if a caller
wants an audit trail for the from-md direction, it is their own
responsibility to keep one, @awacloud/oconv emits none.
Loss ledger — codes in use
Losses are data, appended to losses: in document order (reader losses
lead, writer losses follow) — never a side-channel log line. A detected
loss is always recorded; a construct the pipeline cannot detect is not (the
loss matrix lists the known silent cases). detail is a
string on every reader code and on the fromMd reader and docx/odt writer
codes below; the .pdf typesetter's
layout/*, text/unencodable and inline/* codes carry an object detail
instead, and the per-block records among them also name the block's index
and kind beside code. The fromMd and convert facades return those
records unchanged, and so does the worker reply. Codes emitted by the seven
shipped readers and the md writer:
| Code | Meaning |
|---|---|
block/dropped |
An IR node (or a source element with no IR equivalent) could not be represented; detail names the source element/kind. |
link/target-missing |
A hyperlink run had no resolvable target; the visible text is kept, the link is not. |
list/numbering-unresolved |
The list's ordered-vs-bullet status could not be resolved (docx: no matching numbering.xml entry; odt: the list style falls outside the resolution scope below) — falls back to bullet. |
list/nesting-flattened |
(docx only) A nested list level (ilvl > 0) was flattened into its enclosing list — recorded once per document. |
heading/level-clamped |
(odt only) An outlineLevel above 6 was clamped to 6. |
heading/subtitle-degraded |
(docx only) A paragraph whose style is the built-in Subtitle (resolved by its w:name through styles.xml) was kept as a plain paragraph — no heading level is guessed for it; detail is the style ID, one record per subtitle paragraph. Pinned by src/read/docx-to-ir.test.js, including the committed French-styled fixture docx-fr-styles.docx. |
image/unresolved |
(odt only) An inline image reference could not be resolved through this reader's frozen dependency list. |
image/bytes-unavailable |
(docx only) A drawing's r:embed relationship never resolved to image bytes; detail names the drawing. Pinned by src/read/docx-to-ir.test.js. |
inline/dropped |
(odt only, this reader's own sense) An unmapped inline XML element inside a run; detail names the element. Same code name as the fromMd reader's inline/dropped below — a different domain, the same "unmapped node, kept as loss" shape. |
inline/flattened |
(odt only) An element inside a <text:span> (a field, a reference, a note) holds text: the text is kept as a plain run carrying the span's emphasis, the element's meaning is not; detail names the element. A frame holding text inside a span records image/unresolved instead, and an element holding no text records inline/dropped (or image/unresolved for a frame). |
run/format-unresolved |
(odt only) A <text:span> run's bold/italic/strike/monospace styling could not be resolved — the style falls outside the resolution scope below. |
Read-side resolution scope for the two odt *-unresolved codes: semantic
resolution covers content.xml automatic styles AND styles.xml
office:styles styles that are fully mapped (a single style, no
style:parent-style-name chain); monospace resolves only through a
declared fixed-pitch or generic-modern font face. Foreign documents styled
through parent chains or partially-mapped styles still yield
run/format-unresolved / list/numbering-unresolved on read — that is
why these two codes survive the tier-2 raise rather than being retired with
the four write-side codes below.
Remaining reader codes:
| Code | Meaning |
|---|---|
sheet/formula-as-value |
(xlsx/ods) A formula cell was reduced to its cached/computed value only. |
sheet/format-dropped |
(xlsx/ods) Recorded once per sheet, first styled cell seen. |
sheet/merge-dropped |
(xlsx/ods) Recorded once per sheet, first merged/spanned range seen. |
sheet/chart-dropped |
(xlsx/ods) Recorded once per sheet, first embedded chart/drawing seen. |
slides/untitled |
(pptx/odp) The slide carries no title placeholder/frame, or its only one has no non-empty text. |
slides/media-dropped |
(pptx/odp) A non-text shape/frame (picture, table, chart, image, object) was dropped at tier 1; detail names its kind. |
slides/notes-omitted |
(odp; documented no-op for pptx) The slide has speaker notes but includeNotes was false (the default). |
text/undecodable |
(pdf) ≥1 character code in a text run resolved to no Unicode — counted, never silently kept. |
text/font-unresolved |
(pdf) A show operator ran with no resolvable current font. |
text/width-approximated |
(pdf) A font's glyph widths could not be read from the font, so a declared fallback width (500/1000 em) placed its text — once per font resource per page, detail = the resource name. The text is kept; only the inferred word spaces around it are approximate. |
image/dropped |
(pdf, to-md) An image XObject draw or inline image, including one drawn inside a Form XObject — tier 1 keeps no images. Also emitted by the fromMd docx and odt writers (detail = the IR image's name), only when no bytes are reachable for the image (neither opts.assets[name] nor the docx reader's escapes.docx.bytes) — when bytes are reachable the writer embeds the image and records image/size-defaulted instead (odt: in a default 5.08cm × 3.81cm (2 in × 1.5 in) frame). |
xobject/form-dropped |
(pdf) A Form XObject's stream could not be used (not a stream, undecodable, unparsable), so its text is not extracted; detail = <name>: <reason>. A usable form is executed in place and records nothing. |
xobject/form-cycle |
(pdf) A Form XObject drawn from inside itself, directly or through other forms, was not re-entered; detail = the resource name. |
xobject/form-depth |
(pdf) A Form XObject that would nest deeper than 12 forms was not executed; detail = the resource name. |
xobject/form-budget |
(pdf) A page ran a budget of formOpBudget operators inside Form XObjects (default 1,000,000, configurable through the toMd formOpBudget input); its later form draws were not executed. Recorded once per page; detail = the first skipped resource name. |
content/undecodable |
(pdf) A content stream failed to resolve, decode or parse and was skipped; detail is stream <objNum>: <cause>, the cause being the thrown error's message collapsed to one line and capped at 160 characters. |
struct/dropped |
(pdf, tagged fast path) A Table/L struct container was flattened; its descendant text is kept as paragraphs. |
xobject/form-dropped, xobject/form-cycle, xobject/form-depth,
xobject/form-budget, content/undecodable, text/font-unresolved and
text/width-approximated are each pinned by a dedicated producing test,
src/read/pdf/text-extract.test.js.
Codes emitted by the fromMd reader (src/read/md-to-ir.js) and the
three fromMd writers (src/write/ir-to-docx.js,
src/write/ir-to-odt.js, and src/write/ir-to-pdf.js + its
src/write/pdf/* typesetting-layer modules):
| Code | Meaning |
|---|---|
frontmatter/stripped |
A leading ---/+++/;;; front-matter fence was stripped before parsing — never parsed, never carried into the target container; detail names the fence language (yaml/toml/json). See "Profile v1 as INPUT" above. |
list/start-dropped |
An ordered list started above 1; the IR list node has no start prop to carry it. |
list/task-marker-dropped |
A GFM task-list checkbox ([x]/[ ]) was stripped from a list item. |
table/align-dropped |
A GFM table declared column alignment; the IR row/cell model carries none. |
inline/linebreak-degraded |
A hard line break (trailing double space or backslash) degraded to a single space. |
inline/dropped |
An unmapped inline node kind (detail = its type, e.g. html_inline), or a non-empty link title (detail: 'link-title' — the link text/target are kept). Same code name as the odt reader's own inline/dropped above — a different domain, the same "unmapped node, kept as loss" shape. |
list/depth-clamped |
(docx target) A list nested deeper than ilvl 8 was clamped to 8. |
block/degraded |
(docx and odt targets) A code block (detail: 'codeBlock'), a blockquote (detail: 'blockquote'), or — docx only — a non-paragraph/non-list block inside a list item (detail: 'listItem-child:<kind>') was written in a structurally simpler form. |
image/size-defaulted |
(docx and odt targets) Image bytes WERE reachable and the image is PLACED at a default box — @awacloud/ooxml's 2 in × 4:3 box for docx, a 5.08cm × 3.81cm (2 in × 1.5 in) frame for odt. Neither writer applies the image's intrinsic size; detail names the image. A degrade, never a drop. |
The .pdf target's bounded typesetter (src/write/ir-to-pdf.js +
src/write/pdf/*) emits its own vocabulary — 10 layout/* codes,
text/unencodable and 2 inline/* codes — the full options/font-route
context for each lives in docs/pdf-writer.md; this
table states only the vocabulary:
| Code | Meaning |
|---|---|
layout/unhandled-block |
An IR block kind the typesetter has no delegate for was skipped entirely. |
layout/font-fallback |
A style class had no explicit AND no default-face bytes on a mixed/embedded font call — fell back to Standard 14 for that class. |
layout/s14-variant-metrics-approx |
DEFENSIVE FALLBACK — fires only if the fonts package regresses to a shared width table across the four Standard 14 variants; not reachable against the real @awacloud/fonts, whose Standard 14 tables give each variant its own widths (only a stubbed measurer reaches it, src/write/ir-to-pdf.test.js). |
text/unencodable |
Route-conditional: a non-WinAnsi code point degrades to ? on the Standard 14 route; a code point missing from the resolved face's own cmap draws .notdef on the explicit/default-face route — one collapsed record per document. |
layout/line-overflow |
An unbreakable token (or an overlong code line) wider than the column overflowed it. |
layout/image-dropped |
An image was not placed — a non-JPEG encoding or a CMYK/YCCK JPEG (reason: 'unsupported-encoding'), or no bytes at all (reason: 'no-bytes'); the [image: <alt>] placeholder still draws. |
layout/list-empty |
A list node had no listItem child. |
layout/table-empty |
A table node had zero rows, or zero rows left after the header row. |
layout/table-scaled |
A table (or a cell) wider than the column was scaled proportionally to fit. |
layout/table-clipped |
A table (or cell) stayed too narrow even at minimum width. |
layout/block-clipped |
A block taller than one page was clipped at the page bottom — a hard page break; the typesetter never splits a block across pages. |
inline/strike-dropped |
A text block carried at least one struck run; the typesetter draws no strikethrough rule, so the text is drawn plain. One record per block, detail = { runs, text } (the number of struck runs and the first one's text, up to 40 characters). Recorded, not rendered. |
inline/code-emphasis-dropped |
A text block carried at least one monospace run that is also bold or italic; the monospace class wins and the emphasis is not drawn. One record per block, detail = { runs, text }. Recorded, not rendered. |
Retired codes (odt target)
Four (odt target only) codes shipped in earlier ledgers and are
retired: @awacloud/odf gained a typed semantic write model, so
src/write/ir-to-odt.js now writes every
construct they used to record as lost, and md→odt reaches tier 2. They
are listed here — not silently deleted — because documents converted before
the retirement carry them in their losses: front matter, and a consumer reading
an archived profile-v1 document must still be able to look them up.
| Retired code | What it meant | Now |
|---|---|---|
run/format-unwritable |
A run's bold/italic/strike/code flags could not be written — no run-emphasis vocabulary in the typed model; detail was the flags as a csv. |
Written as a typed span (code → monospace), no loss. |
link/target-unwritable |
A run's link target could not be written — no text:a node in the public render dispatch; the link text was kept, detail was the dropped target. |
Written as a typed link run → text:a, no loss. |
list/ordered-unwritable |
An ordered list was written as a generic, unstyled list — no ordered/bulleted flag in the typed model. | Written with ordered:true + numFormat:'1', no loss. |
block/table-degraded |
A table was written as tab-joined row paragraphs — no text-body table render case; detail was <rows>x<cols>. |
Written as a typed table/row/cell with headerRows, no loss. |
Proof of retirement: src/write/ir-to-odt.test.js § "exhaustive absence of
the four retired loss codes", and tests/roundtrip.integration.test.js
leg 4 (md-structural.md → odt with an EMPTY ledger).
Retired code (docx target)
run/code-degraded shipped in earlier ledgers and is retired:
@awacloud/ooxml round-trips rPr.font through <w:rFonts>
(ooxml/src/docx/properties.js), so src/write/ir-to-docx.js now writes
font: 'Courier New' for a code run and src/read/docx-to-ir.js reads it
back via a frozen monospace name allowlist — md → docx → md keeps inline
code as backticks. Listed here — not silently deleted — because documents
converted before the retirement carry the code in their losses: front matter.
| Retired code | What it meant | Now |
|---|---|---|
run/code-degraded |
(docx target) An inline-code run was written as a plain run — formatting dropped, text kept; detail was the first 40 chars of the run text. |
Written with rPr.font: 'Courier New', no loss. |
Proof of retirement: src/write/ir-to-docx.test.js § "exhaustive absence
of run/code-degraded", and tests/roundtrip.integration.test.js leg 1
(md-structural.md → docx → md, empty fromMd ledger, backticks
recovered).
The full preserved/degraded/dropped picture per pair is
docs/loss-matrix.md — this table is the vocabulary,
that page is the coverage claim.
Zero network, zero external tooling
Nothing in the reader → pivot IR → writer path performs I/O, touches the
network, or shells out to an external tool. The Web APIs used beyond
plain ES are crypto.subtle.digest (source hashing, toMd only) and
performance.now() (call-duration timing — src/oconv.js, src/worker.js)
— available in a browser tab, in Bun, and inside a Worker alike. This is a hard property
for an air-gapped deployment, not a soft goal: bun test packages/front/office/oconv/ runs the whole suite, including the fidelity
harness, with no network access.
Versioning
profile and ir are independent keys on purpose: the wire profile can
stay v1 across pivot IR revisions (oconv-ir/v1 → v2 → …) as long as
the front-matter/body contract itself does not change; a profile-breaking
change gets its own profile: v2.
See also
docs/loss-matrix.md— per-pair fidelity coverage.../README.md— package overview, Quick Start, worker usage.src/write/ir-to-md.js— the writer this page documents (its own file-header JSDoc is the profile's original, code-adjacent spec).