This guide walks through reading a docx, traversing its content, mutating it, and writing it back — first with the docx-large bundle, then with docx-full.
Prerequisites: @awacloud/ooxml and @awacloud/fw installed and the install and runtime registration of Getting started; the bundles come from @awacloud/ooxml/bundles/docx-large and @awacloud/ooxml/bundles/docx-full.
Bootstrapping with ModuleRuntime
register() takes one descriptor per call — use registerAll(array)
for a batch. The extras are not re-exported by name from the
@awacloud/ooxml root; import the whole extras array instead:
// app.js
import fw from '@awacloud/fw';
import { fw_require, modules, extras } from '@awacloud/ooxml';
import { docxLargeBundle } from '@awacloud/ooxml/bundles/docx-large';
fw.runtime.registerAll(fw_require);
fw.runtime.registerAll(modules);
fw.runtime.registerAll(extras); // every opt-in extra; a bundle only
// resolves the ones it declares
fw.runtime.register(docxLargeBundle);
const word = fw.runtime.resolve('docxLargeBundle'); // enriched docx instance
Reading
const bytes = await fetch('/sample.docx')
.then(r => r.arrayBuffer())
.then(b => new Uint8Array(b));
const result = word.read(bytes);
result shape (selected fields):
{
document: { type: 'document', body: [paragraph | table], sectPr? },
package: /* raw OPC package */,
headers: { /* rId → header tree */ },
footers: { /* rId → footer tree */ },
images: { /* rId → { partName, data, contentType } */ },
styles: { styles: [Style], … },
settings: { /* settings.xml typed */ },
hyperlinks: { /* rId → { target, external } */ },
unmodelledParts: [ /* { partName, contentType } — never written back */ ]
}
See docx model for full shapes.
Walking the tree
function walkBody(body) {
for (const node of body) {
if (node.type === 'paragraph') {
for (const child of node.children) {
if (child.type === 'run') {
const text = child.children
.filter(c => c.type === 'text')
.map(c => c.value).join('');
console.log('run:', text, child.rPr);
}
}
} else if (node.type === 'table') {
for (const row of node.rows)
for (const cell of row.cells)
walkBody(cell.children); // recurse — cells contain paragraphs
}
}
}
walkBody(result.document.body);
Concrete input → parsed object
Input XML excerpt (inside word/document.xml):
<w:p>
<w:pPr><w:jc w:val="center"/></w:pPr>
<w:r>
<w:rPr><w:b/><w:caps/><w:lang w:val="en-US"/></w:rPr>
<w:t>Hello, world.</w:t>
</w:r>
</w:p>
After read() (with docx-large wired):
{
type: 'paragraph',
pPr: { align: 'center' },
children: [{
type: 'run',
rPr: { bold: true, caps: true, lang: { val: 'en-US' } },
children: [{ type: 'text', value: 'Hello, world.' }]
}]
}
The caps and lang fields are typed thanks to wmlRunFormatting (included in docx-large).
Mutating + writing back
// Add a new paragraph at the end.
result.document.body.push({
type: 'paragraph',
pPr: { align: 'right' },
children: [{
type: 'run',
rPr: { italic: true, color: '0070C0' },
children: [{ type: 'text', value: 'Appended.' }]
}]
});
// `write` takes the DOCUMENT (`result.document`), not the whole read()
// result — pass through the auxiliary parts you want re-emitted.
const out = word.write(result.document, {
styles: result.styles, settings: result.settings,
headers: result.headers, footers: result.footers,
hyperlinks: result.hyperlinks
});
// → Uint8Array — feed to a Blob to download.
const blob = new Blob([out], {
type: 'application/vnd.openxmlformats-officedocument.wordprocessingml.document'
});
window.URL.createObjectURL(blob);
Inside the parts the model reads, untyped elements are kept in _extras
and re-emitted; write() produces the parts its model carries, and a part
read() did not model (a theme, the font table, document properties, …)
is not written back — read() lists it in result.unmodelledParts.
Tables
result.document.body.push({
type: 'table',
rows: [{
type: 'row',
cells: [
{ type: 'cell', tcPr: { width: 4000, vAlign: 'center' },
children: [{ type: 'paragraph', children: [
{ type: 'run', children: [{ type: 'text', value: 'A1' }] }] }] },
{ type: 'cell',
children: [{ type: 'paragraph', children: [
{ type: 'run', children: [{ type: 'text', value: 'B1' }] }] }] }
]
}]
});
With docx-large, cell.tcPr exposes typed borders, vAlign, mar, gridSpan, vMerge — see wmlTableProperties.
With docx-full
import { docxFullBundle } from '@awacloud/ooxml/bundles/docx-full';
// `extras` (registered above) already covers docx-full's own extras
// (wmlVmlLegacy, dmlShapesAdvanced, transitional, legacyVml, wmlMisc,
// mathMisc, dmlMainMisc) — only the bundle descriptor itself is new:
fw.runtime.register(docxFullBundle);
const wordFull = fw.runtime.resolve('docxFullBundle');
// Now legacy VML pictures, custom-geometry shapes,
// and all long-tail elements roundtrip with typed fields.
const result = wordFull.read(legacyBytes);
Use docx-full when round-tripping fixtures from Word 2003 / converted documents that include <w:pict>, <o:OLEObject>, or <a:custGeom> shapes.