<!-- llms-explorer concept facts · https://llms-explorer.com/tree/word-document-manipulation/ · pack 2026-09-08 · ~4405 tokens -->

# Word Document Manipulation

> A .docx file is a ZIP archive containing XML files.

Parent: [Document & File Formats](https://llms-explorer.com/tree/document-file-formats/) · 24 facets · 83 facts · page: https://llms-explorer.com/tree/word-document-manipulation/

## Overview

- A .docx file is a ZIP archive containing XML files. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#overview)

## Converting .doc to .docx

- Legacy .doc files must be converted before editing: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#converting-doc-to-docx)

## Accepting Tracked Changes

- To produce a clean document with all tracked changes accepted (requires LibreOffice): — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#accepting-tracked-changes)

## Creating New Documents

- Generate .docx files with JavaScript, then validate. Install: npm install -g docx — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#creating-new-documents)

## Validation

- After creating the file, validate it. If validation fails, unpack, fix the XML, and repack. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#validation)

## Page Size

- Common page sizes (DXA units, 1440 DXA = 1 inch): — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#page-size)
- Landscape orientation: docx-js swaps width/height internally, so pass portrait dimensions and let it handle the swap: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#page-size)

## Styles (Override Built-in Headings)

- Use Arial as the default font (universally supported). Keep titles black for readability. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#styles-override-built-in-headings)

## Tables

- CRITICAL: Tables need dual widths - set both columnWidths on the table AND width on each cell. Without both, tables render incorrectly on some platforms. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
- Table width calculation: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
- Always use WidthType.DXA - WidthType.PERCENTAGE breaks in Google Docs. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
  - Always use WidthType.DXA - never WidthType.PERCENTAGE (incompatible with Google Docs) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
  - Table width must equal the sum of columnWidths — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
  - Cell width must match corresponding columnWidth — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
  - Cell margins are internal padding - they reduce content area, not add to cell width — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)
  - For full-width tables: use content width (page width minus left and right margins) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tables)

## Multi-Column Layouts

- Force a column break with a new section using type: SectionType.NEXT_COLUMN. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#multi-column-layouts)

## Critical Rules for docx-js

- Set page size explicitly - docx-js defaults to A4; use US Letter (12240 x 15840 DXA) for US documents — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Landscape: pass portrait dimensions - docx-js swaps width/height internally; pass short edge as width, long edge as height, and set orientation: PageOrientation.LANDSCAPE — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Never use \n - use separate Paragraph elements — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Never use unicode bullets - use LevelFormat.BULLET with numbering config — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- PageBreak must be in Paragraph - standalone creates invalid XML — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- ImageRun requires type - always specify png/jpg/etc — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Always set table width with DXA - never use WidthType.PERCENTAGE (breaks in Google Docs) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Tables need dual widths - columnWidths array AND cell width, both must match — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Table width = sum of columnWidths - for DXA, ensure they add up exactly — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Always add cell margins - use margins: { top: 80, bottom: 80, left: 120, right: 120 } for readable padding — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Use ShadingType.CLEAR - never SOLID for table shading — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Never use tables as dividers/rules - cells have minimum height and render as empty boxes (including in headers/footers); use border: { bottom: { style: BorderStyle.SINGLE, size: 6, color: "2E75B6", space: 1 } } on a Paragraph instead. For two-column footers, use tab stops (see Tab Stops section), not tables — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- TOC requires HeadingLevel only - no custom styles on heading paragraphs — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Override built-in styles - use exact IDs: "Heading1", "Heading2", etc. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)
- Include outlineLevel - required for TOC (0 for H1, 1 for H2, etc.) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#critical-rules-for-docx-js)

## Editing Existing Documents

- Follow all 3 steps in order. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#editing-existing-documents)

## Step 1: Unpack

- Extracts XML, pretty-prints, merges adjacent runs, and converts smart quotes to XML entities (&#x201C; etc.) so they survive editing. Use --merge-runs false to skip run merging. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-1-unpack)

## Step 2: Edit XML

- Edit files in unpacked/word/. See XML Reference below for patterns. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-2-edit-xml)
- Use "Claude" as the author for tracked changes and comments, unless the user explicitly requests use of a different name. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-2-edit-xml)
- Use the Edit tool directly for string replacement. Do not write Python scripts. Scripts introduce unnecessary complexity. The Edit tool shows exactly what is being replaced. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-2-edit-xml)
- CRITICAL: Use smart quotes for new content. When adding text with apostrophes or quotes, use XML entities to produce smart quotes: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-2-edit-xml)
- Adding comments: Use comment.py to handle boilerplate across multiple XML files (text must be pre-escaped XML): — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-2-edit-xml)
- Then add markers to document.xml (see Comments in XML Reference). — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-2-edit-xml)

## Step 3: Pack

- Validates with auto-repair, condenses XML, and creates DOCX. Use --validate false to skip. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-3-pack)
- Auto-repair will fix: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-3-pack)
  - durableId >= 0x7FFFFFFF (regenerates valid ID) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-3-pack)
  - Missing xml:space="preserve" on <w:t> with whitespace — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-3-pack)
- Auto-repair won't fix: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-3-pack)
  - Malformed XML, invalid element nesting, missing relationships, schema violations — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#step-3-pack)

## Common Pitfalls

- Replace entire <w:r> elements: When adding tracked changes, replace the whole <w:r>...</w:r> block with <w:del>...<w:ins>... as siblings. Don't inject tracked change tags inside a run. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#common-pitfalls)
- Preserve <w:rPr> formatting: Copy the original run's <w:rPr> block into your tracked change runs to maintain bold, font size, etc. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#common-pitfalls)

## Schema Compliance

- Element order in <w:pPr>: <w:pStyle>, <w:numPr>, <w:spacing>, <w:ind>, <w:jc>, <w:rPr> last — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#schema-compliance)
- Whitespace: Add xml:space="preserve" to <w:t> with leading/trailing spaces — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#schema-compliance)
- RSIDs: Must be 8-digit hex (e.g., 00AB1234) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#schema-compliance)

## Tracked Changes

- Inside <w:del>: Use <w:delText> instead of <w:t>, and <w:delInstrText> instead of <w:instrText>. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tracked-changes)
- Minimal edits - only mark what changes: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tracked-changes)
- Deleting entire paragraphs/list items - when removing ALL content from a paragraph, also mark the paragraph mark as deleted so it merges with the next paragraph. Add <w:del/> inside <w:pPr><w:rPr>: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tracked-changes)
- Without the <w:del/> in <w:pPr><w:rPr>, accepting changes leaves an empty paragraph/list item. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tracked-changes)
- Rejecting another author's insertion - nest deletion inside their insertion: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tracked-changes)
- Restoring another author's deletion - add insertion after (don't modify their deletion): — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#tracked-changes)

## Comments

- After running comment.py (see Step 2), add markers to document.xml. For replies, use --parent flag and nest markers inside the parent's. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#comments)
- CRITICAL: <w:commentRangeStart> and <w:commentRangeEnd> are siblings of <w:r>, never inside <w:r>. — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#comments)

## Images

- Add image file to word/media/ — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#images-1)
- Add relationship to word/_rels/document.xml.rels: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#images-1)
- Add content type to [Content_Types].xml: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#images-1)
- Reference in document.xml: — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#images-1)

## Dependencies

- pandoc: Text extraction — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#dependencies)
- docx: npm install -g docx (new documents) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#dependencies)
- LibreOffice: PDF conversion (auto-configured for sandboxed environments via scripts/office/soffice.py) — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#dependencies)
- Poppler: pdftoppm for images — [source](https://llms-explorer.com/sources/mdb-context-hub/docx/#dependencies)

## Where this helps

- Programmatically generating a new .docx report or document from data using docx-js, when a Word template alone won't cut it. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Editing an existing .docx to add tracked changes or comments while preserving the original author's formatting and structure. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Converting legacy .doc files or accepting/rejecting tracked changes in a batch document-processing pipeline. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Debugging a .docx that renders incorrectly on one platform, such as Google Docs, but not another, tracing it back to a table width or shading configuration. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Project ideas

- Build a docx-js generator for a recurring report format, applying explicit US Letter page size, DXA-based table widths, and heading-style overrides from the start. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Write an unpack → edit XML → pack pipeline for batch-inserting tracked changes and comments across many .docx files, using comment.py for the comment boilerplate. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Build a validation step into a document pipeline that runs after every generation or edit and fails the build if the resulting .docx doesn't validate. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Create a multi-column report template using SectionType.NEXT_COLUMN column breaks instead of manual layout hacks. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Antipatterns

- Using WidthType.PERCENTAGE for table widths instead of WidthType.DXA — it breaks rendering in Google Docs. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Using tables as visual dividers or horizontal rules — cells have a minimum height and render as empty boxes; a paragraph border does the job instead. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Injecting tracked-change tags inside an existing <w:r> run instead of replacing the whole run with sibling <w:del>/<w:ins> elements. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Writing Python scripts to edit document XML instead of using direct string replacement, which adds unnecessary complexity and hides exactly what changed. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Known issues

- Auto-repair on packing fixes a narrow set of issues, such as a bad durableId or missing xml:space="preserve", but explicitly won't fix malformed XML, invalid element nesting, missing relationships, or schema violations — those need manual correction before packing. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Deleting an entire paragraph's content requires also marking the paragraph mark itself as deleted, or accepting the tracked change leaves a stray empty paragraph behind. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Table rendering correctness depends on both the table's columnWidths and each cell's own width matching exactly — mismatched values can render incorrectly on some platforms even though the file validates. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- docx-js defaults to A4 page size, so any document intended for a US audience needs the page size set explicitly or it will silently render on the wrong paper size. — [source](https://llms-explorer.com/tree/word-document-manipulation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Context files

- [Word Document Manipulation](https://llms-explorer.com/downloads/sources/mdb-context-hub/docx.md)
