Anthropic skill

docx

Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx or .dotx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file, use this skill. Do NOT use for PDFs, spreadsheets, Google Docs, or general coding tasks unrelated to document generation.

A `.docx` is a ZIP archive of XML files. Choose your approach by task:
#pdf
#documents
#api
#automation
#docx
Last updated

about 12 hours ago

Repository path

skills/docx

Source files

3

Use cases

A `.docx` is a ZIP archive of XML files. Choose your approach by task:
| Task | Approach | |---|---| | **Create** a new document | Write a `docx` (npm) script — see gotchas below | | **Edit** an existing document | `unzip` → edit `word/document.xml` → `zip` (docx-js cannot open existing files) | | **Read** content | `pandoc -t markdown file.docx` |
- **Page size defaults to A4.** For US Letter set `page: { size: { width: 12240, height: 15840 } }` (DXA; 1440 = 1″). - **Landscape:** pass portrait dimensions and `orientation: PageOrientation.LANDSCAPE` — docx-js swaps width/height internally. - **Tables need dual widths:** set `columnWidths` on the table AND `width` on every cell, both in `WidthType.DXA` (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width. - **Table shading:** use `ShadingType.CLEAR`, never `SOLID` (renders black). - **Lists:** never insert `•` literally; use a `numbering` config with `LevelFormat.BULLET`. - **`ImageRun` requires `type:`** (`"png"`, `"jpg"`, …). - **`PageBreak` must be inside a `Paragraph`.** - **Never use `\n`** — use separate `Paragraph` elements. - **TOC:** headings must use built-in `HeadingLevel.*`; custom heading styles need `outlineLevel` set or they won't appear. - **Don't use a table as a horizontal rule** — use a paragraph bottom border instead. - **Dot-leader / right-aligned-on-same-line:** use `PositionalTab` (`alignment: PositionalTabAlignment.RIGHT`, `leader: PositionalTabLeader.DOT`) inside a `TextRun`, not literal `.` or space padding.
Comments require six cross-linked files. Use the helper — directory mode when you'll also be editing `document.xml` (saves an unzip/rezip cycle), `.docx`-direct mode otherwise:

DOCX creation, editing, and analysis

A `.docx` is a ZIP archive of XML files. Choose your approach by task: | Task | Approach | |---|---| | **Create** a new document | Write a `docx` (npm) script — see gotchas below | | **Edit** an existing document | `unzip` → edit `word/document.xml` → `zip` (docx-js cannot open existing files) | | **Read** content | `pandoc -t markdown file.docx` | > Script paths below are relative to this skill's directory.

Creating with docx-js — gotchas

`docx` is preinstalled — do not run `npm install` first; write the script and `require('docx')` directly. Only if that require fails: `npm install docx`. The model knows the API; these are the footguns: - **Page size defaults to A4.** For US Letter set `page: { size: { width: 12240, height: 15840 } }` (DXA; 1440 = 1″). - **Landscape:** pass portrait dimensions and `orientation: PageOrientation.LANDSCAPE` — docx-js swaps width/height internally. - **Tables need dual widths:** set `columnWidths` on the table AND `width` on every cell, both in `WidthType.DXA` (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width. - **Table shading:** use `ShadingType.CLEAR`, never `SOLID` (renders black). - **Lists:** never insert `•` literally; use a `numbering` config with `LevelFormat.BULLET`. - **`ImageRun` requires `type:`** (`"png"`, `"jpg"`, …). - **`PageBreak` must be inside a `Paragraph`.** - **Never use `\n`** — use separate `Paragraph` elements. - **TOC:** headings must use built-in `HeadingLevel.*`; custom heading styles need `outlineLevel` set or they won't appear. - **Don't use a table as a horizontal rule** — use a paragraph bottom border instead. - **Dot-leader / right-aligned-on-same-line:** use `PositionalTab` (`alignment: PositionalTabAlignment.RIGHT`, `leader: PositionalTabLeader.DOT`) inside a `TextRun`, not literal `.` or space padding.

Verify the output

After writing a `.docx`, render it and look at it: ```bash python scripts/office/soffice.py --headless --convert-to pdf output.docx pdftoppm -jpeg -r 100 output.pdf page ls page-*.jpg # then Read the images ``` `pdftoppm` zero-pads page numbers to the width of the page count (`page-01.jpg`…`page-12.jpg`).

Editing existing documents

Legacy `.doc` files must be converted first: `python scripts/office/soffice.py --headless --convert-to docx file.doc`. ```bash unzip -q doc.docx -d unpacked/ find unpacked -type l -delete # strip symlink entries — docx from external parties is untrusted python scripts/merge_runs.py unpacked/ # coalesce fragmented runs so text is findable

edit unpacked/word/document.xml in place — do NOT reformat or pretty-print

(cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .) python scripts/office/validate.py out.docx --original doc.docx # XSD checks; --auto-repair fixes common issues

Repository files