A hand-written, dependency-minimal Rich Text Format codec: reads RTF into the shared document-schema.js content pivot, and writes deterministic, 7-bit-ASCII RTF back out. Built against Microsoft's own RTF Specification, version 1.9.1 and Zod 4, with no third-party RTF library.
Status: under active development. The read and write paths described below are implemented and tested, including against a real-producer corpus: pnpm test:corpus runs the gitignored test/corpus/ suite against LibreOffice-produced RTF (flat-ODT sources spanning runs, colour, headings, lists, tables, and alignment, converted through Writer's own RTF filter by scripts/generate-corpus.mjs) — the corpus validated the reader cleanly across all eight fixtures with no defects found. Scope states exactly what is handled and what is not; nothing in this README describes work that is planned rather than done.
Every construct that remains unhandled is either a gap in document-schema.js rather than in this codec (character scaling, kerning, and run background colour have no field to land in), or something RTF itself does not specify at all beyond its own form-field vocabulary (a docx-style rich-text SDT has no RTF spelling of any kind) — see Deliberately not handled, which says which of the two each row is.
RTF is the cleanest structural fit of any format this family did not already handle. It is a wordprocessing format through and through — paragraphs, runs, character properties, paragraph properties, tables, lists and pictures all have direct ContentDocument equivalents — and it can express more of the wordprocessing variant than markdown can, carrying colour, font family, font size, alignment, vertical position and text direction natively. No document-schema.js model change was needed for it.
What it is not is another XML format. RTF is tokenised plain text with a brace-nested group and destination model, so none of the XML plumbing ooxml.js and odf.js share applies here: this package carries its own byte lexer, its own destination state machine, its own \uN/\ucN Unicode handling with code-page fallback, and its own parsers for the five header mini-formats. The closest relative in this workspace is markdown-codec, which is likewise a hand-written scanner and parser for a non-XML text format rather than a wrapper around a document library.
graph TD
archive("archive-codec")
bytes("byte-codec")
schema("document-schema.js")
rtfcodec("rtf-codec")
archive --> rtfcodec
bytes --> rtfcodec
schema --> rtfcodec
click archive "https://github.com/ExaDev/documents.js/tree/main/packages/archive-codec" "archive-codec"
click bytes "https://github.com/ExaDev/documents.js/tree/main/packages/byte-codec" "byte-codec"
click schema "https://github.com/ExaDev/documents.js/tree/main/packages/document-schema.js" "document-schema.js"
click rtfcodec "https://github.com/ExaDev/documents.js/tree/main/packages/rtf-codec" "rtf-codec"
style rtfcodec fill:#f9a825,stroke:#333,stroke-width:3px
rtf-codec depends on document-schema.js for the content pivot and archive-codec for the [MS-CFB] container an embedded object's \objdata carries — see Embedded objects and Dependency choices. It is reachable from documents.js's conversion engine, and so from document-cli, document-mcp, and the web UI, as an ordinary source and target format.
pnpm add rtf-codecimport { readRtf, writeRtf, readRtfContent, writeRtfContent } from "rtf-codec";
// The tree-form pair, over document-schema.js's DocumentTree — what to reach for by default.
const { documentPackage, diagnostics } = readRtf(await file.bytes());
const bytes = writeRtf(documentPackage);
// The flat pair, over its ContentDocument — the shape the reader itself builds.
const { document } = readRtfContent(await file.bytes());
const flatBytes = writeRtfContent(document);Every entry point takes bytes, not a string. RTF is defined over bytes: \'hh names a raw byte decoded through whichever code page the document declared, and \binN is followed by literally N arbitrary bytes. A caller who has already decoded a .rtf file as UTF-8 has destroyed exactly the information the code-page layer needs. For the one string form that genuinely still holds bytes — a file read with a latin-1/binary reader — rtfBytesFromLatin1 converts it exactly, and throws above U+00FF rather than truncating.
Both encodings are also available as z.codec() pairs, matching the convention markdown-codec and pdf-codec already follow:
import { rtfCodec, rtfContentCodec, RtfBytesSchema } from "rtf-codec";
const documentPackage = rtfCodec.parse(bytes); // bytes -> DocumentTree
const roundTripped = rtfCodec.encode(documentPackage); // DocumentTree -> bytesRtfBytesSchema is a real magic-byte check — the <File> production requires an RTF document to begin {\rtf, so a caller handing the codec a docx or a PDF is refused at the schema boundary rather than deep inside the tokenizer.
Everything here is implemented against Microsoft's own Rich Text Format (RTF) Specification, Version 1.9.1 (March 2008, 278 pages) — the final revision, covering Word 2007. Each source module cites the section it implements by name.
- Primary source:
[MSFT-RTF].pdf, hosted in Microsoft's own Office protocol documentation archive. - Microsoft's original download page: https://www.microsoft.com/en-us/download/details.aspx?id=10725 (Wayback snapshot).
- The version history and the note that 1.9.1 is the final revision: Rich Text Format on Wikipedia (Wayback snapshot).
- Format-preservation context: Library of Congress, Sustainability of Digital Formats — RTF (Wayback snapshot).
The code-page tables in src/codepage.ts were generated, not transcribed: each is bytes([b]).decode(codec) over 0x80..0xFF from Python's own codec library, verified byte-for-byte against it, because a hand-typed 128-entry table is exactly where one transposed character hides until a real document decodes wrong.
The five East Asian double-byte code-page tables in src/codepage-dbcs.ts (932 Shift-JIS, 936 GBK/GB2312, 949 UHC/Hangul, 950 Big5, 1361 Johab) were generated the same way, at larger scale: scripts/generate-dbcs-tables.py decodes every lead-byte/trail-byte pair through Python's own cp932/cp936/cp949/cp950/cp1361 codecs (https://docs.python.org/3/library/codecs.html#standard-encodings) — the exact Windows code pages RTF's own \ansicpgN/\cpgN name by that number, not a nearby web-oriented substitute. That script's own header comment has the full citation, including which pages were cross-checked against a second, independent decoder (Node's ICU-backed TextDecoder, for the four of the five pages the WHATWG Encoding Standard also defines) and where the two genuinely diverge.
Five stages, each its own module, each testable on its own:
| Stage | Module | What it does |
|---|---|---|
| Lex | src/tokenize.ts |
Bytes to a flat token stream: control words (32-letter name cap, 10-digit signed parameter, one-space delimiter), control symbols (no delimiter at all), the \'hh hex byte as its own token kind, \binN's raw byte run, and CR/LF handling. |
| Group | src/group.ts |
Brace matching and destination identification — the two structural facts every stage above the lexer needs. |
| Header | src/header.ts |
The five header mini-formats: \fonttbl, \colortbl, \stylesheet, \listtable and \listoverridetable, plus \info and the document properties, in one pass ahead of the body. |
| Read | src/read.ts |
The destination/group state machine that turns the token stream into a ContentDocument. |
| Write | src/write.ts |
The inverse: mints the header tables from what the document actually uses, then emits a body that references them by index. |
Supporting modules: src/codepage.ts (byte-to-character tables and the \ansicpgN/\fcharsetN/\cpgN precedence, plus the lead-byte state machine the five DBCS pages in src/codepage-dbcs.ts need), src/base64.ts (hex and base64 conversion for picture and object payloads), src/units.ts (twips, half-points, pixels), src/list-id.ts (the opaque numId grammar), src/constructs.ts (the fidelity-construct descriptor shapes and the DTTM bit field), src/cell-format.ts (the <celldef> border, shading, and merge production), src/embedded-object.ts (the \object/\objdata payload — JSON in, real [MS-CFB] compound file out, via archive-codec; see Embedded objects), src/diagnostics.ts (the three-tier diagnostic policy).
"Conventions of an RTF Reader" states the model this reader implements exactly: an opening brace stores the current state on a stack, a closing brace retrieves it, a backslash collects a control word or symbol and dispatches on it, and anything else is text written "to the current destination using the current formatting properties". Four kinds of state ride that stack, as the spec enumerates them — destination, character properties, paragraph properties, table properties — plus the \ucN skip count, which the spec separately requires be stacked.
The destination is not merely a label: it decides what happens to text. Body text becomes runs; a \pict destination's text is hex picture payload; a \fldinst destination's text is a field instruction to be parsed rather than shown; a \listtext destination's text is the flat rendering of a list number that "should be ignored by any reader that understands Word 97 through Word 2007 numbering"; an unrecognised {\* destination's text is discarded whole. That mapping is what lets the reader be a single pass with no lookahead beyond a group's own head.
"There is no RTF table group; instead, tables are specified as paragraph properties." A row is a run of \intbl paragraphs terminated by \cell marks and closed by \row, with the row's own \trowd ... \cellxN definition sitting before it, after it, or — for Word 2002 onward — both. The table builder is therefore driven by the \cell/\row marks in the text stream rather than by nesting, and a table closes when a non-table paragraph arrives.
\uN carries the character and is followed by an ANSI approximation a Unicode-aware reader must skip: "the reader should ignore the next N' characters, where N' corresponds to the last \ucN' value encountered", where "any RTF control word or symbol is considered a single character" and a brace ends the skippable run early. All three of those rules are implemented, including partial consumption of a text run — which is why the main loop carries a byte offset alongside its token index. {\upr {ansi} {\*\ud unicode}} pairs take the \ud half and discard the ANSI one.
On the way out, every non-ASCII character leaves as \uN with a one-character ? fallback under a single \uc1. The writer deliberately does not hunt for a code page that could carry a character as a \'hh byte: the output is then pure 7-bit ASCII whatever the input contained, which is what makes it safe to transmit and trivially diffable, and costs a conforming reader nothing.
| Construct | Handled |
|---|---|
Groups, destinations, {\* ignorable destinations |
Yes — per the spec's own reader conventions |
Control words, control symbols, \'hh, \binN |
Yes |
\uN / \ucN with ANSI fallback skipping, \upr/\ud |
Yes |
| Code pages | \ansi/\mac/\pc/\pca, \ansicpgN, per-font \cpgN/\fcharsetN; the Windows, OEM and Macintosh single-byte pages, the five East Asian DBCS pages (932 Shift-JIS, 936 GBK/GB2312, 949 UHC/Hangul, 950 Big5, 1361 Johab), plus UTF-8 |
\fonttbl |
Face name, family keyword, per-font code page |
\colortbl |
RGB, including a theme colour's own literal RGB; index 0 is the auto colour |
\stylesheet |
Paragraph style names and heading levels (\outlinelevelN or a built-in heading N name) |
\listtable / \listoverridetable |
\lsN → \listidN → the level's \levelnfcN and \levelstartatN, with each \lfolevel's own start-at or whole-level override applied |
\*\revtbl |
The revision authors \revauthN and its siblings index into |
| Sections | \sect, \sectd, the \pgwsxnN/\marg*sxnN geometry family, and the \sbk* break vocabulary |
| Paragraphs | \par, \pard, alignment, indents, spacing, \slN/\slmultN, \pagebb |
| Runs | \b, \i, \ul (every variant), \strike, \super/\sub and \upN/\dnN onto ContentRun.verticalAlign (\nosupersub as the off-spelling), \fN, \fsN, \cfN, \v (dropped as hidden) |
| Text direction | The four scopes RTF states it at: \rtlch/\ltrch onto ContentRun.direction, \rtlpar/\ltrpar onto ContentParagraph.direction, \rtlrow/\ltrrow onto ContentTableRow.direction, and \rtldoc/\ltrdoc onto LayoutMetadata.direction. \rtlsect/\ltrsect stay unmapped: they state a section's own column-snaking direction, a page-layout fact ContentSection carries no field for, not the text-direction scope the four fields name |
| Tables | \trowd, \cellxN, \trleftN, \cell, \row, multi-paragraph cells |
| Table cells | \clbrdrt/l/b/r with the whole <brdr> production, \clvertalc/\clvertalb (and \clvertalt, whose stated default collapses into the absence ContentTableCell.verticalAlign already means) as the cell's vertical alignment, \clcbpatN/\clcfpatN/\clshdngN shading (a real two-colour pattern fill, not just a flat colour — see below), and both merge families (\clvmgf/\clvmrg, \clmgf/\clmrg) |
| Bookmarks | \*\bkmkstart/\*\bkmkend as anchor constructs, with \bkmkcolfN/\bkmkcollN quarantined as residue |
| Revision marks | The whole <chrev> production as provenance constructs: \revised, \deleted, \mvf/\mvt, \crauthN, with authors and \revdttmN dates |
| Form fields | A \field whose \*\fldinst names FORMTEXT/FORMCHECKBOX/FORMDROPDOWN, plus whatever \*\formfield data it carries (\*\ffname as its tag, \*\ffhelptext as its alias when \ffownhelp says it is author-set rather than auto-generated, \ffprot as its lock, a checkbox's own \ffres/\ffdefres as its checked state, a dropdown's own \ffres/\ffdefres as its selected entry alongside its \*\ffl option list), as a contentControl construct (plainText/checkbox/dropDown). A plainText field's own \*\ffdeftext group is recognised but its content is skipped whole rather than captured, since value names the control's CURRENT value, which for a text field is the wrapped-run text already carried in the extent's own children, not \ffdeftext's default/reset text (see Write below for the one direction \ffdeftext does feed value); the same applies to the four other \*\formfield destination strings RTF's own Form Fields table names alongside it (\*\ffformat, \*\ffstattext, \*\ffentrymcr, \*\ffexitmcr), each skipped whole for the identical reason — no ContentControlDescriptor field exists to carry any of them |
| Lists | \lsN, \ilvlN, with the marker type carried through the numId grammar |
| Pictures | \pngblip and \jpegblip, hex or \binN payload, \picwgoalN/\pichgoalN or \picwN/\pichN, \picscalexN/\picscaleyN |
| Hyperlinks | The HYPERLINK field production, including its \l anchor switch |
| Special characters | \tab, \line, \emdash, \endash, \bullet, the quotation marks, \~, \-, \_, \\, \{, \}, and the zero-width and directional marks |
| Page breaks | \page |
\info |
Title, author, subject, keywords |
| Embedded objects | \object\objemb -> ContentEmbeddedObjectBlock when \objdata is this package's own payload (see Embedded objects); a real, foreign OLE object degrades with a diagnostic and its own \result fallback paragraphs, when present |
Everything in the read table above has a write path, with the header tables minted from what the document actually uses: a font table entry per distinct family, a colour table per distinct colour (runs' and cells' alike), a heading N style per distinct heading level, a \listtable/\listoverridetable pair per distinct list, and a \*\revtbl per distinct revision author. Output is deterministic (the same document produces byte-identical bytes) and pure 7-bit ASCII. An embeddedObject block writes unconditionally too now (see Embedded objects), and that holds inside a table cell as well as outside one: writeCellBlocks writes a cell's own content as a run of \intbl <pict>/<obj>/paragraph groups, plus any constructStart/constructEnd bracket a bookmark spans across them with (see Deliberately not handled), so only a table or pageBreak block placed directly in a cell's own content is degraded with a diagnostic rather than embedded.
Two places where the two models genuinely differ in shape, rather than merely in spelling:
- Page geometry is stated twice. The document-level
\paperwNfamily is written once in the header from the first section's own geometry, and the section-level\pgwsxnNfamily per section — so a reader that understands neither multiple sections nor the section family still lays the document out on the right paper. - A merged region is one cell per grid column on both sides. RTF writes one
\cellxNand one\cellper grid column, andContentTablerows are dense in the same way, so a merged region is an anchor carryingcolSpan/rowSpan(\clmgfacross columns,\clvmgfdown rows) plus a block-less entry at every other position it covers (\clmrgalong its own row,\clvmrgfrom an earlier row). A covered entry keeps its own borders, shading and vertical alignment. A table that breaks the grid rule (a covered entry carrying blocks or a span, rows of differing lengths, or a region running past the grid or into another, as document-schema.js'sfindTableGridFaultdefines it) is refused with anRtfTableGridFaultErrorrather than written, since a covered position's content has no cell to go in. The reader and writer both go through the shared grid helpers indocument-schema.js(denseTableRows,walkTableGrid).
A contentControl construct mints a real \*\formfield only for the three controlTypes RTF's own vocabulary actually spells (plainText/checkbox/dropDown); any other controlType (richText, comboBox, date, and the rest) degrades through the same construct-gap diagnostic every other unrepresentable construct uses. Two contentControl extents that cross within the same paragraph — neither nests inside nor around the other, so one starts before the other ends but also ends after it does — have no valid \*\formfield brace sequence at all, since RTF's own destination is a bracket, not a range; the later-opening extent is dropped with its own diagnostic rather than emit output where each extent's closing braces close the other's groups instead of its own. The fields inside a minted \*\formfield are emitted in a fixed order this writer itself chooses — RTF 1.9.1's own "Form Fields" section gives a real Formal Syntax production for <formfield>, and further productions for <formparams>/<formstrings> that mandate a fixed order for their own members (via the spec's plain-juxtaposition operator, not its & "any order" operator) — this writer's own order is a genuine subsequence of that spec-mandated order, not a free house convention; see write.ts's own top-of-file comment on formFieldPayload for the full citation and the exact productions. This writer's own convention puts every numeric flag/index control word (\fftype, \ffownhelp, \ffprot, \ffhaslistbox, \ffdefres/\ffres) before every destination string (\*\ffname, \*\ffdeftext, \*\ffhelptext, then the \*\ffl entries), and \*\ffname itself before \*\ffhelptext within that second group. A dropDown always mints the explicit \ffhaslistbox1, never a bare \ffhaslistbox, whether or not it carries any options: [MS-DOC] 2.9.79 FFDataBits.fHasListBox must be set for a list-type field regardless of how many entries the list holds, and RTF 1.9.1's own Form Fields table states \ffhaslistboxN as a genuine N-parameterised control word ("1 if this field has list box attached to it, 0 otherwise"), so a conformant reader applying RTF's own general Value-word default would read a bare occurrence as 0/false, the opposite of what a dropdown actually has — this codec's own reader never has to apply that default here at all, since it has no \ffhaslistbox case anywhere and simply does not consult the control word on the way in. \ffres/\ffdefres are minted alongside it only when the field's own recorded selection genuinely names one of options, and omitted together in both remaining cases — no selection was ever recorded, or the recorded value names none of options — rather than guess an index that would silently point at the wrong entry. The unmatched-value case is reported through the same diagnostic sink every other unrepresentable construct in this writer uses, since it is real, signalable data loss rather than an absence. The never-selected case is different: a real producer (PHPRtfLite, per this package's own read.test.ts fixtures) spells "no current selection" as \ffres25 (FFDataBits' own undefined-selection sentinel) plus a genuine \ffdefres0, not by omitting both — but this writer cannot emit that exact form without reintroducing the ambiguity an earlier round of it removed, because this codec's own reader deliberately falls a sentinel \ffres25 through to \ffdefres (to recover a real PHPRtfLite checkbox's meaningful reset default rather than reading it as unchecked), and that same fallback would read a written \ffdefres0 back as "option 0 is selected" rather than "nothing is selected". Omitting both fields instead sidesteps that: this reader tolerates the omission cleanly and decodes it as an unset value with no ambiguity, at the cost of not matching the form a real producer would actually write for the identical case. [MS-DOC] 2.9.78 FFData.wDef "MUST exist if and only if" the field is a checkbox or dropdown is a real MS-DOC production rule that this omission does not satisfy: a producer omitting wDef is spec-noncompliant but demonstrably tolerated in practice, since this reader is built to survive real-world RTF, not just conformant RTF. Options beyond [MS-DOC] 2.9.78 FFData.hsttbDropList's own 25-entry limit are truncated, also with a diagnostic — not an arbitrary cutoff, since FFDataBits' own iRes field reserves index 25 as its "undefined selection" sentinel, so a 26th real entry would collide with it. When the recorded value genuinely matched one of the truncated-away entries, the resulting unmatched-value diagnostic names truncation as the reason rather than reusing the generic "does not match any of the field's own options" wording — the value did match something, until the cap removed it, and the two causes are distinguished so the message states the one that actually happened. A plainText control's own value is minted as {\*\ffdeftext ...} (FFData.xstzTextDef, "MUST exist if and only if" the field is a text field), distinct from the field's DISPLAYED text carried by its wrapped runs — a real, reachable case: documents.js's own PDF AcroForm-to-contentControl reconstruction hands a text field exactly this {controlType:'plainText', value, ...} shape for a real /V string. value names the control's CURRENT scalar value and \ffdeftext names its DEFAULT/reset text — a genuinely different fact, not merely a different spelling of the same one — so this mis-slot is reported through the diagnostic sink like every other cross-field case this writer names, even though (unlike those) the string itself is not dropped: it lands in the RTF byte stream, just under a field that reads back as something else (see the round-trip note at the end of this paragraph). \ffprot1 ([MS-DOC] 2.9.79 FFDataBits.fProt), always written with its explicit parameter rather than bare, is minted whenever a control's lock is content or both, since both lock the control's own value; lock: container protects only the control's own removal, a fact RTF's form-field vocabulary has no bit for at all, so a container lock is dropped in full (nothing is written for it) and a both lock's own removal half is dropped alongside the \ffprot1 its content half still writes — each reported through the diagnostic sink with a message naming which of the two actually happened, rather than one message describing both. \ffownhelp1{\*\ffhelptext ...} carries a control's alias, RTF's own closest analogue to docx w:alias/PDF AcroForm's /TU alternate description; the read side honours an explicit \ffownhelp0 too, since FFDataBits.fOwnHelp being 0 means xstzHelpText "contains an empty or auto-generated string" rather than an author-set label, and promoting that text to alias regardless would misrepresent it — but a bare, unparameterised \ffownhelp reads as true rather than following that same 0-default, since LibreOffice's own RTF exporter emits exactly that bare form whenever the control model exposes a HelpText property at all, alongside genuine author-set help text, and the literal spec default would otherwise silently discard it (see read.ts's own comment on applyFormFieldControlWord's "ffownhelp" case for the full citation). A checkbox's own value — distinct from checked — has no RTF spelling at all: a real, reachable case (pdf-codec's own AcroForm reading spreads a checkbox widget's /V export-value name, e.g. 'Yes', onto value alongside the boolean checked derived from that same /V) is reported through the diagnostic sink rather than silently dropped, since unlike a dropDown's value it can never match anything RTF's \ffres/\ffdefres can name. value/checked/options recorded on a controlType that has no concept of them at all — a plainText field's checked or options, a checkbox's options, a dropDown's checked — are each reported through the identical sink rather than the writer silently reading past a field it has no branch for. The plainText value->\ffdeftext minting described above is write-only: this codec's own reader does not restore \ffdeftext back onto value on the way in, since value names the control's CURRENT value and \ffdeftext is the field's DEFAULT/reset text — for a text field the current value is whatever the wrapped runs actually carry, so a document built from a value-carrying plainText descriptor does not read back with that value on a round trip, a mismatch reported through the diagnostic sink at write time (see above) rather than left for a caller to discover only by round-tripping the document themselves.
Each of these is reported through a diagnostic rather than dropped silently — see Diagnostics.
| Construct | Why |
|---|---|
| Headers, footers, footnotes, endnotes, annotations | ContentDocument's flat form has no page-furniture or note position for them. A footnote's real home is document-schema.js's tree-only definitions table, which a codec producing the flat form cannot reach. |
| Content controls beyond RTF's own form-field vocabulary (richText, comboBox, date, picture, repeatingSection, button, index, group) | RTF 1.9.1 specifies nothing for these. It predates OOXML's w:sdt: its "Custom XML Tags" (\xmlopen/\xmlclose) are a bare namespace/name tag with no type, lock, alias or value, and \*\datastore is an opaque blob whose "format ... is unknown to RTF" by the spec's own words. \*\formfield (see the Scope table above) is the one real analogue RTF has, and covers plainText/checkbox/dropDown only. |
Code page 42 (SYMBOL_CHARSET) |
Not an encoding: its bytes are glyph indices into whichever symbol font the run names, so there is no correct Unicode for them without that font's own cmap. |
Metafile and bitmap pictures (\wmetafileN, \emfblip, \dibitmapN, \wbitmapN, \macpict) |
ContentImageBlock carries PNG and JPEG only. |
| A picture with no stated size | ContentImageBlock requires a positive width and height, and deriving them from the payload would need an image decoder this package deliberately does not carry. |
Nested tables (\nestcell/\nestrow) |
Read as ordinary cell content; the inner table's own structure is not reconstructed. |
Drawing objects (\do, \shp) |
A schema gap, not a container one. These are Word's own native in-document vector-drawing layer, not an OLE embed — there is no raw drawing-shape ContentBlock for a wordprocessing section's block flow to land in (only ContentEmbeddedObjectBlock, which names a whole embedded document, not a shape), so a \do/\shp construct is dropped regardless of the container work Embedded objects below did for \object. |
A real, foreign \object's OLE data |
This package's own \objdata payload round-trips fully (see Embedded objects); a real Word-authored OLESaveToStream structure (an actual embedded .xls/.doc/OLE-control payload) has no decoder here and degrades with a diagnostic, recovering \result's own fallback paragraphs when present instead of the real object. |
Superscript/subscript (\super, \sub, \upN, \dnN), character scaling, kerning, background colour |
Character scaling (\charscalexN), kerning (\kerningN), background colour (\cbN/\highlightN) A schema gap, not an RTF one. ContentRun carries no field for any of the three. (The vertical-position family that shared this row — \super/\sub/\upN/\dnN — now reads and writes through ContentRun.verticalAlign.) |
Cell vertical alignment (\clvertalt/\clvertalc/\clvertalb), diagonal cell borders (\cldglu/\cldgll) |
Diagonal cell borders (\cldglu/\cldgll) ContentTableCell.borders has no diagonal member, so a diagonal rule is not a side. (Cell vertical alignment, which shared this row, now reads and writes through ContentTableCell.verticalAlign.) |
| A bookmark whose two halves straddle a table cell wall | document-schema.js ratifies this as a drop rather than a shape to repair: each block list is its own bracket scope, and pairing across two of them would need the marker ids its contract deliberately refuses. |
| A contentControl extent that crosses another contentControl extent in the same paragraph | RTF's \*\formfield destination can only nest, never cross: a {\field...} group nests cleanly inside another one's own \fldrslt, but two extents that start-before/end-after each other have no valid brace sequence, since whichever closes second would close the other's own braces instead of its own. |
Two block-scoped bookmarks whose real ranges are disjoint but whose halves both land in the same paragraph (one's \bkmkend, the other's \bkmkstart) |
A paragraph is this reader's finest addressable block-extent unit, so the shared paragraph's own block index gets claimed by both bookmarks' extents even though their real text never overlapped — a genuine crossing pair, not a nesting one, and document-schema.js's own construct group has no shape for that (a tree would need two subtrees crossing, which no tree holds). The later-starting bookmark is dropped with a diagnostic; the earlier one keeps its extent intact. |
A table or pageBreak block placed directly inside a table cell |
writeCellBlocks writes a cell's own content as a run of \intbl <pict>/<obj>/paragraph groups, plus any bracketing constructStart/constructEnd markers spliced in inline exactly as writeBlock does at the top level; an image/embeddedObject block borrows the identical \intbl shell a paragraph gets (a \pict/\object group inside an ordinary \intbl paragraph is a real, already-round-trippable shape — see the tests around it), but a nested table needs its own \itapN row grammar this writer does not build, and a mid-row \page would \pard-reset the row's own \intbl state, so those two kinds are still dropped. |
RTF 1.9.1's own "Objects" section states an \object's grammar precisely: '{' \object (<objtype> & ...) <objdata> <result> '}', where \objdata is '{\*' \objdata (<objalias>? & <objsect>?) <data> '}' and <data> is the identical (\binN #BDATA) | #SDATA production \pict's own payload uses. The spec is direct about what that data actually is: "When the object is an OLE embedded or linked object, the data part of the object is the structure produced by the OLESaveToStream function." [MS-OLEDS] 2.2.5 names that structure precisely, and it has four fields, not two: an ObjectHeader (OLEVersion, a FormatID of 0x00000002, then ClassName/TopicName/ItemName as length-prefixed ANSI strings), a NativeDataSize and the NativeData that size names (the real OLE compound file is NativeData, one field inside the structure, not the whole of \objdata), and — mandatory, not optional — Presentation: "This MUST be a MetaFilePresentationObject, a BitmapPresentationObject, a DIBPresentationObject, a StandardClipboardFormatPresentationObject, or a RegisteredClipboardFormatPresentationObject." A reader that stops once NativeData ends is reading a truncated prefix of an EmbeddedObject, not a real one — a genuine OLE1.0 consumer keeps looking for Presentation and hits EOF instead. This package used to drop \object in both directions because it had no way to build or read a compound file at all; archive-codec now ships exactly that (writeCompoundFile/readCompoundFile, [MS-CFB]), plus the Package stream wrapper (writeOlePackage/readOlePackage) real Word/PowerPoint embeds use inside it — the same pair doc-codec/xls-codec/ppt-codec already depend on archive-codec for. The ObjectHeader/NativeDataSize/Presentation envelope those bytes ride inside is this module's own responsibility (src/embedded-object.ts's writeObjectHeader/readObjectHeader and writePresentationObject/skipPresentationObject), since archive-codec's own charter is container structure below OLE's own object-embedding framing, never that framing itself. Presentation is written as the smallest of the five legitimate shapes — a StandardClipboardFormatPresentationObject (2.2.3.2) wrapping a minimal 1x1 monochrome CF_DIB image — since the other four each need a full metafile/bitmap-record writer this module has no other use for; a real OLE1.0 consumer can genuinely decode and display it as the object's placeholder preview, not merely treat it as size-matching filler.
What rides inside the compound file is this codec's own JSON, not a foreign format's bytes. rtf-codec cannot depend on ooxml.js/odf.js — format codecs are peers in this family, never one another's dependency — so a wordprocessing/presentation/spreadsheet/drawing/formula ContentEmbeddedObjectBlock's own nested ContentDocument cannot be re-serialised into a real docx/pptx/xlsx/odf/MathML byte stream the way a genuine OLE server would. What this codec can write and read back losslessly is its own ContentDocument (a plain, Zod-validated, JSON-serialisable value), so src/embedded-object.ts packages that JSON as the Package stream's own "file" — the identical slot a real embed's actual docx/xlsx bytes would occupy — wraps it in a real [MS-CFB] compound file, wraps that as an ObjectHeader+NativeDataSize+NativeData+Presentation EmbeddedObject structure, and hex-encodes the whole thing as \objdata. objectKind, frame, and the anchor fields all ride the same envelope, so nothing about the embed's position or kind depends on \object's own \objw/\objh/\objclass control words — those are still written, purely as the size hint and class label a reader that cannot decode \objdata at all would fall back to, matching the spec's own advice that a producer supply them "to maintain backward compatibility."
Reading is honest about what it can and cannot decode. A real Word-authored \object — an actual embedded .xls range, a Windows Media Player control, an Equation Editor formula — carries a real OLESaveToStream NativeData with no JSON envelope inside it, and decoding that would need this package to understand every OLE server's own on-disk format, which is out of scope for the same reason a metafile picture is: no image/OLE decoder lives here. readEmbeddedObjectData tries the one path it can (ObjectHeader parse -> NativeData slice -> compound file -> Package stream -> JSON -> ContentEmbeddedObjectSchema -> confirm Presentation is genuinely present and well-formed) and returns undefined for anything else, degrading with RtfDiagnosticCodes.EMBEDDED_OBJECT_UNREADABLE rather than throwing — one unreadable \object must not fail the whole document. A payload whose NativeData is genuinely this package's own JSON but whose mandatory Presentation field is missing or malformed is rejected too, at the same tier: it is not a shape this writer ever produced, so it is not treated as a truncated match. On that degrade path, \object's own \result fallback — the rendered preview a non-\object-aware reader would show instead, per the spec's own advice that this "allows RTF readers that do not understand objects ... to use the current result, in place of the object, to maintain appearance" — is exactly what this reader needs too, so \result's own paragraphs (RTF 1.9.1's <result> = '{' \result <para>+ '}') are read as ordinary body content in the object's place; the \objw/\objh size hint that has nowhere left to land is folded into the EMBEDDED_OBJECT_UNREADABLE diagnostic's own message instead of being silently discarded, reporting whichever of the two is actually present when only one was stated. \result's own content is rendered into a totally isolated scratch accumulator the moment its group is seen, sharing no paragraph, block list, or table state with whatever the surrounding document was already accumulating around \object, and is spliced into \object's own place — in the open table cell's own blocks or the section's, whichever \object itself sits in — only once \object's group closes and \objdata is confirmed never to have decoded; if \objdata does decode, the scratch content is discarded outright rather than spliced in. That isolation is what makes the outcome correct regardless of which of <objdata>/<result> the producer wrote first, without predicting it via a lookahead that would have to independently re-derive every rule the real, group-aware parse already applies: earlier revisions tried a flat token scan ahead of \objdata's own group (which disagreed with the real parse on a payload nesting a spec-legal {\*\objalias ...}/{\*\objsect ...} sub-group, since RTF 1.9.1's own <objdata> production allows both directly inside \objdata's braces and the flat scan folded their bytes into the payload it predicted from, where the real read correctly skips them), then a block-index retraction scheme keyed to the section's or cell's own block count at the moment \result opened and closed (which mistook any text the surrounding paragraph was still accumulating, unclosed, when \object began for \result's own content, and left a bare-inline \result with no trailing \par — legal per that same <result> production, since the group's own closing brace can stand in for the final paragraph's \par — entirely unretracted, since it had closed no block for the index range to name). The isolated scratch accumulator has neither failure mode: it force-closes whatever paragraph \result's own content was still accumulating exactly as a table cell's \cell or the document's own end already do, and its finished blocks are identified by construction rather than by an index range into a list something else was concurrently writing to. A second \result sibling (RTF's own grammar allows only one, but a malformed producer can still write two) is recognised as a duplicate — mirroring how a second \objdata sibling is already handled — and discarded with its own diagnostic rather than silently overwriting the first's recovered content. An \object with no \objdata destination at all is a distinct case from one whose \objdata merely fails to decode: \result's content still recovers, but is reported with its own diagnostic, since a substitution a caller cannot see is indistinguishable from a reader that dropped the construct outright.
writeCompoundFile cannot yet set a root-storage CLSID ([MS-CFB] 2.6.1's Root Directory Entry), so even a correctly-framed \objdata payload will not self-identify its object class to a real OLE consumer the way a genuine Windows Packager-authored compound file does — a gap in archive-codec, not in this module's own framing.
import { readRtfContent, writeRtfContent } from "rtf-codec";
const written = writeRtfContent({
kind: "wordprocessing",
metadata: {},
sections: [
{
pageSize: { widthPt: 612, heightPt: 792 },
margins: { topPt: 72, rightPt: 72, bottomPt: 72, leftPt: 72 },
blocks: [
{
kind: "embeddedObject",
objectKind: "spreadsheet",
frame: { xPt: 0, yPt: 0, widthPt: 200, heightPt: 100 },
document: { kind: "spreadsheet", metadata: {}, sheets: [] },
},
],
},
],
});
// \objdata's own NativeData field (inside the ObjectHeader/NativeDataSize/NativeData/Presentation envelope this writer produces) is a real [MS-CFB] compound file — readRtfContent decodes the whole envelope back into the identical objectKind/frame/document.
const { document } = readRtfContent(written);The same three-tier policy markdown-codec and pdf-codec use: throw for input that cannot be processed at all, recover with a diagnostic for input that is malformed in a way the spec's own robustness advice says to survive, and degrade with a diagnostic for a construct read correctly at the token level whose meaning the ContentDocument mapping does not carry.
That third tier does more work here than in the XML formats. The spec requires an unknown control word to be ignored and an unknown {\* destination to be skipped whole, so "I did not understand this" is the format's normal operating mode rather than an error condition — but a reader that silently drops a construct a caller cared about is indistinguishable from one that never saw it. Every drop this package makes deliberately therefore names itself through a code in RtfDiagnosticCodes.
import { readRtf, RtfDiagnosticCodes } from "rtf-codec";
const { documentPackage, diagnostics } = readRtf(bytes);
const droppedPictures = diagnostics.filter(
(diagnostic) =>
diagnostic.code === RtfDiagnosticCodes.UNSUPPORTED_PICTURE_FORMAT,
);A coverage suite proves every code in RtfDiagnosticCodes is reachable by producing each one from a real input, and fails if any has no fixture — so the table cannot grow an entry nothing can emit, and a construct that stops being dropped has its code removed rather than left behind as a promise the package no longer keeps.
The throw tier is RtfNotAnRtfDocumentError (no {\rtf header), RtfInputTooLargeError and RtfNestingLimitExceededError (the two resource guards, both configurable through ReadRtfOptions), and RtfUnsupportedDocumentKindError on the write side, since RTF is a wordprocessing format and a presentation, spreadsheet, drawing or formula document has no RTF spelling.
document-schema.js for the content pivot, archive-codec for the [MS-CFB] container Embedded objects needs, and zod. Every third-party RTF library is banned by name in this package's own eslint.config.ts — rtf-parser, rtf.js, rtf-stream-parser, node-rtf, jsrtf, @shelf/rtf-to-html — the same bet markdown-codec makes against micromark/remark/marked and pdf-codec makes against pdf-lib/pdfjs-dist. Depending on one would defeat the reason this package exists. archive-codec is not such a library: it is zero document-format knowledge, the same sibling doc-codec/xls-codec/ppt-codec already depend on for their own [MS-CFB] container, not an RTF-aware dependency this package's own bet is against.
iconv-lite is banned for a second reason on top of that: it is Node-only (it is built on Buffer), so depending on it would break this package's Worker isomorphism. The code-page tables in src/codepage.ts exist instead.
It depends on byte-codec for one thing only: bytesToBase64, the family's single base64 encoder (ExaDev/documents.js#1282), which the read path turns a picture's recovered bytes into a ContentImageBlock payload with. What stays hand-written in src/base64.ts is the hex conversion RTF's own #SDATA payload actually needs, which no other package has a use for, and a base64 decoder that answers a different question from byte-codec's: it returns undefined for a character outside the alphabet, so a single malformed image degrades with a diagnostic instead of throwing and failing the whole write.
Like every foundation and format-codec package in this family, rtf-codec is Worker-isomorphic: its published src/ imports no node:* module and uses no Buffer, so one artifact behaves identically in a Node host, a browser, and a Cloudflare Worker. The ban is enforced by isomorphic: true in this package's eslint.config.ts, and pnpm test:workers proves it at runtime by running the public surface inside workerd.
Two places would have been tempting to write with a Node-only shortcut, and the workers suite exercises both: src/base64.ts's hand-written hex and base64 decoding (Buffer.from(text, "base64") is the one-liner they exist instead of) and src/codepage.ts's own tables.
Both of the channels document-schema.js defines are live here.
Channel 1, the harmonised construct vocabulary. Bookmarks read and write as anchor descriptors, and the whole <chrev> revision-mark family as provenance descriptors. Which of the two flat encodings a construct takes is decided by what it actually spans, exactly as the schema requires: a bookmark opening and closing inside one paragraph is a RunConstructExtent on that paragraph, one spanning whole paragraphs is a constructStart/constructEnd marker pair, and a revision mark — being a character property — is always the former. A form field (\*\formfield, gated on FORMTEXT/FORMCHECKBOX/FORMDROPDOWN) reads and writes as a contentControl construct the same way: always a RunConstructExtent, since one {\field ...} group is always inline and never spans a paragraph boundary. Content controls beyond that (a docx-style rich-text SDT and the rest) are still absent because RTF has no spelling for them at all; see the gap table above.
Channel 2, the residue channel. SourceFormatSchema gained its rtf member, so this codec can now quarantine what no semantic field carries. \bkmkcolfN/\bkmkcollN — a bookmark's table-column range — ride the anchor descriptor's own source, and the writer restores them verbatim inside its {\*\bkmkstart …} when the residue names rtf as its format, leaving another format's residue untouched. That decidability is the whole point of the format field.
pnpm install
pnpm build # tsdown -> ESM + CJS + .d.ts in dist/
pnpm typecheck # tsc for the web program and the node program, plus attw --pack
pnpm lint # eslint . --fix --cache --max-warnings 0
pnpm test # vitest run --project unit
pnpm test:workers # the same code inside workerd, the real Cloudflare Workers runtime
pnpm test:smoke # rebuilds dist/ and exercises the built ESM and CJS artifactsTo run a single test file: pnpm vitest run src/read.test.ts.
Release, CI, and commit-message conventions are workspace-wide, not package-local — see the monorepo root README for the mechanism.
Conventional Commits, enforced workspace-wide by commitlint through a root commit-msg hook. Work inside packages/rtf-codec/; see CONTRIBUTING.md for the shared git hooks and history conventions.
rtf-codec/base64 still exists and still exports this package's own base64ToBytes (which returns undefined rather than throwing, see above) and its hex conversion, but it no longer exports bytesToBase64. That encoder was this package's own copy of one every codec carried separately, and it lives in byte-codec now, as one implementation (ExaDev/documents.js#1282).
Import from byte-codec directly:
import { base64ToBytes, bytesToBase64 } from "byte-codec";MIT