上一章用状态机把 rehypeCodeGroups 讲过一遍。若仍觉得「知道在合并,但树在脑子里对不上」,本章只做一件事:拿真实编译结果,对着源码逐步跑。
这是主线之外的扩展篇。读完应能自己在纸上模拟:给定一份 MDX,tree.children 每一步怎么变。rehypePrism 不在这里。
下面的树来自与站点相同的前半段管线(remark-gfm → rehype-slug),在即将进入 rehypeCodeGroups 时截取。解析器还会给每个节点挂 position(行列号),插件完全不用它,文中一律省略。
插件接到的是什么§
lib/mdx.tsx 里写的是:
rehypePlugins: [
[rehypeSlug, { prefix: "" }],
rehypeCodeGroups, // ← 工厂本身,不是工厂的返回值
rehypePrism,
]unified 会先调用 rehypeCodeGroups(),得到真正改树的函数,再把整棵 HAST 喂进去:
export function rehypeCodeGroups() {
return function (tree: Root) {
const walk = (parent) => { /* … */ };
walk(tree);
};
}所以「给到此函数」的 tree,类型是 HAST 的 Root:{ type: "root", children: [...] }。它已经不是 Markdown 字符串,也还不是 React 元素——只是一棵描述 HTML 的对象树。
两个会反复出现的节点形状:
type | 含义 | 插件怎么认 |
|---|---|---|
"element" | 一个标签,如 p / pre / code | isElement:有 tagName |
"text" | 纯文本,含块与块之间的 "\n" | isBlankText:整段都是空白 |
围栏在这棵树上的固定形态是:
pre
└─ code.language-xxx
├─ children: [ { type: "text", value: "源码\n" } ]
└─ data.meta: "file=\"hi.js\"" ← 语言后面那串,还没解析
group / file 不在 pre 上,而在内层 code.data.meta。这是 remark/mdx 的约定,所以函数里总是先 node.children.find(isElement) 再 metaOf。
样本一:最简 MDX(没有 group)§
作者文件(可当成一篇只有正文的稿):
你好。
```js file="hi.js"
console.log(1)
```进入 rehypeCodeGroups 时,完整 tree 如下(已去掉 position):
{
"type": "root",
"children": [
{
"type": "element",
"tagName": "p",
"properties": {},
"children": [{ "type": "text", "value": "你好。" }]
},
{ "type": "text", "value": "\n" },
{
"type": "element",
"tagName": "pre",
"properties": {},
"children": [
{
"type": "element",
"tagName": "code",
"properties": { "className": ["language-js"] },
"children": [{ "type": "text", "value": "console.log(1)\n" }],
"data": { "meta": "file=\"hi.js\"" }
}
]
}
]
}先把 root.children 看成一张有下标的表——process 只看这一层:
下标 i | 节点 | 一句话 |
|---|---|---|
| 0 | <p>你好。</p> | 普通段落 |
| 1 | 文本 "\n" | 块之间的换行,不是空白「段落」,但是空白文本 |
| 2 | <pre><code class="language-js">…</code></pre> | 唯一的围栏 |
函数不会先把整棵树「理解成文章」。它只是从 i = 0 走到末尾。
调用关系:三层盒子§
把源码叠成三层,后面逐步跑时只在最内层打转:
rehypeCodeGroups() 工厂,站点启动编译时调用一次
└─ function (tree) transformer,每篇文章一次
└─ walk(parent) 先 process 当前层 children,再对子元素递归
└─ process(nodes) 真正的线性扫描:一边读 nodes,一边写 out
walk 的两行:
parent.children = process(parent.children);
for (const c of parent.children) {
if (isElement(c)) walk(c);
}对样本一:先 walk(root) → process([p, \n, pre]) 得到新的 root.children;再进入 <p> 和 <pre> 各走一遍。<p> 里只有文本,process 原样送出;<pre> 里是 <code>,不是 <pre>,同样原样送出。合并只发生在「同一层里相邻的 <pre>」,递归是为了列表、引用块里的围栏也能被扫到。
下面只盯第一次 process(root.children)。
样本一逐步执行§
开始:nodes 长度 3,out = [],i = 0。
i = 0:段落,原样搬§
node 是 <p>。isPre(node) 为假(tagName !== "pre")。
if (!isPre(node)) {
out.push(node);
i += 1;
continue;
}结果:out = [p],i = 1。对象没被复制,只是引用推进了新数组。
i = 1:换行文本,同样原样搬§
node 是 { type: "text", value: "\n" },不是 element,更不是 pre。同一分支:out = [p, \n],i = 2。
i = 2:碰到 <pre>,进入围栏分支§
isPre(node) 为真。下一句:
const code0 = node.children.find(isElement);pre.children 只有一个元素,就是那颗 <code>。code0 找到了。
若找不到(空的 <pre>),会原样 push 后 continue——防御畸形树。样本一用不到。
读 meta、拼 file0§
const meta0 = metaOf(code0);
const file0 = { name: meta0.file, lang: langOf(code0), code: codeText(code0) };metaOf 读 code0.data.meta,得到字符串 file="hi.js",交给 parseMeta:
- 正则切出 token:
file="hi.js" - 含
=,key 为file,value 去掉引号得hi.js - 没有
group/preview/title
于是:
{
"meta0": { "group": null, "file": "hi.js", "preview": false, "title": null },
"file0": { "name": "hi.js", "lang": "js", "code": "console.log(1)\n" }
}langOf 看 className 里以 language- 开头的项,切掉前缀并小写 → "js"。
codeText 递归拼文本节点 → "console.log(1)\n"(围栏末行换行保留在原文里)。
没有 group:只装饰,不向后看§
if (!meta0.group) {
decoratePre(node, node.children, [file0], { group: null, preview: meta0.preview });
out.push(node);
i += 1;
continue;
}decoratePre 做两件事:
- 在
pre.properties上写入data-files(JSON 字符串)。group为 null、preview为 false,所以不写data-group、data-preview。 pre.children = codes。这里传入的是原来的node.children,所以内层结构不变,仍是一个<code>。
out 变成 [p, \n, 已装饰的 pre],i = 3。循环结束。
process 返回 out,赋回 root.children。树还是三个顶层节点,但第三个 <pre> 多了属性:
{
"type": "root",
"children": [
{ "type": "element", "tagName": "p", "…": "你好。" },
{ "type": "text", "value": "\n" },
{
"type": "element",
"tagName": "pre",
"properties": {
"data-files": "[{\"name\":\"hi.js\",\"lang\":\"js\",\"code\":\"console.log(1)\\n\"}]"
},
"children": [
{
"type": "element",
"tagName": "code",
"properties": { "className": ["language-js"] },
"children": [{ "type": "text", "value": "console.log(1)\n" }],
"data": { "meta": "file=\"hi.js\"" }
}
]
}
]
}要点:没有 group 不是「什么都不做」。单文件也必须带 data-files,后面的 CodeBlock 才能用同一套解析复制源码。子节点尚未高亮——那是下一个插件的事。
单文件路径不会 delete child.data。meta 仍挂在 code 上。有 group 的合并路径才会清掉(见下一节)。对客户端几乎无影响:React 不会把 data 字段渲染成 DOM 属性。
样本二:相邻同组(真正合并)§
作者:
先看一段说明。
```js group="demo" file="a.js"
const a = 1
```
```css group="demo" file="b.css"
.a { color: red }
```进入插件时的 root.children:
| 下标 | 节点 | code.data.meta |
|---|---|---|
| 0 | <p>先看一段说明。</p> | — |
| 1 | 文本 "\n" | — |
| 2 | <pre> / language-js | group="demo" file="a.js" |
| 3 | 文本 "\n" | — |
| 4 | <pre> / language-css | group="demo" file="b.css" |
完整树:
{
"type": "root",
"children": [
{
"type": "element",
"tagName": "p",
"properties": {},
"children": [{ "type": "text", "value": "先看一段说明。" }]
},
{ "type": "text", "value": "\n" },
{
"type": "element",
"tagName": "pre",
"properties": {},
"children": [
{
"type": "element",
"tagName": "code",
"properties": { "className": ["language-js"] },
"children": [{ "type": "text", "value": "const a = 1\n" }],
"data": { "meta": "group=\"demo\" file=\"a.js\"" }
}
]
},
{ "type": "text", "value": "\n" },
{
"type": "element",
"tagName": "pre",
"properties": {},
"children": [
{
"type": "element",
"tagName": "code",
"properties": { "className": ["language-css"] },
"children": [{ "type": "text", "value": ".a { color: red }\n" }],
"data": { "meta": "group=\"demo\" file=\"b.css\"" }
}
]
}
]
}i = 0、i = 1 与样本一相同:段落和换行进 out。关键从 i = 2 开始。
i = 2:第一个带 group 的 pre§
取出 code0、meta0、file0:
{
"meta0": { "group": "demo", "file": "a.js", "preview": false, "title": null },
"file0": { "name": "a.js", "lang": "js", "code": "const a = 1\n" }
}meta0.group 为 "demo",不走单文件分支,进入吞并:
const group = meta0.group; // "demo"
const pres = [node]; // 第一个 pre(下标 2)
const files = [file0]; // 已有 a.js
const codes = []; // 先空着,合并结束再装填
let preview = meta0.preview; // false
let k = i + 1; // 从 3 往后看注意三个数组职责不同:
| 数组 | 装什么 | 最后干什么 |
|---|---|---|
pres | 被吞掉的那些 <pre> 节点 | 从里面抠 <code> |
files | { name, lang, code } 原文清单 | JSON.stringify 进 data-files |
codes | 实际的 <code> 元素 | 成为第一个 <pre> 的新 children |
k 循环:决定组有多长§
for (; k < nodes.length; k++) {
const n = nodes[k];
if (isBlankText(n)) continue;
if (!isPre(n)) break;
const c = n.children.find(isElement);
if (!c) break;
const m = metaOf(c);
if (m.group !== group) break;
pres.push(n);
files.push({ name: m.file, lang: langOf(c), code: codeText(c) });
if (m.preview) preview = true;
}k = 3:nodes[3] 是 "\n"。isBlankText 为真(type === "text" 且整段空白)。continue——不 break,组还没结束。k 由 for 头加到 4。
这就是「相邻可以夹空白」的全部实现:空白被跳过,扫描继续。若这里是 <p>,会走到 !isPre(n) 然后 break,两个围栏就不会合并。
k = 4:nodes[4] 是第二个 <pre>。不是空白;是 pre;内有 <code>;metaOf 得到 group === "demo",与当前组相同。于是:
pres变成[pre-js, pre-css]files多一项{ name: "b.css", lang: "css", code: ".a { color: red }\n" }- 该块没有
preview,preview仍为 false
k = 5:k < nodes.length 为假,循环结束。此时 k 停在 5(第一个「不能并入」的位置;这里恰好是数组尾)。
end:再吃掉组后面的空白§
let end = k;
while (end < nodes.length && isBlankText(nodes[end])) end += 1;k 已是 5,没有更多节点,end = 5。若组后面还有 "\n",这里会把它们算进「已消费区间」,避免合并后 out 里留下一段多余空白。
装填 codes,只留第一个 pre§
for (const p of pres) {
for (const child of p.children) {
if (isElement(child)) {
delete child.data;
codes.push(child);
}
}
}
decoratePre(node, codes, files, { group, preview });
out.push(node);
i = end;pres 里两个 pre 各有一个 <code>,codes 成为 [code-js, code-css]。各自的 data.meta 被删掉——契约已经写进 files,不必把 meta 再带到客户端。
decoratePre 作用在第一个 pre(i = 2 那个)上:
data-group="demo"data-files为两个文件的 JSONchildren换成两个<code>
第二个 <pre> 不会 push 进 out。i = end = 5,外层 while 结束。
顶层从 5 个节点变成 3 个:
{
"type": "root",
"children": [
{ "type": "element", "tagName": "p", "…": "先看一段说明。" },
{ "type": "text", "value": "\n" },
{
"type": "element",
"tagName": "pre",
"properties": {
"data-group": "demo",
"data-files": "[{\"name\":\"a.js\",\"lang\":\"js\",\"code\":\"const a = 1\\n\"},{\"name\":\"b.css\",\"lang\":\"css\",\"code\":\".a { color: red }\\n\"}]"
},
"children": [
{
"type": "element",
"tagName": "code",
"properties": { "className": ["language-js"] },
"children": [{ "type": "text", "value": "const a = 1\n" }]
},
{
"type": "element",
"tagName": "code",
"properties": { "className": ["language-css"] },
"children": [{ "type": "text", "value": ".a { color: red }\n" }]
}
]
}
]
}之后 walk 进入这个 <pre>,process 看到的 children 是两个 <code>:都不是 pre,原样写回。树稳定。
客户端将看到一个 pre(映射成一个 CodeBlock):Tab 文案来自 data-files[].name,当前高亮 DOM 来自 children[activeIdx]。两个来源下标对齐,是因为装填 codes 与装填 files 用的是同一段 pres 顺序。
样本三:同名 group 被段落切开§
进入时:
| 下标 | 节点 |
|---|---|
| 0 | <pre> group=a file=a.js |
| 1 | "\n" |
| 2 | <p>中间打断。</p> |
| 3 | "\n" |
| 4 | <pre> group=a file=b.js |
i = 0 进入吞并,k = 1 跳过空白,k = 2 是 <p> → !isPre → break。pres 里只有第一块。各自装饰成独立的单组 pre(仍带 data-group="a",但 data-files 只有一项)。
group 不是全局钥匙,只是当前层、当前扫描窗口里的相等判断。
把 process 收成一张决策表§
对 nodes[i],从上到下只走一条:
| 判断 | 动作 | i 怎么走 |
|---|---|---|
不是 <pre> | out.push | i + 1 |
是 <pre> 但没有元素子节点 | out.push | i + 1 |
有 code,但 meta.group 为空 | decoratePre(单文件 data-files)后 push | i + 1 |
有 group | 用 k 吞相邻同组;收集 codes/files;装饰第一个 pre;丢掉其余 pre 与组后空白 | i = end |
k 循环内部:
| 判断 | 动作 |
|---|---|
| 空白文本 | 跳过,继续 |
非 <pre> | 组结束 |
<pre> 无 code | 组结束 |
group 字符串不等(含 null) | 组结束 |
| 其余 | 并入 pres / files,组内任一 preview 则整组可预览 |
为何不在原数组上 splice§
process 另建 out,最后一次性替换。若边扫边从 nodes 里删掉已吞的 <pre>,k 与 i 会互相踩。新数组只含「还要出现在这一层的节点」:段落、换行、以及每个组留下的那一个 <pre>。
和主线的关系§
本章把 rehypeCodeGroups 从「合并算法」落实到「这一次 process 的内存怎么变」。上一章仍负责双轨契约、parseMeta 全表、以及和 rehypePrism 的顺序。下一章回到主线:app/ 如何把散篇与系列章节接到同一条 MdxContent。