feat(doc): 新增 DOCX 转 Markdown Skill
This commit is contained in:
@@ -64,6 +64,15 @@ plugins/<plugin>/skills/<skill>/
|
|||||||
|
|
||||||
所有新增迁移、来源更新和已迁移 Skill 的重新评估,都必须按以下顺序处理。用户未明确确认前,只能进行只读分析,不得创建、复制、改写或删除目标 Skill 文件。
|
所有新增迁移、来源更新和已迁移 Skill 的重新评估,都必须按以下顺序处理。用户未明确确认前,只能进行只读分析,不得创建、复制、改写或删除目标 Skill 文件。
|
||||||
|
|
||||||
|
### 批次策略
|
||||||
|
|
||||||
|
- 迁移前期采用“特殊样本”方式,不以固定数量为目标,而是优先覆盖不同结构、依赖、权限和风险类型。
|
||||||
|
- 特殊样本至少覆盖纯指令、脚本与生成物、只读 Git、有副作用 Git、模板化文档、大型参考资料和复合工作流;缺少某一类型时不得宣告样本阶段完成。
|
||||||
|
- 每个特殊样本原则上单独完成迁移前确认、实现验证和提交确认,以便及时修正规则。
|
||||||
|
- 特殊样本全部通过后,应先总结可复用的命名、目录、改写、验证和状态同步规则,再进入批量阶段。
|
||||||
|
- 批量阶段按同一插件、相近能力和相同风险类型组织,每批建议 6~12 个 Skill;高度同质时可整组处理,但不得跨越不同权限边界强行合批。
|
||||||
|
- 批量迁移仍执行完整的两次确认。任何样本失败或出现新类型,都应暂停扩批并补充对应特殊样本。
|
||||||
|
|
||||||
### 1. 阅读迁移计划和现状
|
### 1. 阅读迁移计划和现状
|
||||||
|
|
||||||
- 完整阅读 `migration/MIGRATION_PLAN.md`、`migration/README.md` 和 `migration/source-lock.json`。
|
- 完整阅读 `migration/MIGRATION_PLAN.md`、`migration/README.md` 和 `migration/source-lock.json`。
|
||||||
@@ -85,6 +94,8 @@ plugins/<plugin>/skills/<skill>/
|
|||||||
|
|
||||||
来源作用必须有文件依据;无法从静态内容确认的运行行为应明确标记为未验证。
|
来源作用必须有文件依据;无法从静态内容确认的运行行为应明确标记为未验证。
|
||||||
|
|
||||||
|
同一目标能力存在多个来源时,应在一个迁移方案中共同评估,明确合并、取舍或排除关系,不按来源机械创建多个目标 Skill。
|
||||||
|
|
||||||
### 3. 给出目标名称、作用和组织结构
|
### 3. 给出目标名称、作用和组织结构
|
||||||
|
|
||||||
针对每个来源 Skill,先给出建议方案:
|
针对每个来源 Skill,先给出建议方案:
|
||||||
|
|||||||
@@ -13,7 +13,7 @@ CraftKit 是一组面向 Codex 插件市场的中性 Skill 工具。项目从通
|
|||||||
| 插件 | 用途 | 当前状态 |
|
| 插件 | 用途 | 当前状态 |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `dev` | 软件设计、编码、审查与测试 | 已初始化,暂无 Skill |
|
| `dev` | 软件设计、编码、审查与测试 | 已初始化,暂无 Skill |
|
||||||
| `doc` | 文档转换、整理与写作 | 已迁移 `format-md` |
|
| `doc` | 文档转换、整理与写作 | 已迁移 `format-md`、`docx-to-md` |
|
||||||
| `git` | 分支、提交、变更提取与集成 | 已初始化,暂无 Skill |
|
| `git` | 分支、提交、变更提取与集成 | 已初始化,暂无 Skill |
|
||||||
| `knowledge` | 交接、复盘、经验与工作总结 | 已初始化,暂无 Skill |
|
| `knowledge` | 交接、复盘、经验与工作总结 | 已初始化,暂无 Skill |
|
||||||
| `skill` | Skill 创建、迁移、检查与同步 | 已初始化,暂无 Skill |
|
| `skill` | Skill 创建、迁移、检查与同步 | 已初始化,暂无 Skill |
|
||||||
|
|||||||
+34
-18
@@ -36,6 +36,8 @@
|
|||||||
|
|
||||||
## 4. 迁移批次
|
## 4. 迁移批次
|
||||||
|
|
||||||
|
迁移分为“特殊样本”和“同质批量”两个阶段。前者用于覆盖迁移机制的不同风险类型,后者在规则稳定后提高吞吐量。批次数量不固定:特殊样本原则上逐个迁移,批量阶段每批建议 6~12 个同质 Skill。
|
||||||
|
|
||||||
### 第 0 批:迁移基础设施
|
### 第 0 批:迁移基础设施
|
||||||
|
|
||||||
第 0 批属于仓库内部维护工具,不发布为市场 Skill:
|
第 0 批属于仓库内部维护工具,不发布为市场 Skill:
|
||||||
@@ -47,29 +49,41 @@
|
|||||||
|
|
||||||
完成门槛:能够对全部来源 Skill 生成稳定清单,并在不复制来源内容的情况下更新状态。
|
完成门槛:能够对全部来源 Skill 生成稳定清单,并在不复制来源内容的情况下更新状态。
|
||||||
|
|
||||||
### 第 1 批:低风险样板
|
### 第 1 阶段:特殊样本
|
||||||
|
|
||||||
先完成唯一标准样板,再按插件推进,避免一次跨模块迁移导致集中返工:
|
特殊样本不追求模块连续性,而是覆盖不同迁移机制:
|
||||||
|
|
||||||
1. `doc/format-md`
|
1. `doc/format-md`:纯指令型样本,已完成。
|
||||||
2. `doc/report`、`doc/requirements`
|
2. `doc/docx-to-md`:多来源合并、脚本、依赖和生成物样本,已完成。
|
||||||
3. `git/commit-msg`
|
3. `git/commit-msg`:只读 Git 状态分析样本。
|
||||||
4. `knowledge/handoff`
|
4. `git/branch`:修改仓库状态和二次授权边界样本。
|
||||||
|
5. `knowledge/handoff`:模板化文档与项目上下文样本。
|
||||||
|
6. `skill/guidance`:大型参考资料、索引和渐进式加载样本。
|
||||||
|
7. `skill/migrate`:脚本、检查、同步和状态追踪组成的复合工作流样本。
|
||||||
|
|
||||||
完成门槛:`doc/format-md` 固化目录模板、验证命令、行为用例和迁移记录格式后,才能继续本批其余 Skill。
|
完成门槛:每种样本均通过对应验证,并形成可复用的命名、目录、独立实现、测试、扫描和状态同步规则。出现未覆盖的新结构或权限类型时,应补充样本,不直接扩批。
|
||||||
|
|
||||||
### 第 2 批:文档转换
|
### 第 2 阶段:同质批量迁移
|
||||||
|
|
||||||
|
特殊样本完成后,按同一插件、相近能力和相同风险类型组织批量迁移:
|
||||||
|
|
||||||
|
- 每批建议 6~12 个 Skill,高度同质时可以整组处理。
|
||||||
|
- 不把只读能力与有副作用能力、纯指令与复杂脚本、普通迁移与公开资料重建强行合为一批。
|
||||||
|
- 每批仍需迁移前确认和提交前确认,并统一更新 README、迁移计划和来源状态。
|
||||||
|
- 任一验收门禁失败时暂停该批,不继续扩大范围。
|
||||||
|
|
||||||
|
### 文档转换批次
|
||||||
|
|
||||||
所属插件:`doc`
|
所属插件:`doc`
|
||||||
|
|
||||||
1. `docx-to-markdown`
|
1. `docx-to-md`
|
||||||
2. `markdown-to-docx`
|
2. `markdown-to-docx`
|
||||||
3. `excel-to-markdown`
|
3. `excel-to-markdown`
|
||||||
4. `archive-docs`
|
4. `archive-docs`
|
||||||
|
|
||||||
重点验证图片、表格、合并单元格、编码、覆盖策略和路径安全。脚本及测试数据必须独立创建。
|
重点验证图片、表格、合并单元格、编码、覆盖策略和路径安全。脚本及测试数据必须独立创建。
|
||||||
|
|
||||||
### 第 3 批:Git 工作流
|
### Git 工作流批次
|
||||||
|
|
||||||
所属插件:`git`
|
所属插件:`git`
|
||||||
|
|
||||||
@@ -78,9 +92,9 @@
|
|||||||
3. `export-changes`
|
3. `export-changes`
|
||||||
4. `integrate-branch`
|
4. `integrate-branch`
|
||||||
|
|
||||||
`summarize-commit` 已在样板批次完成。涉及提交、合并和远端操作的 Skill 必须保留明确授权边界,并保护脏工作区。
|
`commit-msg` 将作为只读 Git 特殊样本先行完成。涉及提交、合并和远端操作的 Skill 必须保留明确授权边界,并保护脏工作区。
|
||||||
|
|
||||||
### 第 4 批:知识管理
|
### 知识管理批次
|
||||||
|
|
||||||
所属插件:`knowledge`
|
所属插件:`knowledge`
|
||||||
|
|
||||||
@@ -92,7 +106,7 @@
|
|||||||
|
|
||||||
来源中与经验初始化、提升和回扫相关的多个能力统一合并为 `maintain-lessons`,通过模式区分具体工作。
|
来源中与经验初始化、提升和回扫相关的多个能力统一合并为 `maintain-lessons`,通过模式区分具体工作。
|
||||||
|
|
||||||
### 第 5 批:开发主流程
|
### 开发主流程批次
|
||||||
|
|
||||||
所属插件:`dev`
|
所属插件:`dev`
|
||||||
|
|
||||||
@@ -119,7 +133,7 @@ plan-change
|
|||||||
|
|
||||||
多个来源中的代码检查、代码审查能力合并为 `review-code`,通过工作区、提交和分支三种模式覆盖。
|
多个来源中的代码检查、代码审查能力合并为 `review-code`,通过工作区、提交和分支三种模式覆盖。
|
||||||
|
|
||||||
### 第 6 批:项目规范与前端辅助
|
### 项目规范与前端辅助批次
|
||||||
|
|
||||||
1. `dev/select-component`
|
1. `dev/select-component`
|
||||||
2. `dev/design-style`
|
2. `dev/design-style`
|
||||||
@@ -129,7 +143,7 @@ plan-change
|
|||||||
|
|
||||||
这里只实现读取和维护“当前项目自身规范”的机制,不随 CraftKit 提供任何来源项目规范。
|
这里只实现读取和维护“当前项目自身规范”的机制,不随 CraftKit 提供任何来源项目规范。
|
||||||
|
|
||||||
### 第 7 批:基于公开资料重建
|
### 基于公开资料重建批次
|
||||||
|
|
||||||
以下能力不从来源文本改写,而是根据官方资料重新设计:
|
以下能力不从来源文本改写,而是根据官方资料重新设计:
|
||||||
|
|
||||||
@@ -142,7 +156,7 @@ plan-change
|
|||||||
|
|
||||||
中性规格必须记录所采用的公开标准、官方文档和许可证信息。
|
中性规格必须记录所采用的公开标准、官方文档和许可证信息。
|
||||||
|
|
||||||
### 第 8 批:暂缓或排除
|
### 暂缓或排除
|
||||||
|
|
||||||
默认排除:
|
默认排除:
|
||||||
|
|
||||||
@@ -170,7 +184,7 @@ plan-change
|
|||||||
| 多套前端设计 | `design-frontend` |
|
| 多套前端设计 | `design-frontend` |
|
||||||
| 多套后端实现 | `implement-backend` |
|
| 多套后端实现 | `implement-backend` |
|
||||||
| 多套前端实现 | `implement-frontend` |
|
| 多套前端实现 | `implement-frontend` |
|
||||||
| 多套 Word 转 Markdown | `docx-to-markdown` |
|
| 多套 Word 转 Markdown | `docx-to-md` |
|
||||||
| 经验初始化、提升、回扫 | `maintain-lessons` |
|
| 经验初始化、提升、回扫 | `maintain-lessons` |
|
||||||
| 规范检索和索引维护 | `retrieve-guidance`、`maintain-guidance` |
|
| 规范检索和索引维护 | `retrieve-guidance`、`maintain-guidance` |
|
||||||
|
|
||||||
@@ -211,4 +225,6 @@ plan-change
|
|||||||
- [x] 实现 Skill 与敏感内容检查脚本。
|
- [x] 实现 Skill 与敏感内容检查脚本。
|
||||||
- [x] 实现迁移台账预览和更新脚本。
|
- [x] 实现迁移台账预览和更新脚本。
|
||||||
- [x] 迁移并验证首个样板 `doc/format-md`。
|
- [x] 迁移并验证首个样板 `doc/format-md`。
|
||||||
- [ ] 样板通过后推进第 1 批其余 Skill。
|
- [x] 迁移并验证脚本型特殊样本 `doc/docx-to-md`。
|
||||||
|
- [ ] 完成其余特殊样本并总结批量迁移规则。
|
||||||
|
- [ ] 按插件和风险类型推进同质批量迁移。
|
||||||
|
|||||||
@@ -55,7 +55,10 @@
|
|||||||
"source-a:fdde0da43c1ec4c7": {
|
"source-a:fdde0da43c1ec4c7": {
|
||||||
"sourcePathHash": "fdde0da43c1ec4c7fddad96b9172435af6be96ceafa2cc824f70e04a3cf11085",
|
"sourcePathHash": "fdde0da43c1ec4c7fddad96b9172435af6be96ceafa2cc824f70e04a3cf11085",
|
||||||
"sourceSha256": "46a02dbff9c3fff57a0802853e2424983927f53c7f1cc3d8a42542403e62a161",
|
"sourceSha256": "46a02dbff9c3fff57a0802853e2424983927f53c7f1cc3d8a42542403e62a161",
|
||||||
"status": "pending"
|
"status": "migrated",
|
||||||
|
"target": "plugins/doc/skills/docx-to-md",
|
||||||
|
"targetVersion": "0.1.0",
|
||||||
|
"reviewedAt": "2026-08-25"
|
||||||
},
|
},
|
||||||
"source-a:4efb6f9dd0f39dff": {
|
"source-a:4efb6f9dd0f39dff": {
|
||||||
"sourcePathHash": "4efb6f9dd0f39dff13339e62ff24b672552efcbcedb007e9ea86fe41950adbb0",
|
"sourcePathHash": "4efb6f9dd0f39dff13339e62ff24b672552efcbcedb007e9ea86fe41950adbb0",
|
||||||
@@ -140,7 +143,10 @@
|
|||||||
"source-b:d6d4166800584a61": {
|
"source-b:d6d4166800584a61": {
|
||||||
"sourcePathHash": "d6d4166800584a61ad6cbf96af57bb8ee269a16babba753f267b48efd66f6d2d",
|
"sourcePathHash": "d6d4166800584a61ad6cbf96af57bb8ee269a16babba753f267b48efd66f6d2d",
|
||||||
"sourceSha256": "60f8708db24f368ecec55bd3f326683fac60bd6776ac69e0fafaf4c77333e703",
|
"sourceSha256": "60f8708db24f368ecec55bd3f326683fac60bd6776ac69e0fafaf4c77333e703",
|
||||||
"status": "pending"
|
"status": "migrated",
|
||||||
|
"target": "plugins/doc/skills/docx-to-md",
|
||||||
|
"targetVersion": "0.1.0",
|
||||||
|
"reviewedAt": "2026-08-25"
|
||||||
},
|
},
|
||||||
"source-b:023767d9c7be6295": {
|
"source-b:023767d9c7be6295": {
|
||||||
"sourcePathHash": "023767d9c7be62956d12fd01c94ac676b9fe6719741acd6f412c8b446efa2760",
|
"sourcePathHash": "023767d9c7be62956d12fd01c94ac676b9fe6719741acd6f412c8b446efa2760",
|
||||||
|
|||||||
@@ -0,0 +1,117 @@
|
|||||||
|
import importlib.util
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
SCRIPT = (
|
||||||
|
Path(__file__).parents[2]
|
||||||
|
/ "plugins"
|
||||||
|
/ "doc"
|
||||||
|
/ "skills"
|
||||||
|
/ "docx-to-md"
|
||||||
|
/ "scripts"
|
||||||
|
/ "convert.py"
|
||||||
|
)
|
||||||
|
SPEC = importlib.util.spec_from_file_location("docx_to_md_convert", SCRIPT)
|
||||||
|
MODULE = importlib.util.module_from_spec(SPEC)
|
||||||
|
assert SPEC.loader is not None
|
||||||
|
sys.modules[SPEC.name] = MODULE
|
||||||
|
SPEC.loader.exec_module(MODULE)
|
||||||
|
|
||||||
|
|
||||||
|
class DocxToMarkdownTest(unittest.TestCase):
|
||||||
|
"""验证转换器的核心内容、覆盖保护和输入边界。"""
|
||||||
|
|
||||||
|
def make_docx(self, path: Path) -> None:
|
||||||
|
"""创建只包含公开 OOXML 结构的最小测试文档。"""
|
||||||
|
|
||||||
|
document = """<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"
|
||||||
|
xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships"
|
||||||
|
xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main">
|
||||||
|
<w:body>
|
||||||
|
<w:p><w:pPr><w:pStyle w:val="Heading1"/></w:pPr><w:r><w:t>测试标题</w:t></w:r></w:p>
|
||||||
|
<w:p><w:r><w:rPr><w:b/></w:rPr><w:t>加粗正文</w:t></w:r>
|
||||||
|
<w:hyperlink r:id="rLink"><w:r><w:t>示例链接</w:t></w:r></w:hyperlink></w:p>
|
||||||
|
<w:p><w:pPr><w:numPr><w:ilvl w:val="0"/><w:numId w:val="1"/></w:numPr></w:pPr>
|
||||||
|
<w:r><w:t>列表项目</w:t></w:r></w:p>
|
||||||
|
<w:p><w:r><w:drawing><a:blip r:embed="rImage"/></w:drawing></w:r></w:p>
|
||||||
|
<w:tbl>
|
||||||
|
<w:tr><w:tc><w:p><w:r><w:t>名称</w:t></w:r></w:p></w:tc><w:tc><w:p><w:r><w:t>值</w:t></w:r></w:p></w:tc></w:tr>
|
||||||
|
<w:tr><w:tc><w:p><w:r><w:t>A</w:t></w:r></w:p></w:tc><w:tc><w:p><w:r><w:t>1</w:t></w:r></w:p></w:tc></w:tr>
|
||||||
|
</w:tbl>
|
||||||
|
</w:body>
|
||||||
|
</w:document>"""
|
||||||
|
styles = """<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<w:styles xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">
|
||||||
|
<w:style w:type="paragraph" w:styleId="Heading1"><w:name w:val="heading 1"/></w:style>
|
||||||
|
</w:styles>"""
|
||||||
|
numbering = """<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<w:numbering xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">
|
||||||
|
<w:abstractNum w:abstractNumId="0"><w:lvl w:ilvl="0"><w:numFmt w:val="bullet"/></w:lvl></w:abstractNum>
|
||||||
|
<w:num w:numId="1"><w:abstractNumId w:val="0"/></w:num>
|
||||||
|
</w:numbering>"""
|
||||||
|
relationships = """<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships">
|
||||||
|
<Relationship Id="rLink" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/hyperlink" Target="https://example.com" TargetMode="External"/>
|
||||||
|
<Relationship Id="rImage" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/image" Target="media/test.png"/>
|
||||||
|
</Relationships>"""
|
||||||
|
with zipfile.ZipFile(path, "w") as archive:
|
||||||
|
archive.writestr("word/document.xml", document)
|
||||||
|
archive.writestr("word/styles.xml", styles)
|
||||||
|
archive.writestr("word/numbering.xml", numbering)
|
||||||
|
archive.writestr("word/_rels/document.xml.rels", relationships)
|
||||||
|
archive.writestr("word/media/test.png", b"\x89PNG\r\n\x1a\n")
|
||||||
|
archive.writestr("word/comments.xml", "<comments/>")
|
||||||
|
|
||||||
|
def test_convert_common_content_and_report_warning(self) -> None:
|
||||||
|
"""常见结构应转换,无法处理的批注应明确告警。"""
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as temp:
|
||||||
|
root = Path(temp)
|
||||||
|
source = root / "sample.docx"
|
||||||
|
output = root / "result"
|
||||||
|
self.make_docx(source)
|
||||||
|
|
||||||
|
result = MODULE.DocxConverter(source, output).convert()
|
||||||
|
markdown = result.output_file.read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
self.assertIn("# 测试标题", markdown)
|
||||||
|
self.assertIn("**加粗正文**", markdown)
|
||||||
|
self.assertIn("[示例链接](https://example.com)", markdown)
|
||||||
|
self.assertIn("- 列表项目", markdown)
|
||||||
|
self.assertIn("| 名称 | 值 |", markdown)
|
||||||
|
self.assertIn("
|
||||||
|
self.assertEqual(result.images, 1)
|
||||||
|
self.assertTrue(any("批注" in warning for warning in result.warnings))
|
||||||
|
|
||||||
|
def test_existing_output_requires_force(self) -> None:
|
||||||
|
"""默认不得写入已有输出目录,显式覆盖后才可继续。"""
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as temp:
|
||||||
|
root = Path(temp)
|
||||||
|
source = root / "sample.docx"
|
||||||
|
output = root / "result"
|
||||||
|
self.make_docx(source)
|
||||||
|
output.mkdir()
|
||||||
|
|
||||||
|
with self.assertRaises(FileExistsError):
|
||||||
|
MODULE.DocxConverter(source, output).convert()
|
||||||
|
|
||||||
|
result = MODULE.DocxConverter(source, output, force=True).convert()
|
||||||
|
self.assertTrue(result.output_file.exists())
|
||||||
|
|
||||||
|
def test_cli_rejects_non_docx(self) -> None:
|
||||||
|
"""命令行入口应拒绝扩展名不正确的文件。"""
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as temp:
|
||||||
|
source = Path(temp) / "sample.txt"
|
||||||
|
source.write_text("not docx", encoding="utf-8")
|
||||||
|
self.assertEqual(MODULE.main([str(source)]), 2)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
---
|
||||||
|
name: docx-to-md
|
||||||
|
description: 将 Word .docx 文档转换为 Markdown,并提取基础结构、表格、链接和图片。适用于需要可读文本版本的 Word 文档;旧版 .doc、版式级还原或 Word 内容编辑不应触发本 Skill。
|
||||||
|
---
|
||||||
|
|
||||||
|
# Word 转 Markdown
|
||||||
|
|
||||||
|
使用 `scripts/convert.py` 将 `.docx` 转换为 Markdown。转换目标是保留可表达的内容结构,不承诺版式或内容无损还原。
|
||||||
|
|
||||||
|
## 执行边界
|
||||||
|
|
||||||
|
- 只接受存在的 `.docx` 文件,不把扩展名不同的文件当作 Word 文档处理。
|
||||||
|
- 默认输出到输入文件旁的同名目录;目录已存在时停止,不自动覆盖或删除内容。
|
||||||
|
- 只有用户明确同意覆盖时才传入 `--force`。该选项仅覆盖同名生成文件,不清理目录中的其他文件。
|
||||||
|
- 不自动安装依赖、访问网络或调用办公软件。
|
||||||
|
- 不擅自删除封面、目录、批注或其他语义内容,也不在转换时改写正文。
|
||||||
|
|
||||||
|
## 转换流程
|
||||||
|
|
||||||
|
1. 确认输入文件和输出位置;用户未指定时使用默认位置。
|
||||||
|
2. 执行:
|
||||||
|
|
||||||
|
```text
|
||||||
|
python scripts/convert.py <input.docx> [--output-dir <directory>] [--force]
|
||||||
|
```
|
||||||
|
|
||||||
|
3. 检查命令返回码、Markdown 文件、图片目录和警告清单。
|
||||||
|
4. 抽查标题、段落、列表、表格、链接和图片引用是否可读。
|
||||||
|
5. 向用户报告输出路径、转换统计和无法可靠处理的内容。
|
||||||
|
|
||||||
|
## 支持范围
|
||||||
|
|
||||||
|
转换器处理正文段落、常见标题样式、基础有序/无序列表、普通表格、外部超链接、粗体、斜体和嵌入图片。
|
||||||
|
|
||||||
|
以下内容可能降级或只产生警告:批注、修订、复杂编号、合并单元格、文本框、页眉页脚、脚注、公式、SmartArt、图表和内嵌附件。发现警告时不得声称转换完整;需要版式级保真时应改用专门的文档工具。
|
||||||
|
|
||||||
|
转换完成后,只有用户另外要求整理排版时,才使用 `format-md` 处理生成的 Markdown。
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
interface:
|
||||||
|
display_name: "DOCX to Markdown"
|
||||||
|
short_description: "将 Word 文档转换为 Markdown 并提取图片"
|
||||||
|
default_prompt: "使用 $docx-to-md 将这份 Word 文档转换为 Markdown,并报告转换警告。"
|
||||||
@@ -0,0 +1,329 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""使用 Python 标准库将 DOCX 的常见内容转换为 Markdown。"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import zipfile
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
from pathlib import Path, PurePosixPath
|
||||||
|
from xml.etree import ElementTree as ET
|
||||||
|
|
||||||
|
|
||||||
|
NS = {
|
||||||
|
"w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main",
|
||||||
|
"r": "http://schemas.openxmlformats.org/officeDocument/2006/relationships",
|
||||||
|
"a": "http://schemas.openxmlformats.org/drawingml/2006/main",
|
||||||
|
"pr": "http://schemas.openxmlformats.org/package/2006/relationships",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def qname(prefix: str, local: str) -> str:
|
||||||
|
"""生成 ElementTree 使用的完整命名空间标签。"""
|
||||||
|
|
||||||
|
return f"{{{NS[prefix]}}}{local}"
|
||||||
|
|
||||||
|
|
||||||
|
def escape_markdown(text: str) -> str:
|
||||||
|
"""转义会破坏普通行结构的 Markdown 字符。"""
|
||||||
|
|
||||||
|
return text.replace("\\", "\\\\").replace("|", "\\|")
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Result:
|
||||||
|
"""记录转换产物与可向用户展示的统计信息。"""
|
||||||
|
|
||||||
|
output_file: Path
|
||||||
|
images_dir: Path
|
||||||
|
paragraphs: int = 0
|
||||||
|
tables: int = 0
|
||||||
|
images: int = 0
|
||||||
|
warnings: list[str] = field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
|
class DocxConverter:
|
||||||
|
"""读取 DOCX 包并转换常见 OOXML 块级元素。"""
|
||||||
|
|
||||||
|
def __init__(self, source: Path, output_dir: Path, force: bool = False) -> None:
|
||||||
|
self.source = source
|
||||||
|
self.output_dir = output_dir
|
||||||
|
self.force = force
|
||||||
|
self.output_file = output_dir / f"{source.stem}.md"
|
||||||
|
self.images_dir = output_dir / "images"
|
||||||
|
self.relationships: dict[str, tuple[str, str]] = {}
|
||||||
|
self.heading_styles: dict[str, int] = {}
|
||||||
|
self.number_formats: dict[str, str] = {}
|
||||||
|
self.image_names: dict[str, str] = {}
|
||||||
|
self.result = Result(self.output_file, self.images_dir)
|
||||||
|
|
||||||
|
def convert(self) -> Result:
|
||||||
|
"""验证目标、解析文档并写入 Markdown 与图片。"""
|
||||||
|
|
||||||
|
if self.output_dir.exists() and not self.force:
|
||||||
|
raise FileExistsError(f"输出目录已存在:{self.output_dir}")
|
||||||
|
self.output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
with zipfile.ZipFile(self.source) as archive:
|
||||||
|
names = set(archive.namelist())
|
||||||
|
if "word/document.xml" not in names:
|
||||||
|
raise ValueError("文件不是有效的 DOCX:缺少 word/document.xml")
|
||||||
|
|
||||||
|
self._load_relationships(archive, names)
|
||||||
|
self._load_styles(archive, names)
|
||||||
|
self._load_numbering(archive, names)
|
||||||
|
self._detect_unsupported_parts(names)
|
||||||
|
|
||||||
|
root = ET.fromstring(archive.read("word/document.xml"))
|
||||||
|
body = root.find("w:body", NS)
|
||||||
|
if body is None:
|
||||||
|
raise ValueError("文件不是有效的 DOCX:缺少正文节点")
|
||||||
|
|
||||||
|
blocks: list[str] = []
|
||||||
|
for child in body:
|
||||||
|
if child.tag == qname("w", "p"):
|
||||||
|
rendered = self._render_paragraph(child, archive)
|
||||||
|
if rendered:
|
||||||
|
blocks.append(rendered)
|
||||||
|
self.result.paragraphs += 1
|
||||||
|
elif child.tag == qname("w", "tbl"):
|
||||||
|
rendered = self._render_table(child, archive)
|
||||||
|
if rendered:
|
||||||
|
blocks.append(rendered)
|
||||||
|
self.result.tables += 1
|
||||||
|
|
||||||
|
markdown = "\n\n".join(blocks).strip() + "\n"
|
||||||
|
self.output_file.write_text(markdown, encoding="utf-8")
|
||||||
|
return self.result
|
||||||
|
|
||||||
|
def _load_relationships(self, archive: zipfile.ZipFile, names: set[str]) -> None:
|
||||||
|
"""读取超链接和媒体关系,后续按关系 ID 解析目标。"""
|
||||||
|
|
||||||
|
path = "word/_rels/document.xml.rels"
|
||||||
|
if path not in names:
|
||||||
|
return
|
||||||
|
root = ET.fromstring(archive.read(path))
|
||||||
|
for rel in root.findall("pr:Relationship", NS):
|
||||||
|
rel_id = rel.get("Id")
|
||||||
|
target = rel.get("Target")
|
||||||
|
rel_type = rel.get("Type", "").rsplit("/", 1)[-1]
|
||||||
|
if rel_id and target:
|
||||||
|
self.relationships[rel_id] = (rel_type, target)
|
||||||
|
|
||||||
|
def _load_styles(self, archive: zipfile.ZipFile, names: set[str]) -> None:
|
||||||
|
"""识别标题样式 ID;同时兼容英文 Heading 与中文标题名称。"""
|
||||||
|
|
||||||
|
if "word/styles.xml" not in names:
|
||||||
|
return
|
||||||
|
root = ET.fromstring(archive.read("word/styles.xml"))
|
||||||
|
for style in root.findall("w:style", NS):
|
||||||
|
if style.get(qname("w", "type")) != "paragraph":
|
||||||
|
continue
|
||||||
|
style_id = style.get(qname("w", "styleId"), "")
|
||||||
|
name_node = style.find("w:name", NS)
|
||||||
|
style_name = name_node.get(qname("w", "val"), "") if name_node is not None else ""
|
||||||
|
match = re.search(r"(?:heading|标题)\s*([1-6])", style_name, re.IGNORECASE)
|
||||||
|
if not match:
|
||||||
|
match = re.search(r"heading([1-6])", style_id, re.IGNORECASE)
|
||||||
|
if match:
|
||||||
|
self.heading_styles[style_id] = int(match.group(1))
|
||||||
|
|
||||||
|
def _load_numbering(self, archive: zipfile.ZipFile, names: set[str]) -> None:
|
||||||
|
"""建立编号实例到列表类型的基础映射。"""
|
||||||
|
|
||||||
|
if "word/numbering.xml" not in names:
|
||||||
|
return
|
||||||
|
root = ET.fromstring(archive.read("word/numbering.xml"))
|
||||||
|
abstract_formats: dict[str, str] = {}
|
||||||
|
for abstract in root.findall("w:abstractNum", NS):
|
||||||
|
abstract_id = abstract.get(qname("w", "abstractNumId"), "")
|
||||||
|
level = abstract.find("w:lvl", NS)
|
||||||
|
fmt = level.find("w:numFmt", NS) if level is not None else None
|
||||||
|
if fmt is not None:
|
||||||
|
abstract_formats[abstract_id] = fmt.get(qname("w", "val"), "bullet")
|
||||||
|
|
||||||
|
for number in root.findall("w:num", NS):
|
||||||
|
number_id = number.get(qname("w", "numId"), "")
|
||||||
|
abstract = number.find("w:abstractNumId", NS)
|
||||||
|
if abstract is not None:
|
||||||
|
abstract_id = abstract.get(qname("w", "val"), "")
|
||||||
|
self.number_formats[number_id] = abstract_formats.get(abstract_id, "bullet")
|
||||||
|
|
||||||
|
def _detect_unsupported_parts(self, names: set[str]) -> None:
|
||||||
|
"""对存在但当前不能可靠转换的部件给出明确警告。"""
|
||||||
|
|
||||||
|
checks = {
|
||||||
|
"word/comments.xml": "文档包含批注,当前转换器不会输出批注内容",
|
||||||
|
"word/footnotes.xml": "文档包含脚注,当前转换器不会输出脚注内容",
|
||||||
|
"word/endnotes.xml": "文档包含尾注,当前转换器不会输出尾注内容",
|
||||||
|
}
|
||||||
|
for path, warning in checks.items():
|
||||||
|
if path in names:
|
||||||
|
self.result.warnings.append(warning)
|
||||||
|
if any(name.startswith("word/header") for name in names):
|
||||||
|
self.result.warnings.append("文档包含页眉,当前转换器不会输出页眉内容")
|
||||||
|
if any(name.startswith("word/footer") for name in names):
|
||||||
|
self.result.warnings.append("文档包含页脚,当前转换器不会输出页脚内容")
|
||||||
|
if any(name.startswith("word/embeddings/") for name in names):
|
||||||
|
self.result.warnings.append("文档包含内嵌附件,当前转换器不会提取附件")
|
||||||
|
|
||||||
|
def _render_paragraph(self, paragraph: ET.Element, archive: zipfile.ZipFile) -> str:
|
||||||
|
"""转换段落、超链接、基础行内格式和段落内图片。"""
|
||||||
|
|
||||||
|
parts: list[str] = []
|
||||||
|
for child in paragraph:
|
||||||
|
if child.tag == qname("w", "r"):
|
||||||
|
parts.append(self._render_run(child, archive))
|
||||||
|
elif child.tag == qname("w", "hyperlink"):
|
||||||
|
text = "".join(self._render_run(run, archive) for run in child.findall("w:r", NS))
|
||||||
|
rel_id = child.get(qname("r", "id"))
|
||||||
|
relation = self.relationships.get(rel_id or "")
|
||||||
|
if text and relation and relation[0] == "hyperlink":
|
||||||
|
parts.append(f"[{text}]({relation[1]})")
|
||||||
|
else:
|
||||||
|
parts.append(text)
|
||||||
|
|
||||||
|
text = "".join(parts).strip()
|
||||||
|
if not text:
|
||||||
|
return ""
|
||||||
|
|
||||||
|
properties = paragraph.find("w:pPr", NS)
|
||||||
|
if properties is not None:
|
||||||
|
style = properties.find("w:pStyle", NS)
|
||||||
|
style_id = style.get(qname("w", "val"), "") if style is not None else ""
|
||||||
|
if style_id in self.heading_styles:
|
||||||
|
return f"{'#' * self.heading_styles[style_id]} {text}"
|
||||||
|
|
||||||
|
numbering = properties.find("w:numPr", NS)
|
||||||
|
if numbering is not None:
|
||||||
|
level = numbering.find("w:ilvl", NS)
|
||||||
|
number = numbering.find("w:numId", NS)
|
||||||
|
depth = int(level.get(qname("w", "val"), "0")) if level is not None else 0
|
||||||
|
number_id = number.get(qname("w", "val"), "") if number is not None else ""
|
||||||
|
marker = "-" if self.number_formats.get(number_id, "bullet") == "bullet" else "1."
|
||||||
|
return f"{' ' * depth}{marker} {text}"
|
||||||
|
|
||||||
|
return text
|
||||||
|
|
||||||
|
def _render_run(self, run: ET.Element, archive: zipfile.ZipFile) -> str:
|
||||||
|
"""转换单个文字区段,并在当前位置追加图片引用。"""
|
||||||
|
|
||||||
|
chunks: list[str] = []
|
||||||
|
for child in run:
|
||||||
|
if child.tag == qname("w", "t"):
|
||||||
|
chunks.append(child.text or "")
|
||||||
|
elif child.tag in {qname("w", "br"), qname("w", "cr")}:
|
||||||
|
chunks.append(" \n")
|
||||||
|
elif child.tag == qname("w", "tab"):
|
||||||
|
chunks.append(" ")
|
||||||
|
elif child.tag == qname("w", "drawing"):
|
||||||
|
chunks.extend(self._render_images(child, archive))
|
||||||
|
|
||||||
|
text = "".join(chunks)
|
||||||
|
if not text:
|
||||||
|
return ""
|
||||||
|
|
||||||
|
properties = run.find("w:rPr", NS)
|
||||||
|
if properties is not None:
|
||||||
|
if properties.find("w:b", NS) is not None:
|
||||||
|
text = f"**{text}**"
|
||||||
|
if properties.find("w:i", NS) is not None:
|
||||||
|
text = f"*{text}*"
|
||||||
|
return text
|
||||||
|
|
||||||
|
def _render_images(self, drawing: ET.Element, archive: zipfile.ZipFile) -> list[str]:
|
||||||
|
"""提取 drawing 关系指向的包内媒体,使用内容哈希稳定命名。"""
|
||||||
|
|
||||||
|
rendered: list[str] = []
|
||||||
|
for blip in drawing.findall(".//a:blip", NS):
|
||||||
|
rel_id = blip.get(qname("r", "embed"))
|
||||||
|
relation = self.relationships.get(rel_id or "")
|
||||||
|
if not relation or relation[0] != "image":
|
||||||
|
continue
|
||||||
|
target = PurePosixPath("word") / PurePosixPath(relation[1])
|
||||||
|
normalized = str(PurePosixPath(*[part for part in target.parts if part not in {".", ".."}]))
|
||||||
|
try:
|
||||||
|
data = archive.read(normalized)
|
||||||
|
except KeyError:
|
||||||
|
self.result.warnings.append(f"图片关系无法读取:{relation[1]}")
|
||||||
|
continue
|
||||||
|
if normalized not in self.image_names:
|
||||||
|
suffix = Path(normalized).suffix.lower() or ".bin"
|
||||||
|
filename = f"image-{hashlib.sha256(data).hexdigest()[:12]}{suffix}"
|
||||||
|
self.images_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
(self.images_dir / filename).write_bytes(data)
|
||||||
|
self.image_names[normalized] = filename
|
||||||
|
self.result.images += 1
|
||||||
|
rendered.append(f"")
|
||||||
|
return rendered
|
||||||
|
|
||||||
|
def _render_table(self, table: ET.Element, archive: zipfile.ZipFile) -> str:
|
||||||
|
"""将普通表格转换为 Markdown;复杂合并关系以警告提示。"""
|
||||||
|
|
||||||
|
rows: list[list[str]] = []
|
||||||
|
merged = False
|
||||||
|
for row in table.findall("w:tr", NS):
|
||||||
|
cells: list[str] = []
|
||||||
|
for cell in row.findall("w:tc", NS):
|
||||||
|
if cell.find(".//w:gridSpan", NS) is not None or cell.find(".//w:vMerge", NS) is not None:
|
||||||
|
merged = True
|
||||||
|
paragraphs = [self._render_paragraph(p, archive) for p in cell.findall("w:p", NS)]
|
||||||
|
content = "<br>".join(part for part in paragraphs if part)
|
||||||
|
cells.append(escape_markdown(content))
|
||||||
|
rows.append(cells)
|
||||||
|
|
||||||
|
if not rows:
|
||||||
|
return ""
|
||||||
|
width = max(len(row) for row in rows)
|
||||||
|
padded = [row + [""] * (width - len(row)) for row in rows]
|
||||||
|
header = padded[0]
|
||||||
|
lines = [f"| {' | '.join(header)} |", f"| {' | '.join(['---'] * width)} |"]
|
||||||
|
lines.extend(f"| {' | '.join(row)} |" for row in padded[1:])
|
||||||
|
if merged:
|
||||||
|
self.result.warnings.append("文档包含合并单元格,Markdown 表格可能无法保持原结构")
|
||||||
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
||||||
|
"""解析命令行参数。"""
|
||||||
|
|
||||||
|
parser = argparse.ArgumentParser(description="将 DOCX 的常见内容转换为 Markdown")
|
||||||
|
parser.add_argument("input", type=Path, help="输入 .docx 文件")
|
||||||
|
parser.add_argument("--output-dir", type=Path, help="输出目录,默认位于输入文件旁的同名目录")
|
||||||
|
parser.add_argument("--force", action="store_true", help="允许覆盖同名生成文件,但不删除其他文件")
|
||||||
|
return parser.parse_args(argv)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: list[str] | None = None) -> int:
|
||||||
|
"""命令行入口,失败时返回非零状态。"""
|
||||||
|
|
||||||
|
args = parse_args(argv)
|
||||||
|
source = args.input.resolve()
|
||||||
|
if not source.is_file():
|
||||||
|
print(f"错误:输入文件不存在:{source}", file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
if source.suffix.lower() != ".docx":
|
||||||
|
print("错误:只支持 .docx 文件", file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
|
||||||
|
output_dir = (args.output_dir or source.with_suffix("")).resolve()
|
||||||
|
try:
|
||||||
|
result = DocxConverter(source, output_dir, args.force).convert()
|
||||||
|
except (FileExistsError, ValueError, zipfile.BadZipFile, OSError) as error:
|
||||||
|
print(f"错误:{error}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
print(f"Markdown:{result.output_file}")
|
||||||
|
print(f"段落:{result.paragraphs},表格:{result.tables},图片:{result.images}")
|
||||||
|
if result.warnings:
|
||||||
|
print("警告:")
|
||||||
|
for warning in dict.fromkeys(result.warnings):
|
||||||
|
print(f"- {warning}")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
Reference in New Issue
Block a user