使用Sphinx与Crowdin国际化Python库文档:长段落拆分问题
解决Sphinx生成.pot文件中方法docstring合并的问题
要实现每个方法的docstring对应独立的msgid和空msgstr,必须编写自定义脚本——Sphinx没有内置工具处理这种拆分需求。下面是具体的实现思路和示例脚本:
核心思路
- 解析.pot文件,提取所有标记为方法docstring的注释行(
#:开头的行) - 从Python源码中提取对应方法的原始docstring内容
- 为每个方法生成独立的.pot条目,包含对应的
#:行、msgid(填充docstring)和空msgstr
示例脚本实现
步骤1:解析.pot文件提取方法信息
import re from pathlib import Path # 替换为你的.pot文件路径 pot_file = Path("./docs/_build/gettext/index.pot") pot_content = pot_file.read_text(encoding="utf-8") # 匹配#:行的正则,提取文件路径、类名+方法名、行号 method_pattern = re.compile(r'#: (.+):docstring of ([a-zA-Z0-9_.]+):(\d+)') method_list = [] for line in pot_content.splitlines(): line = line.strip() match = method_pattern.match(line) if match: file_path = match.group(1) full_method = match.group(2) # 拆分类名和方法名(支持多级模块类,如disnake.abc.GuildChannel.clone) class_name, method_name = full_method.rsplit('.', 1) method_list.append({ "source_line": line, "file_path": file_path, "class_name": class_name.split('.')[-1], # 取最后一级类名(适配嵌套类可调整) "method_name": method_name, "full_class": class_name })
步骤2:从源码提取方法docstring
用ast模块解析源码,避免导入模块时的依赖问题:
import ast def extract_docstring(source_file, target_class, target_method): """从Python文件中提取指定类的指定方法的docstring""" with open(source_file, encoding="utf-8") as f: tree = ast.parse(f.read(), filename=str(source_file)) # 遍历模块中的类 for node in ast.walk(tree): if isinstance(node, ast.ClassDef) and node.name == target_class: # 遍历类中的方法(包含普通方法和异步方法) for item in node.body: if isinstance(item, (ast.FunctionDef, ast.AsyncFunctionDef)) and item.name == target_method: if item.docstring: # 转义双引号,处理多行字符串 return item.docstring.replace('"', '\\"').replace('\n', '\\n') return ""
步骤3:生成独立的.pot条目
# 去重:避免同一方法多次处理 processed = set() new_pot_entries = [] for method in method_list: key = (method["file_path"], method["full_class"], method["method_name"]) if key in processed: continue processed.add(key) # 转换为绝对路径(根据你的项目结构调整) source_file = Path(method["file_path"]).resolve() if not source_file.exists(): print(f"跳过不存在的文件:{source_file}") continue docstring = extract_docstring(source_file, method["class_name"], method["method_name"]) if not docstring: print(f"{method['full_class']}.{method['method_name']} 无docstring,跳过") continue # 生成标准.pot条目 entry = f"""{method['source_line']} msgid "{docstring}" msgstr "" """ new_pot_entries.append(entry) # 写入新的.pot文件 new_pot_file = Path("./docs/_build/gettext/separated_docstrings.pot") new_pot_file.write_text("\n\n".join(new_pot_entries), encoding="utf-8") print(f"已生成独立条目.pot文件:{new_pot_file}")
注意事项
- 路径适配:脚本中
source_file = Path(method["file_path"]).resolve()需根据你的项目结构调整,确保能找到对应的Python源码文件 - 嵌套类处理:如果你的项目有嵌套类,需要修改
class_name的拆分逻辑,比如递归解析类的嵌套结构 - 特殊字符处理:docstring中的双引号、换行符必须转义,否则会破坏.pot文件格式
- 去重逻辑:同一方法可能在.pot中多次出现(比如被多个文档引用),必须去重避免重复条目
内容的提问来源于stack exchange,提问作者snipy7374
相关产品推荐
相关产品推荐

