You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Sphinx与Crowdin国际化Python库文档:长段落拆分问题

解决Sphinx生成.pot文件中方法docstring合并的问题

要实现每个方法的docstring对应独立的msgid和空msgstr,必须编写自定义脚本——Sphinx没有内置工具处理这种拆分需求。下面是具体的实现思路和示例脚本:

核心思路

  1. 解析.pot文件,提取所有标记为方法docstring的注释行(#:开头的行)
  2. 从Python源码中提取对应方法的原始docstring内容
  3. 为每个方法生成独立的.pot条目,包含对应的#:行、msgid(填充docstring)和空msgstr

示例脚本实现

步骤1:解析.pot文件提取方法信息

import re
from pathlib import Path

# 替换为你的.pot文件路径
pot_file = Path("./docs/_build/gettext/index.pot")
pot_content = pot_file.read_text(encoding="utf-8")

# 匹配#:行的正则,提取文件路径、类名+方法名、行号
method_pattern = re.compile(r'#: (.+):docstring of ([a-zA-Z0-9_.]+):(\d+)')
method_list = []

for line in pot_content.splitlines():
    line = line.strip()
    match = method_pattern.match(line)
    if match:
        file_path = match.group(1)
        full_method = match.group(2)
        # 拆分类名和方法名(支持多级模块类,如disnake.abc.GuildChannel.clone)
        class_name, method_name = full_method.rsplit('.', 1)
        method_list.append({
            "source_line": line,
            "file_path": file_path,
            "class_name": class_name.split('.')[-1],  # 取最后一级类名(适配嵌套类可调整)
            "method_name": method_name,
            "full_class": class_name
        })

步骤2:从源码提取方法docstring

用ast模块解析源码,避免导入模块时的依赖问题:

import ast

def extract_docstring(source_file, target_class, target_method):
    """从Python文件中提取指定类的指定方法的docstring"""
    with open(source_file, encoding="utf-8") as f:
        tree = ast.parse(f.read(), filename=str(source_file))
    
    # 遍历模块中的类
    for node in ast.walk(tree):
        if isinstance(node, ast.ClassDef) and node.name == target_class:
            # 遍历类中的方法(包含普通方法和异步方法)
            for item in node.body:
                if isinstance(item, (ast.FunctionDef, ast.AsyncFunctionDef)) and item.name == target_method:
                    if item.docstring:
                        # 转义双引号,处理多行字符串
                        return item.docstring.replace('"', '\\"').replace('\n', '\\n')
    return ""

步骤3:生成独立的.pot条目

# 去重:避免同一方法多次处理
processed = set()
new_pot_entries = []

for method in method_list:
    key = (method["file_path"], method["full_class"], method["method_name"])
    if key in processed:
        continue
    processed.add(key)
    
    # 转换为绝对路径(根据你的项目结构调整)
    source_file = Path(method["file_path"]).resolve()
    if not source_file.exists():
        print(f"跳过不存在的文件:{source_file}")
        continue
    
    docstring = extract_docstring(source_file, method["class_name"], method["method_name"])
    if not docstring:
        print(f"{method['full_class']}.{method['method_name']} 无docstring,跳过")
        continue
    
    # 生成标准.pot条目
    entry = f"""{method['source_line']}
msgid "{docstring}"
msgstr ""
"""
    new_pot_entries.append(entry)

# 写入新的.pot文件
new_pot_file = Path("./docs/_build/gettext/separated_docstrings.pot")
new_pot_file.write_text("\n\n".join(new_pot_entries), encoding="utf-8")
print(f"已生成独立条目.pot文件:{new_pot_file}")

注意事项

  • 路径适配:脚本中source_file = Path(method["file_path"]).resolve()需根据你的项目结构调整,确保能找到对应的Python源码文件
  • 嵌套类处理:如果你的项目有嵌套类,需要修改class_name的拆分逻辑,比如递归解析类的嵌套结构
  • 特殊字符处理:docstring中的双引号、换行符必须转义,否则会破坏.pot文件格式
  • 去重逻辑:同一方法可能在.pot中多次出现(比如被多个文档引用),必须去重避免重复条目

内容的提问来源于stack exchange,提问作者snipy7374

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 16:37:34