Python正则表达式解析多行长msgid与msgstr内容至字典
解析PO文件中跨多行的msgid与msgstr内容
需求
需要解析大型PO格式翻译文件,提取msgid标记的源文本和对应的msgstr目标文本,完全保留所有空格、缩进及换行格式,最终将结果存入字典以便输出至表格。现有代码无法处理跨多行的文本内容,需用Python正则表达式实现。
示例PO文件内容
#: superset-frontend/src/explore/components/controls/DndColumnSelectControl/Option.tsx:68 #: superset-frontend/src/explore/components/controls/OptionControls/index.tsx:323 msgid "" "\n" " This filter was inherited from the dashboard's context.\n" " It won't be saved when saving the chart.\n" " " msgstr "" "\n" " Фильтр был наследован из контекста дашборда.\n" " Это не будет сохранено при сохранении графика.\n" " " #: superset/tasks/schedules.py:184 #, python-format msgid "" "\n" " <b><a href=\"%(url)s\">Explore in Superset</a></b><p></p>\n" " <img src=\"cid:%(msgid)s\">\n" " " msgstr "" "\n" " <b><a href=“%(url)s”>Исследовать в Superset</a></b><p></p>\n" " <img src=“cid:%(msgid)s”>\n" " " #: superset/reports/notifications/email.py:60
用户未完成代码
def Parse(file : io.TextIOWrapper): text = "" for lines in file: if lines.startswith("msgid"): text += f" {lines.strip()}" elif lines.startswith("msgstr"): text += f" {lines.strip()}" else: text += lines.strip() text.split() sourcetxt = {} for index, word in enumerate(text): if word = "msgid":
解决方案
核心思路
利用正则表达式匹配PO文件中完整的msgid-msgstr块,通过非贪婪模式捕获跨多行的文本内容,再用ast.literal_eval解析转义字符(如\n),还原原始格式。
完整代码
import re import ast from typing import Dict def parse_po_file(file_path: str) -> Dict[str, str]: # 读取整个文件内容 with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 正则匹配所有msgid-msgstr块:(?s)让.匹配换行符 pattern = r'(?s)msgid\s*""\s*(.*?)\s*msgstr\s*""\s*(.*?)(?=\n#:|\Z)' matches = re.findall(pattern, content) translation_dict = {} for msgid_parts, msgstr_parts in matches: # 处理msgid内容:分割成带引号的行,解析转义字符后拼接 msgid_lines = [line.strip().strip('"') for line in msgid_parts.split('\n') if line.strip().strip('"')] msgid_content = ''.join([ast.literal_eval(f'"{line}"') for line in msgid_lines]) # 处理msgstr内容:同上 msgstr_lines = [line.strip().strip('"') for line in msgstr_parts.split('\n') if line.strip().strip('"')] msgstr_content = ''.join([ast.literal_eval(f'"{line}"') for line in msgstr_lines]) translation_dict[msgid_content] = msgstr_content return translation_dict # 使用示例 if __name__ == "__main__": result = parse_po_file("your_translation_file.po") # 输出验证 for src, target in result.items(): print("=== 源文本 ===") print(repr(src)) # repr()可以显示换行、空格等不可见字符 print("=== 目标文本 ===") print(repr(target)) print("\n")
代码说明
- 正则模式:
(?s)启用单行模式,使.匹配换行;(.*?)非贪婪捕获msgid和msgstr之间的所有内容;(?=\n#:|\Z)指定终止条件为下一个#:行或文件结尾。 - 转义解析:用
ast.literal_eval将PO文件中的字符串转义(如"\n"转为实际换行符),确保完全保留原始格式。 - 字典存储:最终以
{源文本: 目标文本}的结构存储,方便后续导出至表格。
内容的提问来源于stack exchange,提问作者Nikita
相关产品推荐
相关产品推荐

