You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python与正则表达式提取并分组文件中的文化姓氏

问题描述

手上有大量结构相似的文件,每个文件包含多组文化对应的姓氏集合,单个文件的姓氏条目可达数千条。手动整理效率太低,希望用Python结合正则表达式自动化提取姓氏并按文化分组。

输入文件内容示例:

###Myrman###
360 = { # DUPLICATE §§§§§§
    name="of Myr"
    culture = myrman
}
300507 = {
    name = "of Myr"
    culture = myrman
}
300525 = {
    name = "Trellos"
    culture = myrman
}
300534 = {
    name = "Uteuran"
    culture = myrman
}

##Lysene##
1386 = {
    name="Ormollen"
    culture = lysene
    coat_of_arms = {
        template = 0
        layer = {
            texture = 14
            texture_internal = 9
            emblem = 0
            color = 0
            color = 0
            color = 0
        }
    }
}
300505 = {
    name = "of Lys"
    culture = lysene
}
300523 = {
    name = "Lohar"
    culture = lysene
}
300532 = {
    name = "Assadyrn"
    culture = lysene
}

期望输出格式:

Myrman: ["of Myr", "of Myr", "Trellos", "Uteuran"]

Lysene: ["Ormollen", "of Lys", "Lohar", "Assadyrn"]
实现方案

核心思路是先匹配出每个文化的名称,再提取该文化对应内容块内的所有name字段值,最后按文化分组输出。以下是具体代码:

import re

def extract_surnames(file_content):
    # 匹配文化标题(兼容##或###包裹的格式)
    culture_pattern = re.compile(r'#{2,3}(\w+)#{2,3}')
    # 匹配name字段值,兼容=前后有无空格的情况
    name_pattern = re.compile(r'name\s*=\s*"([^"]+)"')
    
    # 获取所有文化标题的匹配结果
    culture_matches = list(culture_pattern.finditer(file_content))
    surnames_map = {}
    
    for idx, match in enumerate(culture_matches):
        culture_name = match.group(1)
        # 确定当前文化块的结束位置:下一个文化标题的起始,或文件末尾
        end_pos = culture_matches[idx+1].start() if idx < len(culture_matches)-1 else len(file_content)
        # 截取当前文化对应的内容块
        culture_block = file_content[match.end():end_pos]
        # 提取该块内所有name值
        surnames_map[culture_name] = name_pattern.findall(culture_block)
    
    return surnames_map

# 读取目标文件(实际使用时替换为你的文件路径)
with open('surnames_file.txt', 'r', encoding='utf-8') as f:
    content = f.read()

# 提取并按格式输出
result = extract_surnames(content)
for culture, names in result.items():
    names_str = ', '.join(f'"{name}"' for name in names)
    print(f'{culture}: [{names_str}]\n')

代码说明

  1. 文化标题匹配:正则#{2,3}(\w+)#{2,3}可以同时匹配##XXX##和###XXX###两种格式的文化标题,提取出XXX作为文化名称;
  2. 姓氏提取:正则name\s*=\s*"([^"]+)"匹配所有name="XXX"或name = "XXX"格式的字段,提取引号内的姓氏内容;
  3. 内容块分割:通过遍历文化标题的匹配位置,分割出每个文化对应的独立内容块,避免跨文化提取错误;
  4. 格式输出:将提取到的姓氏列表转换为要求的字符串格式,逐组输出。

内容的提问来源于stack exchange,提问作者Keenonthedaywalk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 01:35:16