You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Wiktionary JSON Dump中提取个人姓名记录?

提取Wiktionary中姓名记录的解决方法

方法一:用JSON解析工具(推荐)

正则处理JSON极易出错,优先用编程语言的JSON解析库处理,步骤如下:

  • 读取JSON文件,将其解析为字典/对象结构
  • 遍历每条记录,检查记录中是否包含given name或surname标签(注意统一转小写判断,适配Wiktionary可能的大小写差异)
  • 筛选符合条件的记录并保存

Python示例代码:

import json

# 读取目标JSON文件
with open("wiktionary_proper_nouns.json", "r", encoding="utf-8") as f:
    data = json.load(f)

# 筛选含姓名标签的记录
name_records = []
for record in data:
    # 根据实际JSON结构调整字段名,示例假设标签存在于"tags"字段
    if "tags" in record:
        lower_tags = [tag.lower() for tag in record["tags"]]
        if "given name" in lower_tags or "surname" in lower_tags:
            name_records.append(record)

# 保存筛选结果
with open("extracted_names.json", "w", encoding="utf-8") as f:
    json.dump(name_records, f, ensure_ascii=False, indent=2)

方法二:正则匹配(仅当无法用JSON解析时)

若必须用正则,先明确记录边界格式(比如每条记录以{"word":开头、},结尾),可参考以下思路:

  • 匹配完整记录块:\{[^}]*"given name"[^}]*\}|{[^}]*"surname"[^}]*\}
  • 需开启忽略大小写、匹配多行的正则参数,适配换行和大小写差异

Python正则示例代码:

import re
import json

# 读取文件内容
with open("wiktionary_proper_nouns.json", "r", encoding="utf-8") as f:
    content = f.read()

# 匹配含目标标签的记录
pattern = re.compile(r'\{[^}]*("given name"|"surname")[^}]*\}', re.IGNORECASE | re.DOTALL)
matches = pattern.findall(content)

# 转换为JSON对象(若匹配结果存在语法问题,需手动调整正则)
name_records = [json.loads(match) for match in matches]

# 保存结果
with open("extracted_names.json", "w", encoding="utf-8") as f:
    json.dump(name_records, f, ensure_ascii=False, indent=2)

注意事项

  • 优先使用JSON解析库,正则处理JSON容易因嵌套对象、转义引号等格式问题出错
  • 需确认Wiktionary JSON的具体结构,比如标签字段的实际名称(可能为"pos"或"categories"),调整代码中的字段判断逻辑
  • 测试时先用小样本验证筛选逻辑,确保结果准确

内容的提问来源于stack exchange,提问作者Ne Mo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 03:32:47