如何从Wiktionary JSON Dump中提取个人姓名记录?
提取Wiktionary中姓名记录的解决方法
方法一:用JSON解析工具(推荐)
正则处理JSON极易出错,优先用编程语言的JSON解析库处理,步骤如下:
- 读取JSON文件,将其解析为字典/对象结构
- 遍历每条记录,检查记录中是否包含given name或surname标签(注意统一转小写判断,适配Wiktionary可能的大小写差异)
- 筛选符合条件的记录并保存
Python示例代码:
import json # 读取目标JSON文件 with open("wiktionary_proper_nouns.json", "r", encoding="utf-8") as f: data = json.load(f) # 筛选含姓名标签的记录 name_records = [] for record in data: # 根据实际JSON结构调整字段名,示例假设标签存在于"tags"字段 if "tags" in record: lower_tags = [tag.lower() for tag in record["tags"]] if "given name" in lower_tags or "surname" in lower_tags: name_records.append(record) # 保存筛选结果 with open("extracted_names.json", "w", encoding="utf-8") as f: json.dump(name_records, f, ensure_ascii=False, indent=2)
方法二:正则匹配(仅当无法用JSON解析时)
若必须用正则,先明确记录边界格式(比如每条记录以{"word":开头、},结尾),可参考以下思路:
- 匹配完整记录块:
\{[^}]*"given name"[^}]*\}|{[^}]*"surname"[^}]*\} - 需开启忽略大小写、匹配多行的正则参数,适配换行和大小写差异
Python正则示例代码:
import re import json # 读取文件内容 with open("wiktionary_proper_nouns.json", "r", encoding="utf-8") as f: content = f.read() # 匹配含目标标签的记录 pattern = re.compile(r'\{[^}]*("given name"|"surname")[^}]*\}', re.IGNORECASE | re.DOTALL) matches = pattern.findall(content) # 转换为JSON对象(若匹配结果存在语法问题,需手动调整正则) name_records = [json.loads(match) for match in matches] # 保存结果 with open("extracted_names.json", "w", encoding="utf-8") as f: json.dump(name_records, f, ensure_ascii=False, indent=2)
注意事项
- 优先使用JSON解析库,正则处理JSON容易因嵌套对象、转义引号等格式问题出错
- 需确认Wiktionary JSON的具体结构,比如标签字段的实际名称(可能为"pos"或"categories"),调整代码中的字段判断逻辑
- 测试时先用小样本验证筛选逻辑,确保结果准确
内容的提问来源于stack exchange,提问作者Ne Mo
相关产品推荐
相关产品推荐

