如何提取列表中的手机号,匹配相邻手机号区间内的姓名生成Pandas表格
完整可运行代码
import pandas as pd nameBank = ["John Doe", "Jane Doe", "Patrick Star", "Spongebob Squarepants"] phoneList = [] nameList = [] list1 = ["1234567890", "John doe", "Not a NAME/USELESS FILLERINFO", "2345678901", "jane doe", "Not a NAME/USELESS FILLERINFO", "Not a NAME/USELESS FILLERINFO", "3456789012", "4567890123", "5678901234", "patrick star", "6789012345"] # 辅助判断是否为手机号(适配示例中10位纯数字的规则,可根据实际手机号规则调整) def is_phone_num(item): return isinstance(item, str) and item.isdigit() and len(item) == 10 # 第一步:收集所有手机号及其在list1中的索引 phone_positions = [] for idx, val in enumerate(list1): if is_phone_num(val): phone_positions.append((idx, val)) # 第二步:逐个手机号检索对应区间的匹配姓名 for i in range(len(phone_positions)): current_idx, current_phone = phone_positions[i] phoneList.append(current_phone) # 确定检索区间结束位置 end_idx = phone_positions[i+1][0] if i < len(phone_positions)-1 else len(list1) matched_name = "No Name Found" # 遍历当前手机号到下一个手机号之间的内容做匹配 for content in list1[current_idx+1 : end_idx]: content_lower = content.strip().lower() for official_name in nameBank: if official_name.lower() == content_lower: matched_name = official_name break if matched_name != "No Name Found": break nameList.append(matched_name) # 第三步:生成数据表导出 df = pd.DataFrame({'Phone Number': phoneList, 'Name': nameList}) df.to_csv('results.csv', index=False, encoding='utf-8') print(df)
实现逻辑说明
- 先遍历list1识别所有10位纯数字的手机号,同时记录每个手机号的索引位置,方便后续确定检索区间
- 遍历每个手机号时,检索范围为当前手机号后一位到下一个手机号前一位,最后一个手机号的检索范围为当前手机号后一位到list1末尾
- 姓名匹配不区分大小写,匹配到第一个符合的姓名就存入结果,未匹配则填入
No Name Found,保证手机号和姓名数量一一对应 - 最后用pandas生成数据表导出为csv文件,导出时不保留索引列
内容的提问来源于stack exchange,提问作者Sean
相关产品推荐
相关产品推荐

