Python提取电话后无法写入CSV清洗文件的问题求助
问题解决:保留清洗后的电话数据
你已通过正则从谷歌搜索结果片段中提取电话,但数据清洗后生成的CSV仅保留邮箱,电话丢失。以下是修改代码的具体位置和方法:
1. 同步清洗电话字段
在清洗email的代码块后,添加对phone字段的清洗逻辑,和邮箱的清洗规则保持一致:
order_selection = data[['raw_url','email','phone']].drop_duplicates().dropna() # 清洗email(原代码保留) order_selection['email'] = order_selection['email'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex=True) # 新增:同步清洗phone字段 order_selection['phone'] = order_selection['phone'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex=True) order_selection = order_selection.drop_duplicates().dropna()
2. 分组时同时聚合电话数据
原分组代码仅对email做聚合,需要修改为同时聚合email和phone:
# GROUP BY PART 修改 groupby_order_selection = order_selection.groupby('raw_url').agg({ 'email': list, 'phone': list # 新增:按url分组聚合电话列表 }).reset_index()
3. 分组后同步清洗电话字段
在清洗分组后的email字段后,添加对phone的清洗:
# 清洗分组后的email(原代码保留) groupby_order_selection['email'] = groupby_order_selection['email'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex = True) # 新增:清洗分组后的phone字段 groupby_order_selection['phone'] = groupby_order_selection['phone'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex = True) groupby_order_selection = groupby_order_selection.drop_duplicates().dropna()
修改完成后,导出groupby_order_selection到CSV时,就会同时包含raw_url、email和phone三个字段。
内容的提问来源于stack exchange,提问作者djmystica
相关产品推荐
相关产品推荐

