You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取电话后无法写入CSV清洗文件的问题求助

问题解决:保留清洗后的电话数据

你已通过正则从谷歌搜索结果片段中提取电话,但数据清洗后生成的CSV仅保留邮箱,电话丢失。以下是修改代码的具体位置和方法:

1. 同步清洗电话字段

在清洗email的代码块后,添加对phone字段的清洗逻辑,和邮箱的清洗规则保持一致:

order_selection = data[['raw_url','email','phone']].drop_duplicates().dropna()

# 清洗email(原代码保留)
order_selection['email'] = order_selection['email'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex=True)

# 新增:同步清洗phone字段
order_selection['phone'] = order_selection['phone'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex=True)

order_selection = order_selection.drop_duplicates().dropna()

2. 分组时同时聚合电话数据

原分组代码仅对email做聚合,需要修改为同时聚合email和phone:

# GROUP BY PART 修改
groupby_order_selection = order_selection.groupby('raw_url').agg({
    'email': list,
    'phone': list  # 新增:按url分组聚合电话列表
}).reset_index()

3. 分组后同步清洗电话字段

在清洗分组后的email字段后,添加对phone的清洗:

# 清洗分组后的email(原代码保留)
groupby_order_selection['email'] = groupby_order_selection['email'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex = True)

# 新增:清洗分组后的phone字段
groupby_order_selection['phone'] = groupby_order_selection['phone'].astype(str).str.replace(r"^[][\s]*$|(^\[+|\]+$|')", lambda x: '' if x.group(1) else np.nan , regex = True)

groupby_order_selection = groupby_order_selection.drop_duplicates().dropna()

修改完成后,导出groupby_order_selection到CSV时,就会同时包含raw_url、email和phone三个字段。

内容的提问来源于stack exchange,提问作者djmystica

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 18:00:22