如何合并两个BeautifulSoup循环的数据并对应写入Excel?
解决BeautifulSoup提取数据后CSV每行对应手机号与公司名的问题
你当前代码的问题在于使用两个独立的for循环分别写入手机号和公司名,导致CSV文件中先列完所有手机号,再列所有公司名,无法实现数据的一一对应。要解决这个问题,需要将对应的手机号和公司名配对后,在同一个循环中完成单行写入。
基础解决方案(假设手机号与公司名数量完全匹配)
使用zip()函数将phones和names两个列表按位置配对,在同一个循环中同时提取对应的数据并写入CSV:
import csv import requests from bs4 import BeautifulSoup # 用with语句管理文件,自动处理文件关闭,避免资源泄漏 with open('/path/test.csv', "a", encoding="utf-8", newline='') as file: writer = csv.writer(file) writer.writerow(["Phone Number","Firm Name"]) response = requests.get(url) soup = BeautifulSoup(response.content, "lxml") phones = soup.findAll("div", attrs={"class":"PhonesBox"}) names = soup.findAll("h2", attrs={"class":"CompanyName"}) # 同时遍历配对后的手机号和公司名元素 for phone, name in zip(phones, names): try: # 用strip()清理文本前后的空白字符,让数据更整洁 firm_phone = phone.find("label").text.strip() firm_name = name.find("span").text.strip() # 一行写入一组对应数据 writer.writerow([firm_phone, firm_name]) except AttributeError: # 处理找不到label或span标签的情况,避免程序崩溃 writer.writerow(["", ""])
进阶解决方案(处理数据数量不一致的情况)
如果存在部分公司无手机号,或部分手机号无对应公司名的情况,可使用itertools.zip_longest以较长的列表为基准,缺失数据用默认值填充:
import csv import requests from bs4 import BeautifulSoup from itertools import zip_longest with open('/path/test.csv', "a", encoding="utf-8", newline='') as file: writer = csv.writer(file) writer.writerow(["Phone Number","Firm Name"]) response = requests.get(url) soup = BeautifulSoup(response.content, "lxml") phones = soup.findAll("div", attrs={"class":"PhonesBox"}) names = soup.findAll("h2", attrs={"class":"CompanyName"}) # 以较长的列表为遍历基准,缺失位置用None填充 for phone, name in zip_longest(phones, names, fillvalue=None): try: firm_phone = phone.find("label").text.strip() if phone else "" firm_name = name.find("span").text.strip() if name else "" writer.writerow([firm_phone, firm_name]) except AttributeError: writer.writerow(["", ""])
额外注意事项
- 将
try-except放在循环内部,确保某一组数据提取失败时,不会中断整个数据写入流程。 - 如果是首次写入CSV文件,建议使用
"w"模式(覆盖写入)替代"a"(追加写入),避免重复写入表头。
内容的提问来源于stack exchange,提问作者Doğan Topuz
相关产品推荐
相关产品推荐

