Python中如何通过for循环将两类数据写入CSV的两个独立列
问题原因
你写入CSV出现内容覆盖、无法分两列存储的核心原因有两个:
- 没有暂存全量爬取数据,单次循环写入时默认使用覆盖模式,导致之前的内容被新内容替换
- 没有建立姓名和手机号的一一对应关系,按行写入时无法对应到两列结构
解决方法
你可以采用「暂存数据到字典列表→全部爬取完成后统一导出CSV」的方案实现需求,修改后的完整代码如下,新增内容都加了注释标注:
import json from io import StringIO from bs4 import BeautifulSoup from requests_html import HTMLSession import time from selenium import webdriver import requests import pandas as pd import numpy as np from selenium.webdriver.chrome.options import Options import colorama from colorama import Fore, Back, Style colorama.init(autoreset = False) import selenium from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.action_chains import ActionChains PATH = "C:\Program Files (x86)\chromedriver.exe" chrome_options = Options() chrome_options.add_experimental_option('excludeSwitches', ['enable-logging']) driver = webdriver.Chrome(PATH, options = chrome_options) driver.minimize_window() # ========== 新增:初始化空列表存储所有数据 ========== data_list = [] for b in range(500): driver.implicitly_wait(10) url = "https://www.healthgrades.com/usearch?what=Marriage%20%26%20Family%20Therapy&entityCode=PS303&where=CA&pageNum={}&sort.provider=bestmatch&state=CA".format(b+104) driver.get(url) time.sleep(10) length = len(driver.find_elements_by_xpath("//a[@data-qa-target='provider-details-provider-name']")) for i in range(length): elements = driver.find_elements_by_xpath("//a[@data-qa-target='provider-details-provider-name']") elements[i].click() handles = driver.window_handles driver.switch_to.window(handles[1]) time.sleep(1) # ========== 修改:获取值后存入列表,不再直接打印 ========== name = driver.find_element_by_tag_name("h1").get_attribute("innerText").strip() phone = driver.find_element_by_xpath("//div[@class='summary-standard-button-row-mobile']/a").get_attribute("innerText").strip() # 打印保持和之前一致的效果,可自行删除 print(name) print(phone) # 把单条数据加入总列表 data_list.append({"姓名": name, "联系电话": phone}) driver.close() driver.switch_to.window(handles[0]) time.sleep(1) # ========== 新增:所有数据爬取完成后统一导出CSV ========== df = pd.DataFrame(data_list) # encoding用utf_8_sig是为了避免Windows下打开CSV出现中文乱码 df.to_csv("加州婚姻家庭治疗师联系方式.csv", index=False, encoding="utf_8_sig")
补充说明
- 如果你担心爬取中途程序崩溃丢失已爬数据,可以改用追加模式逐行写入,每次写入时指定
mode='a'和header=False,仅第一次写入时保留表头即可 - 建议给获取姓名、电话的代码段加上
try-except异常捕获,避免个别页面元素缺失导致整个爬虫中断 - 高版本Selenium已经废弃了
find_elements_by_xpath这类写法,如果运行报错可以替换为driver.find_elements(By.XPATH, "对应路径")的格式
内容的提问来源于stack exchange,提问作者Adrian Reichert
相关产品推荐
相关产品推荐

