Python中循环执行Selenium命令处理CSV多URL的问题求助
Twitter数据爬取CSV导出空文件问题排查与修复
我是Python新手,尝试用Selenium爬取Twitter数据:将多个URL保存于CSV文件,期望代码逐个访问这些URL,滚动页面并爬取推文的回复对象、文本、发布日期,最终将所有数据保存到CSV。目前Selenium爬虫逻辑和循环逻辑单独运行均正常,但结合后导出的CSV始终为空,恳请帮忙排查代码问题!
原代码:
#Do imports import csv import time import selenium import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait as wait from selenium.webdriver.common.action_chains import ActionChains import time driver = webdriver.Chrome(executable_path=r"/chromedriver") tweets = [] with open('BKQuotedTweetsURL.csv', 'rt') as BK_csv: BK_url = csv.reader(BK_csv) for row in BK_url: links = row[0] tweets.append(links) #link should be something like "https://.com" for link in tweets: driver.get(link) time.sleep(10) # Get scroll height after first time page load last_height = driver.execute_script("return document.body.scrollHeight") last_elem='' current_elem='' while True: # Scroll down to bottom driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait to load page time.sleep(5) # Calculate new scroll height and compare with last scroll height new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height #update all_tweets to keep loop all_tweets = driver.find_elements(By.XPATH, '//div[@data-testid]//article[@data-testid="tweet"]') for item in all_tweets[1:]: # skip tweet already scrapped print('--- date ---') try: date = item.find_element(By.XPATH, './/time').text except: date = '[empty]' print(date) print('--- text ---') try: text = item.find_element(By.XPATH, './/div[@data-testid="tweetText"]').text except: text = '[empty]' print(text) print('--- replying_to ---') try: replying_to = item.find_element(By.XPATH, './/div[contains(text(), "Replying to")]//a').text except: replying_to = '[empty]' print(replying_to) #Append new tweets replies to tweet array tweets.append([replying_to, text, date]) if (last_elem == current_elem): result = True else: last_elem = current_elem df = pd.DataFrame(tweets, columns=['Replying to', 'Tweet', 'Date of Tweet']) df.to_csv(r'BKURLListComm.csv', index=False, encoding='utf-8') #save a csv file in the downloads folder, change it to your structure and desired folder
核心问题点
变量复用导致数据结构混乱:
一开始用tweets数组存储从CSV读取的URL字符串,后续又往里面追加推文的列表数据([replying_to, text, date]),最终数组里既有字符串又有列表,生成DataFrame时结构完全错误,导致导出的CSV为空或格式异常。滚动与元素获取顺序错误:
代码先滚动到底部再获取推文元素,初始页面加载的推文会被跳过;同时all_tweets[1:]的切片逻辑没有依据,无法正确跳过已爬取的内容,导致重复爬取或漏爬。无效的循环终止判断:
last_elem和current_elem的逻辑完全没有关联到推文元素,无法判断是否已经加载完所有内容,可能提前终止循环或进入死循环。WebDriver初始化方式过时:
executable_path参数已被Selenium弃用,容易引发驱动加载失败的隐性错误。
修复后的代码
import csv import time import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service # 初始化Chrome驱动(使用Service类,适配新版Selenium) service = Service(r"/chromedriver") driver = webdriver.Chrome(service=service) # 分离URL列表和推文数据列表,避免冲突 url_list = [] tweets_data = [] # 读取CSV中的URL with open('BKQuotedTweetsURL.csv', 'r', encoding='utf-8') as BK_csv: BK_url = csv.reader(BK_csv) for row in BK_url: # 跳过空行 if row[0].strip(): url_list.append(row[0]) # 遍历每个URL爬取数据 for link in url_list: driver.get(link) # 等待页面加载完成(可改用WebDriverWait优化) time.sleep(10) last_height = driver.execute_script("return document.body.scrollHeight") # 记录已爬取的推文数量,避免重复爬取 crawled_count = 0 while True: # 先获取当前页面的所有推文,再滚动加载 all_tweets = driver.find_elements(By.XPATH, '//div[@data-testid]//article[@data-testid="tweet"]') # 只处理新增的推文 for item in all_tweets[crawled_count:]: print('--- date ---') try: date = item.find_element(By.XPATH, './/time').text except: date = '[empty]' print(date) print('--- text ---') try: text = item.find_element(By.XPATH, './/div[@data-testid="tweetText"]').text except: text = '[empty]' print(text) print('--- replying_to ---') try: replying_to = item.find_element(By.XPATH, './/div[contains(text(), "Replying to")]//a').text except: replying_to = '[empty]' print(replying_to) # 将数据存入专门的推文列表 tweets_data.append([replying_to, text, date]) # 更新已爬取数量 crawled_count = len(all_tweets) # 滚动加载更多内容 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(5) new_height = driver.execute_script("return document.body.scrollHeight") # 滚动到底部且没有新内容时退出循环 if new_height == last_height: break last_height = new_height # 生成DataFrame并导出CSV df = pd.DataFrame(tweets_data, columns=['Replying to', 'Tweet', 'Date of Tweet']) df.to_csv(r'BKURLListComm.csv', index=False, encoding='utf-8') # 关闭驱动 driver.quit()
关键修改说明
- 分离变量:用
url_list存储待爬取的URL,tweets_data存储爬取到的推文数据,彻底避免数据结构冲突。 - 调整爬取顺序:先获取当前页面的推文再滚动,确保初始加载的内容不会被遗漏;用
crawled_count记录已爬取的推文数量,只处理新增内容。 - 修复驱动初始化:改用
Service类配置Chrome驱动路径,适配新版Selenium的API规范。 - 优化循环终止逻辑:基于滚动高度判断是否加载完所有内容,同时结合已爬取数量确保数据完整。
内容的提问来源于stack exchange,提问作者Zabina
相关产品推荐
相关产品推荐

