You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中循环执行Selenium命令处理CSV多URL的问题求助

Twitter数据爬取CSV导出空文件问题排查与修复

我是Python新手,尝试用Selenium爬取Twitter数据:将多个URL保存于CSV文件,期望代码逐个访问这些URL,滚动页面并爬取推文的回复对象、文本、发布日期,最终将所有数据保存到CSV。目前Selenium爬虫逻辑和循环逻辑单独运行均正常,但结合后导出的CSV始终为空,恳请帮忙排查代码问题!

原代码:

#Do imports
import csv 
import time
import selenium
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait as wait 
from selenium.webdriver.common.action_chains import ActionChains
import time

driver = webdriver.Chrome(executable_path=r"/chromedriver")

tweets = []

with open('BKQuotedTweetsURL.csv', 'rt') as BK_csv:
    BK_url = csv.reader(BK_csv)
    for row in BK_url:
        links = row[0]
        tweets.append(links)

#link should be something like "https://.com"
for link in tweets:
    driver.get(link)
    time.sleep(10)
            
    # Get scroll height after first time page load
    last_height = driver.execute_script("return document.body.scrollHeight")

    last_elem=''
    current_elem=''

    while True:
            
        # Scroll down to bottom
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # Wait to load page
        time.sleep(5)
        # Calculate new scroll height and compare with last scroll height
        new_height = driver.execute_script("return document.body.scrollHeight")
        if new_height == last_height:
           break
        last_height = new_height
            
            
        #update all_tweets to keep loop
        all_tweets = driver.find_elements(By.XPATH, '//div[@data-testid]//article[@data-testid="tweet"]')

        for item in all_tweets[1:]: # skip tweet already scrapped

            print('--- date ---')
            try:
                date = item.find_element(By.XPATH, './/time').text
            except:
                date = '[empty]'
            print(date)

            print('--- text ---')
            try:
                text = item.find_element(By.XPATH, './/div[@data-testid="tweetText"]').text
            except:
                text = '[empty]'
            print(text)
            
            print('--- replying_to ---')
            try:
                replying_to = item.find_element(By.XPATH, './/div[contains(text(), "Replying to")]//a').text
            except:
                replying_to = '[empty]'
            print(replying_to)
            
            #Append new tweets replies to tweet array
            tweets.append([replying_to, text, date])
                       
            if (last_elem == current_elem):
                result = True
            else:
                last_elem = current_elem


df = pd.DataFrame(tweets, columns=['Replying to', 'Tweet', 'Date of Tweet'])
df.to_csv(r'BKURLListComm.csv', index=False, encoding='utf-8') #save a csv file in the downloads folder, change it to your structure and desired folder

核心问题点

  1. 变量复用导致数据结构混乱:
    一开始用tweets数组存储从CSV读取的URL字符串,后续又往里面追加推文的列表数据([replying_to, text, date]),最终数组里既有字符串又有列表,生成DataFrame时结构完全错误,导致导出的CSV为空或格式异常。

  2. 滚动与元素获取顺序错误:
    代码先滚动到底部再获取推文元素,初始页面加载的推文会被跳过;同时all_tweets[1:]的切片逻辑没有依据,无法正确跳过已爬取的内容,导致重复爬取或漏爬。

  3. 无效的循环终止判断:
    last_elem和current_elem的逻辑完全没有关联到推文元素,无法判断是否已经加载完所有内容,可能提前终止循环或进入死循环。

  4. WebDriver初始化方式过时:
    executable_path参数已被Selenium弃用,容易引发驱动加载失败的隐性错误。


修复后的代码

import csv 
import time
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service

# 初始化Chrome驱动(使用Service类,适配新版Selenium)
service = Service(r"/chromedriver")
driver = webdriver.Chrome(service=service)

# 分离URL列表和推文数据列表,避免冲突
url_list = []
tweets_data = []

# 读取CSV中的URL
with open('BKQuotedTweetsURL.csv', 'r', encoding='utf-8') as BK_csv:
    BK_url = csv.reader(BK_csv)
    for row in BK_url:
        # 跳过空行
        if row[0].strip():
            url_list.append(row[0])

# 遍历每个URL爬取数据
for link in url_list:
    driver.get(link)
    # 等待页面加载完成(可改用WebDriverWait优化)
    time.sleep(10)
    
    last_height = driver.execute_script("return document.body.scrollHeight")
    # 记录已爬取的推文数量,避免重复爬取
    crawled_count = 0
    
    while True:
        # 先获取当前页面的所有推文,再滚动加载
        all_tweets = driver.find_elements(By.XPATH, '//div[@data-testid]//article[@data-testid="tweet"]')
        
        # 只处理新增的推文
        for item in all_tweets[crawled_count:]:
            print('--- date ---')
            try:
                date = item.find_element(By.XPATH, './/time').text
            except:
                date = '[empty]'
            print(date)

            print('--- text ---')
            try:
                text = item.find_element(By.XPATH, './/div[@data-testid="tweetText"]').text
            except:
                text = '[empty]'
            print(text)
            
            print('--- replying_to ---')
            try:
                replying_to = item.find_element(By.XPATH, './/div[contains(text(), "Replying to")]//a').text
            except:
                replying_to = '[empty]'
            print(replying_to)
            
            # 将数据存入专门的推文列表
            tweets_data.append([replying_to, text, date])
        
        # 更新已爬取数量
        crawled_count = len(all_tweets)
        
        # 滚动加载更多内容
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(5)
        
        new_height = driver.execute_script("return document.body.scrollHeight")
        # 滚动到底部且没有新内容时退出循环
        if new_height == last_height:
            break
        last_height = new_height

# 生成DataFrame并导出CSV
df = pd.DataFrame(tweets_data, columns=['Replying to', 'Tweet', 'Date of Tweet'])
df.to_csv(r'BKURLListComm.csv', index=False, encoding='utf-8')

# 关闭驱动
driver.quit()

关键修改说明

  • 分离变量:用url_list存储待爬取的URL,tweets_data存储爬取到的推文数据,彻底避免数据结构冲突。
  • 调整爬取顺序:先获取当前页面的推文再滚动,确保初始加载的内容不会被遗漏;用crawled_count记录已爬取的推文数量,只处理新增内容。
  • 修复驱动初始化:改用Service类配置Chrome驱动路径,适配新版Selenium的API规范。
  • 优化循环终止逻辑:基于滚动高度判断是否加载完所有内容,同时结合已爬取数量确保数据完整。

内容的提问来源于stack exchange,提问作者Zabina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 11:40:24