You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium提取Twitter推文时页面加载失败的解决方案

Python Selenium爬取Twitter推文时频繁出现“Can't load page”错误,如何稳定执行?

使用Python Selenium提取Twitter推文时,运行一段时间后会出现“Can't load page”报错。已经尝试在提取后添加time.sleep(),但执行N次后仍然会触发该问题,希望实现稳定爬取N次的效果。原代码如下:

driver = webdriver.Chrome()

driver.get(URL)

while(counter != tweets_count):
        # Get tweets
        for i in range(1,6):
            context = driver.find_element(By.XPATH, "//*[@id='react-root']/div/div/div[2]/main/div/div/div/div[1]/div/div[3]/section/div/div/div[" + str(i) + "]")
            context_list = context.text.split("\n")
            context_list = context_list[4:-5]
            temp = ""

            for j in range(len(context_list)):
                temp += context_list[j]

            tweets.append(temp)
            
            time.sleep(3)
        
        # Scroll down to bottom
        driver.execute_script("window.scrollTo(" + str(first_height) + ", " + str(increase_amount) + ");")
        
        first_height = increase_amount
        increase_amount = increase_amount+increase_amount
        
        # Waiting until page loading completely
        time.sleep(7)
        
        counter += 1
        print(counter)

问题根源分析

Twitter有严格的反爬机制,你的代码存在几个触发反爬的核心问题:

  • 固定翻倍的滚动方式完全不符合真实用户行为,极易被检测
  • 硬编码的XPATH依赖静态页面结构,不仅容易失效,频繁定位也会增加爬虫特征暴露风险
  • 固定时长的time.sleep()无法适配页面真实加载速度,要么等待不足导致元素未加载,要么等待过久触发异常检测
  • 未做任何自动化特征隐藏,Selenium的默认标识会被Twitter直接识别

具体修复方案

1. 替换固定增量滚动为渐进式滚动

真实用户不会每次滚动翻倍距离,改为滚动到当前页面底部,或每次滚动小段距离:

# 滚动到页面底部(推荐)
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")

# 或者每次滚动500px,更贴近用户浏览习惯
# driver.execute_script("window.scrollBy(0, 500);")

2. 使用稳定的特征选择器替代硬编码XPATH

Twitter页面结构动态变化,改用基于官方测试标识的选择器,比如data-testid="tweet":

# 定位所有推文元素
tweet_elements = driver.find_elements(By.CSS_SELECTOR, '[data-testid="tweet"]')
for tweet in tweet_elements[-5:]:  # 只取最新的5条,避免重复提取
    try:
        # 提取推文正文
        content = tweet.find_element(By.CSS_SELECTOR, '[data-testid="tweetText"]').text
        tweets.append(content)
    except Exception:
        continue  # 跳过无正文的推文

3. 用显式等待替代固定sleep

显式等待会在元素加载完成后再执行操作,既提升效率又避免触发反爬:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

wait = WebDriverWait(driver, 15)

# 等待推文加载完成
tweet_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, '[data-testid="tweet"]')))

# 滚动后等待新推文加载
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
wait.until(EC.staleness_of(tweet_elements[-1]))  # 等待旧元素失效,新元素加载

4. 添加自动化特征隐藏

修改Chrome配置,隐藏Selenium的自动化标识:

from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
options.add_argument("--start-maximized")  # 模拟最大化窗口
options.add_argument("--disable-infobars")

driver = webdriver.Chrome(options=options)
# 移除navigator.webdriver标识
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

5. 增加加载失败重试机制

当检测到“Can't load page”时,自动刷新页面并重试当前步骤:

while counter != tweets_count:
    try:
        # 等待页面正常加载
        tweet_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, '[data-testid="tweet"]')))
        
        # 提取推文逻辑...
        
        # 滚动逻辑...
        
        counter += 1
        print(f"已完成第 {counter} 次爬取")
    except Exception:
        if "Can't load page" in driver.page_source:
            print("页面加载失败,刷新重试...")
            driver.refresh()
            time.sleep(3)  # 刷新后短暂等待
        else:
            break

修改后的完整示例代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options
import time

options = Options()
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
options.add_argument("--start-maximized")
options.add_argument("--disable-infobars")

driver = webdriver.Chrome(options=options)
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

URL = "你的Twitter目标URL"
driver.get(URL)

wait = WebDriverWait(driver, 15)
tweets = []
counter = 0
tweets_count = 10  # 自定义需要的爬取次数

while counter != tweets_count:
    try:
        # 等待推文加载完成
        tweet_elements = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, '[data-testid="tweet"]')))
        
        # 提取最新5条推文,避免重复
        for tweet in tweet_elements[-5:]:
            try:
                content = tweet.find_element(By.CSS_SELECTOR, '[data-testid="tweetText"]').text
                tweets.append(content)
                time.sleep(1)  # 模拟用户阅读间隔
            except Exception:
                continue
        
        # 滚动到页面底部
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # 等待新推文加载
        wait.until(EC.staleness_of(tweet_elements[-1]))
        
        counter += 1
        print(f"已完成第 {counter} 次爬取")
    except Exception:
        if "Can't load page" in driver.page_source:
            print("页面加载失败,刷新重试...")
            driver.refresh()
            time.sleep(3)
        else:
            print("出现未知错误,终止爬取")
            break

driver.quit()
print(f"爬取完成,共获取 {len(tweets)} 条推文")

内容的提问来源于stack exchange,提问作者W39

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 17:57:53