You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取Twitter重复要求登录问题排查与解决指导

问题排查与解决指导

一、代码中的核心问题及修复

你的代码存在几个关键错误,导致无法正确抓取数据甚至报错:

  1. XPath路径错误
    你使用的//span[contains(text(), "@")]是绝对路径,会从页面根节点查找所有匹配元素,而非当前tweet节点下的元素。应改为相对路径.//,限定在当前推文范围内查找。

  2. find_elements误用
    find_elements返回的是元素列表,直接调用.text会报错。每个推文对应一个用户和一段文本,应使用find_element(单数)获取单个元素。

  3. 列表追加逻辑颠倒
    你写的user.append(user_data)是把列表追加到变量里,正确的应该是user_data.append(user),将获取到的用户信息添加到结果列表中。

  4. 固定等待不可靠
    time.sleep(5)是固定等待,无法保证元素加载完成。应使用WebDriverWait实现显式等待,提升稳定性。

二、登录状态不持久的解决方法

Selenium默认启动全新的Chrome匿名会话,不会复用你之前登录的浏览器配置。要保留登录状态,需指定Chrome的用户数据目录:

  1. 找到Chrome用户数据目录

    • Windows: C:\Users\你的用户名\AppData\Local\Google\Chrome\User Data
    • macOS: ~/Library/Application Support/Google/Chrome
    • Linux: ~/.config/google-chrome
  2. 在ChromeOptions中配置
    添加参数指定用户数据目录,让Selenium复用已登录的浏览器 profile。

三、修复后的完整代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By

# 配置Chrome选项,复用已登录的用户profile
options = Options()
# 替换为你的Chrome用户数据目录路径
options.add_argument(r"user-data-dir=C:\Users\你的用户名\AppData\Local\Google\Chrome\User Data")
options.add_argument("--profile-directory=Default")
# 隐藏Selenium自动化特征,避免被Twitter检测
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(options=options)
driver.get('https://twitter.com/search?q=python&src=typed_query')
driver.maximize_window()

# 显式等待推文元素加载完成
wait = WebDriverWait(driver, 10)
tweets = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//article[@role="article"]')))

user_data = []
text_data = []

for tweet in tweets:
    try:
        # 相对路径查找当前推文内的用户handle
        user = tweet.find_element(By.XPATH, './/span[contains(text(), "@")]').text
        # 相对路径查找当前推文内的文本内容
        text = tweet.find_element(By.XPATH, './/div[@lang]').text
        user_data.append(user)
        text_data.append(text)
    except Exception as e:
        # 跳过加载异常的推文
        continue

driver.quit()

df_tweets = pd.DataFrame({'user': user_data, 'text': text_data})
df_tweets.to_csv('tweets.csv', index=False)
print(df_tweets)

额外注意事项

  • 确保替换代码中的用户数据目录路径为你自己的路径。
  • 如果仍被要求登录,先手动打开Chrome登录Twitter,关闭后再运行脚本。
  • Twitter有反爬机制,频繁抓取可能会触发限制,建议添加合理的间隔时间。

内容的提问来源于stack exchange,提问作者Win123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 13:08:20