You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python工具爬取混淆网页内容?含Twitter爬取失效场景

爬取Twitter混淆后用户加入日期的解决方案

Got it, let's tackle this problem—Twitter's dynamic CSS class names definitely throw a wrench into the old scraping approach. Here are a few solid methods using the tools you mentioned to extract those join dates:

1. Requests + BeautifulSoup

Since fixed class names are no longer reliable, we can target elements based on stable patterns like text content or SVG structure (the calendar-style SVG next to the join date has consistent attributes).

import requests
from bs4 import BeautifulSoup

# Replace with your target user's profile URL
TARGET_URL = "https://twitter.com/realDonaldTrump"
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# Fetch the page
response = requests.get(TARGET_URL, headers=HEADERS)
soup = BeautifulSoup(response.text, "html.parser")

# Method 1: Target spans containing the "Joined" keyword
join_date_spans = soup.find_all("span", string=lambda text: text and "Joined" in text.strip())
for span in join_date_spans:
    print("Join Date:", span.get_text(strip=True))

# Method 2: Locate via the calendar SVG's attributes
calendar_svgs = soup.find_all("svg", attrs={"viewBox": "0 0 24 24"})
for svg in calendar_svgs:
    parent_span = svg.find_parent("span")
    if parent_span and "Joined" in parent_span.get_text(strip=True):
        print("Join Date (via SVG):", parent_span.get_text(strip=True))

Why this works: The "Joined" text string and SVG viewBox attribute are static, even as Twitter's CSS classes get randomized.

2. Scrapy

If you're working at scale, Scrapy's robust selectors make this straightforward. We'll use XPath to target the same stable patterns:

import scrapy

class TwitterJoinDateSpider(scrapy.Spider):
    name = "twitter_join_date"
    start_urls = ["https://twitter.com/realDonaldTrump"]
    custom_settings = {
        "USER_AGENT": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }

    def parse(self, response):
        # Extract via "Joined" text match
        join_dates = response.xpath('//span[contains(normalize-space(text()), "Joined")]/text()').getall()
        for date in join_dates:
            yield {"username": response.url.split("/")[-1], "join_date": date.strip()}

        # Alternative: Extract via calendar SVG parent span
        svg_parent_spans = response.xpath('//svg[@viewBox="0 0 24 24"]/parent::span')
        for span in svg_parent_spans:
            join_text = span.xpath('normalize-space(text())').get()
            if "Joined" in join_text:
                yield {"username": response.url.split("/")[-1], "join_date": join_text}

Pro tip: Add delays or rotate proxies in your Scrapy settings to avoid hitting Twitter's rate limits.

3. Selenium (For JS-Rendered Content)

If the page relies heavily on JavaScript to load content (common for logged-in views or dynamic profiles), Selenium simulates a real browser to fully render the page:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize driver (make sure ChromeDriver is in your PATH)
driver = webdriver.Chrome()
TARGET_URL = "https://twitter.com/realDonaldTrump"

try:
    driver.get(TARGET_URL)
    # Wait up to 10 seconds for the join date elements to load
    join_date_elements = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.XPATH, '//span[contains(text(), "Joined")]'))
    )
    for element in join_date_elements:
        print("Join Date:", element.text.strip())
finally:
    driver.quit()

Bonus: Selenium makes it easy to handle Twitter's login flow if you need access to private profiles—just add steps to enter your credentials before navigating to the target page.

Important Notes

  • Anti-scraping measures: Twitter actively blocks scrapers, so always use a realistic user-agent, add random delays between requests, and consider using authenticated sessions (via cookies or Selenium login) to avoid being blocked.
  • Rate limits: Even with authenticated requests, Twitter has strict rate limits—keep this in mind if you're planning large-scale scraping.

内容的提问来源于stack exchange,提问作者copy data

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 13:24:58