You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests和Scrapy Selector爬取SeekingAlpha评论的HTML差异及爬取问题

Why the HTTP Response HTML Differs from Browser & How to Fetch Comments on Seeking Alpha

Let's break down your problem step by step:

1. Why Doesn't the HTTP Response Match the Browser's HTML?

The core issue here is dynamic content rendering. When you use requests.get(), you're only fetching the initial static HTML sent by the server. However, modern websites like Seeking Alpha load many parts of the page (including comments) using JavaScript after the initial page loads.

  • The browser executes this JavaScript, which makes additional AJAX requests to fetch comment data, then renders that data into the DOM (the HTML you see in DevTools).
  • requests doesn't execute JavaScript—it just grabs the raw server response, so the comment elements you're targeting (//div[@class='b-b']) don't exist in that initial HTML. That's why your XPath query returns an empty list.

2. How to Fetch the Comment Data

You have two reliable approaches to get the comments:

Instead of scraping rendered HTML, find the API endpoint that the website uses to load comments. This is faster and more stable than rendering the whole page. Here's how to do it:

  1. Open your browser's DevTools (F12) → Go to the Network tab.
  2. Refresh the article page, filter for XHR/Fetch requests.
  3. Look for requests related to comments (they might have keywords like "comments", "thread", or the article ID in the URL).
  4. Copy the request URL, headers, and any required parameters (like articleId).

For your specific article, here's a sample implementation (adjust with the actual API details you find):

import requests

# Replace with the real comments API endpoint from DevTools
comments_url = "https://seekingalpha.com/api/v3/comments?article_id=4312816"
headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
    'accept': 'application/json',
    # Add cookies/authorization headers if required for access
}

response = requests.get(comments_url, headers=headers)
comments_data = response.json()

# Parse the JSON to extract comments
for comment in comments_data['data']:
    comment_text = comment['attributes']['body']
    comment_author = comment['attributes']['author']['username']
    print(f"{comment_author}: {comment_text}")

Option 2: Use a Headless Browser to Render JavaScript

If you prefer scraping the rendered HTML, use a tool that simulates a browser (like Playwright, Selenium, or Scrapy Splash) to execute JavaScript and load comments. Here's an example with Playwright:

First, install Playwright:

pip install playwright
playwright install chrome

Then, use it to fetch the fully rendered page:

from playwright.sync_api import sync_playwright
from scrapy import Selector

url = "https://seekingalpha.com/article/4312816-exxon-mobil-dividend-problems"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
    page.goto(url)
    # Wait for comments to load (adjust the selector to match the actual comment container)
    page.wait_for_selector("div.b-b")
    
    # Get the fully rendered HTML
    rendered_html = page.content()
    sel = Selector(text=rendered_html)
    
    # Extract comments
    comments = sel.xpath("//div[@class='b-b']")
    for comment in comments:
        comment_text = ' '.join(comment.xpath(".//text()").getall())
        print(comment_text)
    
    browser.close()

Key Tips to Avoid Blocks

  • Use a Realistic User-Agent: Don't just send the AppleWebKit string—use a full, modern browser user-agent to avoid being flagged as a bot.
  • Respect Rate Limits: Add small delays between requests to avoid overwhelming the server.
  • Check for Authentication: Some parts of Seeking Alpha require a logged-in session (via cookies) to access comments. If you get unauthorized errors, you may need to log in programmatically first.

内容的提问来源于stack exchange,提问作者mk_sch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:12:14