使用Requests和Scrapy Selector爬取SeekingAlpha评论的HTML差异及爬取问题
Let's break down your problem step by step:
1. Why Doesn't the HTTP Response Match the Browser's HTML?
The core issue here is dynamic content rendering. When you use requests.get(), you're only fetching the initial static HTML sent by the server. However, modern websites like Seeking Alpha load many parts of the page (including comments) using JavaScript after the initial page loads.
- The browser executes this JavaScript, which makes additional AJAX requests to fetch comment data, then renders that data into the DOM (the HTML you see in DevTools).
requestsdoesn't execute JavaScript—it just grabs the raw server response, so the comment elements you're targeting (//div[@class='b-b']) don't exist in that initial HTML. That's why your XPath query returns an empty list.
2. How to Fetch the Comment Data
You have two reliable approaches to get the comments:
Option 1: Directly Call the Comments API (Recommended)
Instead of scraping rendered HTML, find the API endpoint that the website uses to load comments. This is faster and more stable than rendering the whole page. Here's how to do it:
- Open your browser's DevTools (F12) → Go to the Network tab.
- Refresh the article page, filter for XHR/Fetch requests.
- Look for requests related to comments (they might have keywords like "comments", "thread", or the article ID in the URL).
- Copy the request URL, headers, and any required parameters (like
articleId).
For your specific article, here's a sample implementation (adjust with the actual API details you find):
import requests # Replace with the real comments API endpoint from DevTools comments_url = "https://seekingalpha.com/api/v3/comments?article_id=4312816" headers = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'accept': 'application/json', # Add cookies/authorization headers if required for access } response = requests.get(comments_url, headers=headers) comments_data = response.json() # Parse the JSON to extract comments for comment in comments_data['data']: comment_text = comment['attributes']['body'] comment_author = comment['attributes']['author']['username'] print(f"{comment_author}: {comment_text}")
Option 2: Use a Headless Browser to Render JavaScript
If you prefer scraping the rendered HTML, use a tool that simulates a browser (like Playwright, Selenium, or Scrapy Splash) to execute JavaScript and load comments. Here's an example with Playwright:
First, install Playwright:
pip install playwright playwright install chrome
Then, use it to fetch the fully rendered page:
from playwright.sync_api import sync_playwright from scrapy import Selector url = "https://seekingalpha.com/article/4312816-exxon-mobil-dividend-problems" with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") page.goto(url) # Wait for comments to load (adjust the selector to match the actual comment container) page.wait_for_selector("div.b-b") # Get the fully rendered HTML rendered_html = page.content() sel = Selector(text=rendered_html) # Extract comments comments = sel.xpath("//div[@class='b-b']") for comment in comments: comment_text = ' '.join(comment.xpath(".//text()").getall()) print(comment_text) browser.close()
Key Tips to Avoid Blocks
- Use a Realistic User-Agent: Don't just send the AppleWebKit string—use a full, modern browser user-agent to avoid being flagged as a bot.
- Respect Rate Limits: Add small delays between requests to avoid overwhelming the server.
- Check for Authentication: Some parts of Seeking Alpha require a logged-in session (via cookies) to access comments. If you get unauthorized errors, you may need to log in programmatically first.
内容的提问来源于stack exchange,提问作者mk_sch

