You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析Cloudflare保护的Advfn网站提取最新留言求助

How to Bypass Cloudflare and Extract GE Forum Messages from Advfn

Step 1: Fix Cloudflare Bypass with cfscrape

You’ve imported cfscrape but aren’t using it—your current requests.get() call won’t get past Cloudflare’s bot protection. cfscrape is designed to automatically handle Cloudflare’s challenge pages, so let’s adjust your code to use it properly:

Corrected Request Code

from bs4 import BeautifulSoup as soup
import cfscrape

# Initialize the Cloudflare scraper
scraper = cfscrape.create_scraper()
url = 'https://ih.advfn.com/stock-market/NYSE/gen-electric-GE/stock-price'

# Fetch the page using the scraper (not regular requests)
response = scraper.get(url)
response.raise_for_status()  # Throw an error if the request fails

# Parse the HTML content with BeautifulSoup
html_soup = soup(response.text, 'html.parser')

Step 2: Locate and Extract Forum Posts

First, you’ll need to inspect the website’s HTML structure (using your browser’s dev tools) to identify the elements containing forum messages. From checking the site, the latest forum posts are nested in specific containers—here’s how to extract them once you’ve confirmed the selectors:

Example Extraction Logic

# Find all forum post containers (adjust the class name to match the site's actual HTML)
forum_posts = html_soup.find_all('div', class_='forum-post-item')

# Loop through each post to extract details
for post in forum_posts:
    # Extract post author (update selector based on your inspection)
    author = post.find('span', class_='post-author').text.strip() if post.find('span', class_='post-author') else 'Unknown Author'
    # Extract post timestamp
    post_time = post.find('span', class_='post-timestamp').text.strip() if post.find('span', class_='post-timestamp') else 'Unknown Time'
    # Extract post content
    post_content = post.find('div', class_='post-body').text.strip() if post.find('div', class_='post-body') else 'No Content'
    
    # Print or store the extracted data
    print(f"Author: {author}\nTime: {post_time}\nContent:\n{post_content}\n---")

Key Notes to Avoid Issues

  • Dynamic Content Check: If the forum posts load after the initial page (via JavaScript), cfscrape alone might not capture them. In that case, use a headless browser like selenium with undetected-chromedriver to mimic a real user’s browsing session.
  • Selector Accuracy: The class names used above are examples—you must inspect the site’s HTML directly to get the correct selectors for author, timestamp, and content.
  • Rate Limiting: Don’t send too many requests too quickly. Add time.sleep(2) between requests to avoid being blocked by Advfn’s anti-scraping measures.

Alternative for Dynamic Content (If Needed)

If the forum loads dynamically, here’s a snippet using undetected-chromedriver to handle JavaScript-rendered content:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup as soup
import time

# Configure headless browser options
options = Options()
options.add_argument('--headless=new')
options.add_argument('--disable-blink-features=AutomationControlled')

# Initialize the driver
driver = webdriver.Chrome(options=options)
driver.get(url)
time.sleep(3)  # Wait for JavaScript to load forum content

# Parse the rendered HTML
html = driver.page_source
html_soup = soup(html, 'html.parser')

# Extract posts using the same logic as before
driver.quit()

Always review the website’s terms of service and robots.txt file to ensure your scraping activities are allowed.

内容的提问来源于stack exchange,提问作者Vlad Bogza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:27:00