Python解析Cloudflare保护的Advfn网站提取最新留言求助
Step 1: Fix Cloudflare Bypass with cfscrape
You’ve imported cfscrape but aren’t using it—your current requests.get() call won’t get past Cloudflare’s bot protection. cfscrape is designed to automatically handle Cloudflare’s challenge pages, so let’s adjust your code to use it properly:
Corrected Request Code
from bs4 import BeautifulSoup as soup import cfscrape # Initialize the Cloudflare scraper scraper = cfscrape.create_scraper() url = 'https://ih.advfn.com/stock-market/NYSE/gen-electric-GE/stock-price' # Fetch the page using the scraper (not regular requests) response = scraper.get(url) response.raise_for_status() # Throw an error if the request fails # Parse the HTML content with BeautifulSoup html_soup = soup(response.text, 'html.parser')
Step 2: Locate and Extract Forum Posts
First, you’ll need to inspect the website’s HTML structure (using your browser’s dev tools) to identify the elements containing forum messages. From checking the site, the latest forum posts are nested in specific containers—here’s how to extract them once you’ve confirmed the selectors:
Example Extraction Logic
# Find all forum post containers (adjust the class name to match the site's actual HTML) forum_posts = html_soup.find_all('div', class_='forum-post-item') # Loop through each post to extract details for post in forum_posts: # Extract post author (update selector based on your inspection) author = post.find('span', class_='post-author').text.strip() if post.find('span', class_='post-author') else 'Unknown Author' # Extract post timestamp post_time = post.find('span', class_='post-timestamp').text.strip() if post.find('span', class_='post-timestamp') else 'Unknown Time' # Extract post content post_content = post.find('div', class_='post-body').text.strip() if post.find('div', class_='post-body') else 'No Content' # Print or store the extracted data print(f"Author: {author}\nTime: {post_time}\nContent:\n{post_content}\n---")
Key Notes to Avoid Issues
- Dynamic Content Check: If the forum posts load after the initial page (via JavaScript),
cfscrapealone might not capture them. In that case, use a headless browser likeseleniumwithundetected-chromedriverto mimic a real user’s browsing session. - Selector Accuracy: The class names used above are examples—you must inspect the site’s HTML directly to get the correct selectors for author, timestamp, and content.
- Rate Limiting: Don’t send too many requests too quickly. Add
time.sleep(2)between requests to avoid being blocked by Advfn’s anti-scraping measures.
Alternative for Dynamic Content (If Needed)
If the forum loads dynamically, here’s a snippet using undetected-chromedriver to handle JavaScript-rendered content:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup as soup import time # Configure headless browser options options = Options() options.add_argument('--headless=new') options.add_argument('--disable-blink-features=AutomationControlled') # Initialize the driver driver = webdriver.Chrome(options=options) driver.get(url) time.sleep(3) # Wait for JavaScript to load forum content # Parse the rendered HTML html = driver.page_source html_soup = soup(html, 'html.parser') # Extract posts using the same logic as before driver.quit()
Always review the website’s terms of service and
robots.txtfile to ensure your scraping activities are allowed.
内容的提问来源于stack exchange,提问作者Vlad Bogza

