使用requests/BeautifulSoup无法抓取网站定期更新数值的技术问题
Hey there, I’ve dealt with this exact scenario countless times—you see the data right there in DevTools, but requests and BeautifulSoup just can’t pick it up. Let’s walk through the most likely causes and how to fix them:
1. The content is dynamically rendered with JavaScript
When you load the page in a browser, it runs JavaScript to populate the DOM with those numbers. But requests only fetches the raw, unprocessed HTML—no JS execution means no dynamic content.
Solutions:
- Use a browser automation tool like Selenium to mimic a real browser (it runs JS and renders the full page):
from selenium import webdriver from bs4 import BeautifulSoup # Initialize Chrome (make sure you have chromedriver matching your Chrome version) driver = webdriver.Chrome() driver.get("your-target-url-here") # Grab the fully rendered page source page_html = driver.page_source soup = BeautifulSoup(page_html, "html.parser") # Now find your target element advance_td = soup.find("td", style=lambda s: s and "color:green" in s) if advance_td: print(f"Advances: {advance_td.text.strip()}") driver.quit() - Or use
requests-html, a lighter library that supports JS rendering:from requests_html import HTMLSession session = HTMLSession() response = session.get("your-target-url-here") response.html.render() # Triggers JS execution advance_td = response.html.find('td[style*="color:green"]', first=True) print(f"Advances: {advance_td.text.strip()}")
2. Your request headers are too "robot-like"
Many sites serve different content based on the User-Agent header. The default requests header (python-requests/x.x.x) flags you as a bot, so you might get a stripped-down version of the page that doesn’t include the data you want.
Solution:
Spoof a browser’s User-Agent (and other headers if needed):
import requests from bs4 import BeautifulSoup headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9" } response = requests.get("your-target-url-here", headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Try locating the element again with the updated headers target_td = soup.find("td", style=lambda s: s and "border-right: 1px solid #ACA99F" in s and "color:green" in s) print(target_td.text.strip() if target_td else "Element not found—might still be dynamic content")
3. The data is inside an iframe
If your target <td> lives inside an <iframe> tag, requests only fetches the main page’s HTML—it won’t automatically load the iframe’s content.
Solution:
Find the iframe’s src attribute, then request that URL directly:
import requests from bs4 import BeautifulSoup headers = {"User-Agent": "your-browser-user-agent"} # First get the main page to find the iframe main_response = requests.get("your-target-url-here", headers=headers) main_soup = BeautifulSoup(main_response.text, "html.parser") # Extract the iframe source (handle relative paths by adding the base domain) iframe_src = main_soup.find("iframe")["src"] iframe_url = "https://your-target-site-domain.com" + iframe_src if not iframe_src.startswith("http") else iframe_src # Now fetch the iframe content iframe_response = requests.get(iframe_url, headers=headers) iframe_soup = BeautifulSoup(iframe_response.text, "html.parser") # Locate your data in the iframe's HTML advance_td = iframe_soup.find("td", style=lambda s: s and "color:green" in s) print(f"Advances: {advance_td.text.strip()}")
4. Scrape the API directly (best practice!)
Most dynamic sites load data via an API behind the scenes. Instead of scraping HTML, you can hit this API directly for cleaner, more reliable data.
How to find the API:
- Open DevTools (F12) in your browser, go to the Network tab
- Refresh the page, filter requests by XHR/Fetch
- Look for requests that return JSON data containing your Advances/Declines numbers
- Copy that request’s URL and use it with
requests:
import requests headers = {"User-Agent": "your-browser-user-agent"} api_url = "the-api-url-you-found-in-network-tab" response = requests.get(api_url, headers=headers) data = response.json() # Example: if the JSON structure has keys like 'advances', 'declines' print(f"Advances: {data['advances']}, Declines: {data['declines']}, Unchanged: {data['unchanged']}, Total: {data['total']}")
内容的提问来源于stack exchange,提问作者p699

