Python新手求助:原网页HTML获取解析代码突然失效
Hey there! Let's figure out why your web scraping script stopped working—since you suspect the site changed, here are the most likely fixes to get you back up and running:
Common Issues & Fixes
1. Outdated User-Agent
Your old user-agent (Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 ...) is super outdated (Chrome 66 is from 2018!). Many sites block old or suspicious user-agents to prevent scraping.
Fix: Update your user-agent to match a modern browser. For example:
user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
2. Deprecated Libraries (Python 2 vs 3)
The cookielib and urllib2 modules you're using are from Python 2, which is no longer supported. If you're running this on Python 3, those libraries won't work properly.
Fix: Switch to requests (a far more user-friendly HTTP library) which handles cookies automatically. Here's a modern replacement for your code skeleton:
import requests from lxml import etree user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' headers = {'User-Agent': user_agent} # Send request with cookies enabled by default response = requests.get('YOUR_TARGET_URL', headers=headers) # Check if the request succeeded (status code 200 = OK) print(response.status_code) # Parse the HTML tree = etree.HTML(response.content) # Replace with your original XPath selectors target_text = tree.xpath('//YOUR_XPATH_HERE/text()') print(target_text)
3. Changed Website Structure
Even if you get the HTML, your XPath selectors might fail if the site updated its layout.
Fix:
- First, print
response.text(orresponse.content) to see if you're actually getting the full page HTML (not an error page or a blocked message). - Use your browser's Developer Tools (F12) to inspect the new page structure and update your XPath/CSS selectors accordingly.
4. Enhanced Anti-Scraping Measures
If the site added JavaScript rendering (content loads after page load) or CAPTCHAs, your static HTTP request won't capture the data.
Fix:
- For dynamic content, use tools like
seleniumorplaywrightto simulate a real browser. Here's a quick selenium example:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from lxml import etree options = Options() options.add_argument('--headless=new') # Run in background options.add_argument(f'user-agent={user_agent}') driver = webdriver.Chrome(options=options) driver.get('YOUR_TARGET_URL') # Get the fully rendered HTML tree = etree.HTML(driver.page_source) target_text = tree.xpath('//YOUR_XPATH_HERE/text()') print(target_text) driver.quit()
Quick Checks First
Before diving into code changes:
- Test the URL in your browser to confirm it still loads the data you need.
- Check if the site requires you to log in now (your old cookie setup might not work if session rules changed).
内容的提问来源于stack exchange,提问作者Rakesh

