You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:原网页HTML获取解析代码突然失效

Troubleshooting Your Broken Python Web Scraping Code

Hey there! Let's figure out why your web scraping script stopped working—since you suspect the site changed, here are the most likely fixes to get you back up and running:

Common Issues & Fixes

1. Outdated User-Agent

Your old user-agent (Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 ...) is super outdated (Chrome 66 is from 2018!). Many sites block old or suspicious user-agents to prevent scraping.

Fix: Update your user-agent to match a modern browser. For example:

user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'

2. Deprecated Libraries (Python 2 vs 3)

The cookielib and urllib2 modules you're using are from Python 2, which is no longer supported. If you're running this on Python 3, those libraries won't work properly.

Fix: Switch to requests (a far more user-friendly HTTP library) which handles cookies automatically. Here's a modern replacement for your code skeleton:

import requests
from lxml import etree

user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
headers = {'User-Agent': user_agent}

# Send request with cookies enabled by default
response = requests.get('YOUR_TARGET_URL', headers=headers)

# Check if the request succeeded (status code 200 = OK)
print(response.status_code)

# Parse the HTML
tree = etree.HTML(response.content)
# Replace with your original XPath selectors
target_text = tree.xpath('//YOUR_XPATH_HERE/text()')
print(target_text)

3. Changed Website Structure

Even if you get the HTML, your XPath selectors might fail if the site updated its layout.

Fix:

  • First, print response.text (or response.content) to see if you're actually getting the full page HTML (not an error page or a blocked message).
  • Use your browser's Developer Tools (F12) to inspect the new page structure and update your XPath/CSS selectors accordingly.

4. Enhanced Anti-Scraping Measures

If the site added JavaScript rendering (content loads after page load) or CAPTCHAs, your static HTTP request won't capture the data.

Fix:

  • For dynamic content, use tools like selenium or playwright to simulate a real browser. Here's a quick selenium example:
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from lxml import etree

options = Options()
options.add_argument('--headless=new')  # Run in background
options.add_argument(f'user-agent={user_agent}')

driver = webdriver.Chrome(options=options)
driver.get('YOUR_TARGET_URL')

# Get the fully rendered HTML
tree = etree.HTML(driver.page_source)
target_text = tree.xpath('//YOUR_XPATH_HERE/text()')
print(target_text)

driver.quit()

Quick Checks First

Before diving into code changes:

  • Test the URL in your browser to confirm it still loads the data you need.
  • Check if the site requires you to log in now (your old cookie setup might not work if session rules changed).

内容的提问来源于stack exchange,提问作者Rakesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:53:22