使用Python从在线URL的XPath提取特定文本失败求助
问题:Discogs在线页面XPath提取无输出,本地HTML正常
我想从Discogs的艺人页面提取文本,预期输出为:
Artista 1: Total Eclipse (4) Testo elemento 1: Jungle Fever Testo elemento 2: Come Together
但运行脚本后控制台无任何输出,不过脚本在本地HTML文件上能正常工作。使用的Python脚本如下:
import requests from lxml import html # Ask the user to enter the Discogs URL url = input("Enter the Discogs URL: ") # Make an HTTP request to get the HTML content of the page response = requests.get(url) html_content = response.text # Parse HTML using lxml tree = html.fromstring(html_content) # Use a generic XPath to capture as many "h1" and "a" elements as desired elements = tree.xpath('//*[starts-with(local-name(), "div")][4]/div/div[1]/div[1]/div/h1 | //*[starts-with(local-name(), "div")][4]/div/div[2]/div[2]/div/table//*[starts-with(local-name(), "tr")][2]/td[5]/a | //*[starts-with(local-name(), "div")][4]/div/div[2]/div[2]/div/table//*[starts-with(local-name(), "tr")]/td[5]/a') # Counters to keep track of "a" item artists and lyrics artist_counter = 1 text_counter = 1 # Print the text of the "h1" and "a" elements found in sequence for element in elements: if element.tag == 'h1': print(f"Artista {artist_counter}: {element.text.strip()}") artist_counter += 1 else: print(f"Testo elemento {text_counter}: {element.text.strip()}") text_counter += 1
问题原因
- 反爬拦截:Discogs会识别非浏览器请求,直接用
requests.get发送的请求缺少浏览器标识,会被网站拦截,返回空白或反爬页面,导致XPath无法匹配元素。 - XPath路径脆弱:原XPath依赖元素的索引位置(如
div[4]),在线页面的DOM结构可能和本地保存版本有差异,索引变化后XPath直接失效。
修复方案
1. 添加请求头模拟浏览器
给请求添加User-Agent头,伪装成浏览器请求:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers)
2. 优化XPath路径,使用稳定定位
避免依赖元素索引,改用class或属性定位:
- 艺人名称:用
//h1[@class='profile_title']定位 - 作品名称:用
//table[@class='table_responsive']//tr[@class='shortcut_navigable']/td[contains(@class, 'title')]/a定位
修改后的完整脚本:
import requests from lxml import html # Ask the user to enter the Discogs URL url = input("Enter the Discogs URL: ") # 添加浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } # 发送请求并检查状态 response = requests.get(url, headers=headers) if response.status_code != 200: print(f"请求失败,状态码:{response.status_code}") exit() html_content = response.text tree = html.fromstring(html_content) # 定位目标元素 artist = tree.xpath('//h1[@class="profile_title"]/text()') tracks = tree.xpath('//table[@class="table_responsive"]//tr[@class="shortcut_navigable"]/td[contains(@class, "title")]/a/text()') # 输出结果 if artist: print(f"Artista 1: {artist[0].strip()}") for idx, track in enumerate(tracks, 1): print(f"Testo elemento {idx}: {track.strip()}")
内容的提问来源于stack exchange,提问作者Peter Long
相关产品推荐
相关产品推荐

