You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python从在线URL的XPath提取特定文本失败求助

问题:Discogs在线页面XPath提取无输出,本地HTML正常

我想从Discogs的艺人页面提取文本,预期输出为:

Artista 1: Total Eclipse (4)
Testo elemento 1: Jungle Fever
Testo elemento 2: Come Together

但运行脚本后控制台无任何输出,不过脚本在本地HTML文件上能正常工作。使用的Python脚本如下:

import requests
from lxml import html

# Ask the user to enter the Discogs URL
url = input("Enter the Discogs URL: ")

# Make an HTTP request to get the HTML content of the page
response = requests.get(url)
html_content = response.text

# Parse HTML using lxml
tree = html.fromstring(html_content)

# Use a generic XPath to capture as many "h1" and "a" elements as desired
elements = tree.xpath('//*[starts-with(local-name(), "div")][4]/div/div[1]/div[1]/div/h1 | //*[starts-with(local-name(), "div")][4]/div/div[2]/div[2]/div/table//*[starts-with(local-name(), "tr")][2]/td[5]/a | //*[starts-with(local-name(), "div")][4]/div/div[2]/div[2]/div/table//*[starts-with(local-name(), "tr")]/td[5]/a')

# Counters to keep track of "a" item artists and lyrics
artist_counter = 1
text_counter = 1

# Print the text of the "h1" and "a" elements found in sequence
for element in elements:
    if element.tag == 'h1':
        print(f"Artista {artist_counter}: {element.text.strip()}")
        artist_counter += 1
    else:
        print(f"Testo elemento {text_counter}: {element.text.strip()}")
        text_counter += 1

问题原因

  1. 反爬拦截:Discogs会识别非浏览器请求,直接用requests.get发送的请求缺少浏览器标识,会被网站拦截,返回空白或反爬页面,导致XPath无法匹配元素。
  2. XPath路径脆弱:原XPath依赖元素的索引位置(如div[4]),在线页面的DOM结构可能和本地保存版本有差异,索引变化后XPath直接失效。

修复方案

1. 添加请求头模拟浏览器

给请求添加User-Agent头,伪装成浏览器请求:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)

2. 优化XPath路径,使用稳定定位

避免依赖元素索引,改用class或属性定位:

  • 艺人名称:用//h1[@class='profile_title']定位
  • 作品名称:用//table[@class='table_responsive']//tr[@class='shortcut_navigable']/td[contains(@class, 'title')]/a定位

修改后的完整脚本:

import requests
from lxml import html

# Ask the user to enter the Discogs URL
url = input("Enter the Discogs URL: ")

# 添加浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

# 发送请求并检查状态
response = requests.get(url, headers=headers)
if response.status_code != 200:
    print(f"请求失败,状态码:{response.status_code}")
    exit()

html_content = response.text
tree = html.fromstring(html_content)

# 定位目标元素
artist = tree.xpath('//h1[@class="profile_title"]/text()')
tracks = tree.xpath('//table[@class="table_responsive"]//tr[@class="shortcut_navigable"]/td[contains(@class, "title")]/a/text()')

# 输出结果
if artist:
    print(f"Artista 1: {artist[0].strip()}")
for idx, track in enumerate(tracks, 1):
    print(f"Testo elemento {idx}: {track.strip()}")

内容的提问来源于stack exchange,提问作者Peter Long

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 09:01:28