You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

lxml.etree.XPathEvalError无效表达式问题求助

解决lxml XPath无效表达式错误

问题重现

尝试用lxml解析网页时,调用parsed.xpath()一直提示无效表达式,已通过pip安装lxml库,代码和报错信息如下:

原代码:

import requests
import lxml.html as html

HOME_URL = 'https://www.larepublica.co/' #URL de la página home

XPATH_LINK_TO_ARTICLE = "$x('//div[@class='newsV_Title_Img' or @class='V_Title']/text-fill/a/@href').map(x => x.value)"
XPATH_TITLE = "$x('//full-text/span/text()').map(x => x.wholeText)"
XPATH_SUMMARY = "$x('//div[@class='lead']/p/text()').map(x => x.wholeText)"
XPATH_BODY = "$x('//div[@class='html-content']/p[not(@class)]/text()').map(x => x.wholeText)"

#Función para extraer los links de las noticias
def parse_home():
    #Envolver el código dentro de un bloque try except
    try:
        response = requests.get(HOME_URL)
        if response.status_code == 200:
            home = response.content.decode('utf-8')
            parsed = html.fromstring(home) 
            links_to_notices = parsed.xpath(XPATH_LINK_TO_ARTICLE) 
            print(links_to_notices)
        else:
            raise ValueError(f'Error: {response.status_code}') 
    except ValueError as ve: 
        print(ve)

def run():
    parse_home()

if __name__ == '__main__':
    run()

报错信息:

Traceback (most recent call last):
  File "D:\Andres\PLATZI\Curso_Fundamentos_WebScrapping\LaRepublica_Scrapper\scraper.py", line 30, in <module>
    run()
  File "D:\Andres\PLATZI\Curso_Fundamentos_WebScrapping\LaRepublica_Scrapper\scraper.py", line 27, in run
    parse_home()
  File "D:\Andres\PLATZI\Curso_Fundamentos_WebScrapping\LaRepublica_Scrapper\scraper.py", line 19, in parse_home
    links_to_notices = parsed.xpath(XPATH_LINK_TO_ARTICLE) #Obtener lista de links con comandos xtpath del documento contenido en parsed
  File "src\lxml\etree.pyx", line 1599, in lxml.etree._Element.xpath
  File "src\lxml\xpath.pxi", line 305, in lxml.etree.XPathElementEvaluator.__call__
  File "src\lxml\xpath.pxi", line 225, in lxml.etree._XPathEvaluatorBase._handle_result
lxml.etree.XPathEvalError: Invalid expression

问题原因

你直接复制了Chrome浏览器控制台的XPath调用代码,但这些是浏览器DevTools专属的JavaScript语法,lxml的xpath方法只支持标准XPath表达式:

  1. $x()是Chrome提供的快捷查询函数,不属于标准XPath
  2. .map(x => x.value)是JavaScript的数组处理逻辑,lxml的xpath会直接返回匹配结果的列表
  3. 字符串嵌套单引号导致语法冲突(比如'//div[@class='newsV_Title_Img']'里的单引号会被Python解析为字符串结束符)

修正后的代码

import requests
import lxml.html as html

HOME_URL = 'https://www.larepublica.co/'

# 使用标准XPath表达式,修正引号嵌套问题
XPATH_LINK_TO_ARTICLE = "//div[@class='newsV_Title_Img' or @class='V_Title']/text-fill/a/@href"
XPATH_TITLE = "//full-text/span/text()"
XPATH_SUMMARY = "//div[@class='lead']/p/text()"
XPATH_BODY = "//div[@class='html-content']/p[not(@class)]/text()"

def parse_home():
    try:
        response = requests.get(HOME_URL)
        if response.status_code == 200:
            home = response.content.decode('utf-8')
            parsed = html.fromstring(home) 
            links_to_notices = parsed.xpath(XPATH_LINK_TO_ARTICLE) 
            print(links_to_notices)
        else:
            raise ValueError(f'Error: {response.status_code}') 
    except ValueError as ve: 
        print(ve)

def run():
    parse_home()

if __name__ == '__main__':
    run()

关键修正点

  • 移除$x()包裹和.map()逻辑,直接编写纯XPath表达式
  • 用双引号包裹整个XPath字符串,避免内部属性值的单引号冲突(也可以用单引号包裹整个字符串,内部属性值改用双引号)
  • lxml的xpath方法会自动返回匹配到的属性值或文本内容的列表,无需额外处理

内容的提问来源于stack exchange,提问作者Andres Cardona

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 20:36:38