lxml.etree.XPathEvalError无效表达式问题求助
解决lxml XPath无效表达式错误
问题重现
尝试用lxml解析网页时,调用parsed.xpath()一直提示无效表达式,已通过pip安装lxml库,代码和报错信息如下:
原代码:
import requests import lxml.html as html HOME_URL = 'https://www.larepublica.co/' #URL de la página home XPATH_LINK_TO_ARTICLE = "$x('//div[@class='newsV_Title_Img' or @class='V_Title']/text-fill/a/@href').map(x => x.value)" XPATH_TITLE = "$x('//full-text/span/text()').map(x => x.wholeText)" XPATH_SUMMARY = "$x('//div[@class='lead']/p/text()').map(x => x.wholeText)" XPATH_BODY = "$x('//div[@class='html-content']/p[not(@class)]/text()').map(x => x.wholeText)" #Función para extraer los links de las noticias def parse_home(): #Envolver el código dentro de un bloque try except try: response = requests.get(HOME_URL) if response.status_code == 200: home = response.content.decode('utf-8') parsed = html.fromstring(home) links_to_notices = parsed.xpath(XPATH_LINK_TO_ARTICLE) print(links_to_notices) else: raise ValueError(f'Error: {response.status_code}') except ValueError as ve: print(ve) def run(): parse_home() if __name__ == '__main__': run()
报错信息:
Traceback (most recent call last): File "D:\Andres\PLATZI\Curso_Fundamentos_WebScrapping\LaRepublica_Scrapper\scraper.py", line 30, in <module> run() File "D:\Andres\PLATZI\Curso_Fundamentos_WebScrapping\LaRepublica_Scrapper\scraper.py", line 27, in run parse_home() File "D:\Andres\PLATZI\Curso_Fundamentos_WebScrapping\LaRepublica_Scrapper\scraper.py", line 19, in parse_home links_to_notices = parsed.xpath(XPATH_LINK_TO_ARTICLE) #Obtener lista de links con comandos xtpath del documento contenido en parsed File "src\lxml\etree.pyx", line 1599, in lxml.etree._Element.xpath File "src\lxml\xpath.pxi", line 305, in lxml.etree.XPathElementEvaluator.__call__ File "src\lxml\xpath.pxi", line 225, in lxml.etree._XPathEvaluatorBase._handle_result lxml.etree.XPathEvalError: Invalid expression
问题原因
你直接复制了Chrome浏览器控制台的XPath调用代码,但这些是浏览器DevTools专属的JavaScript语法,lxml的xpath方法只支持标准XPath表达式:
$x()是Chrome提供的快捷查询函数,不属于标准XPath.map(x => x.value)是JavaScript的数组处理逻辑,lxml的xpath会直接返回匹配结果的列表- 字符串嵌套单引号导致语法冲突(比如
'//div[@class='newsV_Title_Img']'里的单引号会被Python解析为字符串结束符)
修正后的代码
import requests import lxml.html as html HOME_URL = 'https://www.larepublica.co/' # 使用标准XPath表达式,修正引号嵌套问题 XPATH_LINK_TO_ARTICLE = "//div[@class='newsV_Title_Img' or @class='V_Title']/text-fill/a/@href" XPATH_TITLE = "//full-text/span/text()" XPATH_SUMMARY = "//div[@class='lead']/p/text()" XPATH_BODY = "//div[@class='html-content']/p[not(@class)]/text()" def parse_home(): try: response = requests.get(HOME_URL) if response.status_code == 200: home = response.content.decode('utf-8') parsed = html.fromstring(home) links_to_notices = parsed.xpath(XPATH_LINK_TO_ARTICLE) print(links_to_notices) else: raise ValueError(f'Error: {response.status_code}') except ValueError as ve: print(ve) def run(): parse_home() if __name__ == '__main__': run()
关键修正点
- 移除
$x()包裹和.map()逻辑,直接编写纯XPath表达式 - 用双引号包裹整个XPath字符串,避免内部属性值的单引号冲突(也可以用单引号包裹整个字符串,内部属性值改用双引号)
- lxml的xpath方法会自动返回匹配到的属性值或文本内容的列表,无需额外处理
内容的提问来源于stack exchange,提问作者Andres Cardona
相关产品推荐
相关产品推荐

