Firefox Web Inspector生成的XPath在xmllint中无法生效问题
问题描述
目标
提取Wikidata查询输出结果
执行命令
curl https://query.wikidata.org/#SELECT%20DISTINCT%20%3Fitem%20%3FitemLabel%20WHERE%20%7B%0A%20%20SERVICE%20wikibase%3Alabel%20%7B%20bd%3AserviceParam%20wikibase%3Alanguage%20%22%5BAUTO_LANGUAGE%5D%2Cmul%2Cen%22.%20%7D%0A%20%20%7B%0A%20%20%20%20SELECT%20DISTINCT%20%3Fitem%20WHERE%20%7B%0A%20%20%20%20%20%20%3Fitem%20p%3AP1417%20%3Fstatement0.%0A%20%20%20%20%20%20%3Fstatement0%20ps%3AP1417%20%22topic%2FJacobs-Room%22.%0A%20%20%20%20%7D%0A%20%20%20%20LIMIT%20100%0A%20%20%7D%0A%7D > tmp.txt && xmllint tmp.txt --html --xpath "/html/body/div[2]/div[4]/div/div[1]/div[2]/div[2]/table/tbody/tr/td[1]/a[2]"
错误信息
% Total % Received % Xferd Average Speed Time Time Time Current Dload Upload Total Spent Left Speed 100 18033 0 18033 0 0 35974 0 --:--:-- --:--:-- --:--:-- 35994 tmp.txt:1: HTML parser error : Tag nav invalid ueryservice container-fluid"><div class="row"><nav class="navbar navbar-default" ^ tmp.txt:1: HTML parser error : htmlParseEntityRef: expecting ';' ="https://www.mediawiki.org/w/index.php?title=Talk:Wikidata_Query_Service&action ^ tmp.txt:1: HTML parser error : htmlParseEntityRef: expecting ';' .mediawiki.org/w/index.php?title=Talk:Wikidata_Query_Service&action=edit§ion ^ tmp.txt:1: HTML parser error : Tag nav invalid /div></div></noscript><div class="row"><nav class="navbar navbar-default result" ^ tmp.txt:1: element button: validity error : ID open-example already defined btn-default" id="open-example" data-toggle="modal" data-target="#QueryExamples" ^ tmp.txt:1: element span: validity error : ID examples-label already defined r-open-o"></span> <span data-i18n="wdqs-app-button-examples" id="examples-label" ^ XPath set is empty
补充说明
- 已通过Firefox开发者工具复制目标元素的精确XPath
环境信息
alinuxchap@libertus-desktop:~ $ hostnamectl Static hostname: libertus-desktop Icon name: computer Machine ID: ######## Boot ID: ######## Operating System: Debian GNU/Linux 12 (bookworm) Kernel: Linux 6.12.25+rpt-rpi-v8 Architecture: arm64 alinuxchap@libertus-desktop:~ $ xmllint --version xmllint: using libxml version 20914 compiled with: Threads Tree Output Push Reader Patterns Writer SAXv1 FTP HTTP DTDValid HTML Legacy C14N Catalog XPath XPointer XInclude Iconv ICU ISO8859X Unicode Regexps Automata Schemas Schematron Modules Debug Zlib Lzma alinuxchap@libertus-desktop:~ $ curl --version curl 7.88.1 (aarch64-unknown-linux-gnu) libcurl/7.88.1 OpenSSL/3.0.16 zlib/1.2.13 brotli/1.0.9 zstd/1.5.4 libidn2/2.3.3 libpsl/0.21.2 (+libidn2/2.3.3) libssh2/1.10.0 nghttp2/1.52.0 librtmp/2.3 OpenLDAP/2.5.13 Release-Date: 2023-02-20, security patched: 7.88.1-10+deb12u12 Protocols: dict file ftp ftps gopher gophers http https imap imaps ldap ldaps mqtt pop3 pop3s rtmp rtsp scp sftp smb smbs smtp smtps telnet tftp Features: alt-svc AsynchDNS brotli GSS-API HSTS HTTP2 HTTPS-proxy IDN IPv6 Kerberos Largefile libz NTLM NTLM_WB PSL SPNEGO SSL threadsafe TLS-SRP UnixSockets zstd alinuxchap@libertus-desktop:~ $ bash --version GNU bash, version 5.2.15(1)-release (aarch64-unknown-linux-gnu) Copyright (C) 2022 Free Software Foundation, Inc. License GPLv3+: GNU GPL version 3 or later <http://gnu.org/licenses/gpl.html> This is free software; you are free to change and redistribute it. There is NO WARRANTY, to the extent permitted by law.
解决方法
问题根源
- URL中
#后的查询参数仅在浏览器端处理,curl只会获取Wikidata查询页面的框架,不会触发查询执行,因此页面中不存在目标结果表格。 - xmllint对现代HTML标签(如
<nav>)支持有限,且页面存在未转义的&符号,导致解析报错。
- URL中
推荐方案:调用Wikidata查询API
直接通过API获取结构化数据,避免解析HTML的麻烦,示例命令:curl -G "https://query.wikidata.org/sparql" \ --data-urlencode "query=SELECT DISTINCT ?item ?itemLabel WHERE { SERVICE wikibase:label { bd:serviceParam wikibase:language \"[AUTO_LANGUAGE],mul,en\" . } { SELECT DISTINCT ?item WHERE { ?item p:P1417 ?statement0. ?statement0 ps:P1417 \"topic/Jacobs-Room\". } LIMIT 100 } }" \ --data-urlencode "format=json"该命令会返回JSON格式的查询结果,便于后续处理。
若需处理HTML的替代方案
- 使用更适配现代HTML的解析工具,如
pup或Python的BeautifulSoup库。 - 手动在浏览器中执行查询后保存页面,再进行本地解析,但这种方式稳定性差,不推荐。
- 使用更适配现代HTML的解析工具,如
内容的提问来源于stack exchange,提问作者Signor Pizza
相关产品推荐
相关产品推荐

