You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Firefox Web Inspector生成的XPath在xmllint中无法生效问题

问题描述

目标

提取Wikidata查询输出结果

执行命令

curl https://query.wikidata.org/#SELECT%20DISTINCT%20%3Fitem%20%3FitemLabel%20WHERE%20%7B%0A%20%20SERVICE%20wikibase%3Alabel%20%7B%20bd%3AserviceParam%20wikibase%3Alanguage%20%22%5BAUTO_LANGUAGE%5D%2Cmul%2Cen%22.%20%7D%0A%20%20%7B%0A%20%20%20%20SELECT%20DISTINCT%20%3Fitem%20WHERE%20%7B%0A%20%20%20%20%20%20%3Fitem%20p%3AP1417%20%3Fstatement0.%0A%20%20%20%20%20%20%3Fstatement0%20ps%3AP1417%20%22topic%2FJacobs-Room%22.%0A%20%20%20%20%7D%0A%20%20%20%20LIMIT%20100%0A%20%20%7D%0A%7D > tmp.txt && xmllint tmp.txt --html --xpath "/html/body/div[2]/div[4]/div/div[1]/div[2]/div[2]/table/tbody/tr/td[1]/a[2]"

错误信息

% Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100 18033    0 18033    0     0  35974      0 --:--:-- --:--:-- --:--:-- 35994
tmp.txt:1: HTML parser error : Tag nav invalid
ueryservice container-fluid"><div class="row"><nav class="navbar navbar-default"
                                                                               ^
tmp.txt:1: HTML parser error : htmlParseEntityRef: expecting ';'
="https://www.mediawiki.org/w/index.php?title=Talk:Wikidata_Query_Service&action
                                                                               ^
tmp.txt:1: HTML parser error : htmlParseEntityRef: expecting ';'
.mediawiki.org/w/index.php?title=Talk:Wikidata_Query_Service&action=edit&section
                                                                               ^
tmp.txt:1: HTML parser error : Tag nav invalid
/div></div></noscript><div class="row"><nav class="navbar navbar-default result"
                                                                               ^
tmp.txt:1: element button: validity error : ID open-example already defined
 btn-default" id="open-example" data-toggle="modal" data-target="#QueryExamples"
                                                                               ^
tmp.txt:1: element span: validity error : ID examples-label already defined
r-open-o"></span> <span data-i18n="wdqs-app-button-examples" id="examples-label"
                                                                               ^
XPath set is empty

补充说明

  • 已通过Firefox开发者工具复制目标元素的精确XPath

环境信息

alinuxchap@libertus-desktop:~ $ hostnamectl
 Static hostname: libertus-desktop
       Icon name: computer
      Machine ID: ########
         Boot ID: ########
Operating System: Debian GNU/Linux 12 (bookworm)  
          Kernel: Linux 6.12.25+rpt-rpi-v8
    Architecture: arm64
alinuxchap@libertus-desktop:~ $ xmllint --version
xmllint: using libxml version 20914
   compiled with: Threads Tree Output Push Reader Patterns Writer SAXv1 FTP HTTP DTDValid HTML Legacy C14N Catalog XPath XPointer XInclude Iconv ICU ISO8859X Unicode Regexps Automata Schemas Schematron Modules Debug Zlib Lzma 
alinuxchap@libertus-desktop:~ $ curl --version
curl 7.88.1 (aarch64-unknown-linux-gnu) libcurl/7.88.1 OpenSSL/3.0.16 zlib/1.2.13 brotli/1.0.9 zstd/1.5.4 libidn2/2.3.3 libpsl/0.21.2 (+libidn2/2.3.3) libssh2/1.10.0 nghttp2/1.52.0 librtmp/2.3 OpenLDAP/2.5.13
Release-Date: 2023-02-20, security patched: 7.88.1-10+deb12u12
Protocols: dict file ftp ftps gopher gophers http https imap imaps ldap ldaps mqtt pop3 pop3s rtmp rtsp scp sftp smb smbs smtp smtps telnet tftp
Features: alt-svc AsynchDNS brotli GSS-API HSTS HTTP2 HTTPS-proxy IDN IPv6 Kerberos Largefile libz NTLM NTLM_WB PSL SPNEGO SSL threadsafe TLS-SRP UnixSockets zstd
alinuxchap@libertus-desktop:~ $ bash --version
GNU bash, version 5.2.15(1)-release (aarch64-unknown-linux-gnu)
Copyright (C) 2022 Free Software Foundation, Inc.
License GPLv3+: GNU GPL version 3 or later <http://gnu.org/licenses/gpl.html>

This is free software; you are free to change and redistribute it.
There is NO WARRANTY, to the extent permitted by law.

解决方法

  1. 问题根源

    • URL中#后的查询参数仅在浏览器端处理,curl只会获取Wikidata查询页面的框架,不会触发查询执行,因此页面中不存在目标结果表格。
    • xmllint对现代HTML标签(如<nav>)支持有限,且页面存在未转义的&符号,导致解析报错。
  2. 推荐方案:调用Wikidata查询API
    直接通过API获取结构化数据,避免解析HTML的麻烦,示例命令:

    curl -G "https://query.wikidata.org/sparql" \
         --data-urlencode "query=SELECT DISTINCT ?item ?itemLabel WHERE {
           SERVICE wikibase:label { bd:serviceParam wikibase:language \"[AUTO_LANGUAGE],mul,en\" . }
           {
             SELECT DISTINCT ?item WHERE {
               ?item p:P1417 ?statement0.
               ?statement0 ps:P1417 \"topic/Jacobs-Room\".
             }
             LIMIT 100
           }
         }" \
         --data-urlencode "format=json"
    

    该命令会返回JSON格式的查询结果,便于后续处理。

  3. 若需处理HTML的替代方案

    • 使用更适配现代HTML的解析工具,如pup或Python的BeautifulSoup库。
    • 手动在浏览器中执行查询后保存页面,再进行本地解析,但这种方式稳定性差,不推荐。

内容的提问来源于stack exchange,提问作者Signor Pizza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 01:20:01