You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

排查Bash脚本中XPath无法匹配网页元素的问题

Idealista房源抓取:XPath无法匹配元素的解决方案

问题背景

使用Bash脚本结合Xidel抓取idealista网站房源数据时,脚本中的XPath无法匹配目标页面(https://www.idealista.it/vendita-case/fucecchio-firenze/lista-18.htm)的房源元素,无法提取价格、面积、链接、描述信息。尝试修改根节点XPath仍未生效。

原脚本代码

# Define variables for the URL and browser
sGDomain="idealista"
sGCitta="fucecchio-firenze"
sGTypo="vendita-case"
iGPagina=1

# Start of the loop
while :; do

    # Build the URL with the iGPagina variable
    url="https://www.$sGDomain.it/$sGTypo/$sGCitta/lista-$iGPagina.htm"
    #echo "$url"
    
    # Get the HTML content of the page
    html_content=$(curl -s -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:89.0) Gecko/20100101 Firefox/89.0" "$url")

    echo "$html_content" > htmlcompleto.txt
    
    # Check if the error string is not present in the HTML content
    if [[ ! $html_content =~ "Successiva" ]]; then
        break  # Exit the loop if the error string is not present
    fi
    
    # Use xidel to extract the ads
    xidel_output=$(xidel --silent --xpath '
        //div[contains(@class, "item-info-container")] ! string-join(
            (
                ( "price=" || normalize-space(.//span[contains(@class, "item-price")]/text()[1]) ),
                ( "size="  || normalize-space(.//span[contains(@class, "item-detail") and contains(text(), "m2")]) ),
                ( "link="  || normalize-space(.//a[contains(@class, "item-link")]/@href) ),
                ( "desc="  || normalize-space(.//p[contains(@class, "ellipsis")]) )
            ),
            codepoints-to-string(9)
        )
    ' -)

    # Check if the temporary file exists and delete it if present
    if [ -f "temp.txt" ]; then
        rm temp.txt
    fi

    # Replace special characters from "desc=" to the end of each line in semi.txt
    echo "$xidel_output" | sed -e "s/desc=\(.*\)\(['\"]\)/desc=\1 /g" > semi.txt

    sed -i 's/\([0-9]\{1,\}\)\.\([0-9]\{1,\}\),[0-9]\{2\}/\1\2/g' semi.txt
    sed -i 's/m²//g' semi.txt

    # Concatenate semi.txt with debugtxt.txt for debugging purposes
    cat semi.txt >> debugtxt.txt

    # Connect to the SQLite database
    db_file="immo.db"
 
    # Loop through the lines and insert them into the SQLite database
    while IFS= read -r line; do
        # Extract price, size, link, and description values from the lines using awk
        prezzo=$(echo "$line" | awk -F 'price=' '{print $2}' | awk -F 'size=' '{print $1}')
        size=$(echo "$line" | awk -F 'size=' '{print $2}' | awk -F 'link=' '{print $1}')
        link=$(echo "$line" | awk -F 'link=' '{print $2}' | awk -F 'desc=' '{print $1}')
        descrizione=$(echo "$line" | awk -F 'desc=' '{print $2}')

        # Determine if the description contains "asta"
        if [[ $descrizione =~ "asta" ]]; then
            asta=1
        else
            asta=0
        fi

        # Insert the data into the SQLite database
        sqlite3 "$db_file" "INSERT INTO $sGDomain (prezzo, link, descrizione, metratura, asta) VALUES ('$prezzo', '$link', '$descrizione', '$size', $asta)"
    done < semi.txt

    # Increment the iGPagina variable for the next iteration
    ((iGPagina++))
done

尝试过的XPath表达式

//div[contains(@class, "item-info-container")] ! string-join(
  (
    ( "price=" || normalize-space(.//span[contains(@class, "item-price")]/text()[1]) ),
    ( "size="  || normalize-space(.//span[contains(@class, "item-detail") and contains(text(), "m2")]) ),
    ( "link="  || normalize-space(.//a[contains(@class, "item-link")]/@href) ),
    ( "desc="  || normalize-space(.//p[contains(@class, "ellipsis")]) )
  ),
  codepoints-to-string(9)
)

以及将根节点改为//div[contains(@class, "items-container items-list")]的版本。


问题分析

核心原因是idealista的页面DOM结构已更新,原XPath依赖的class名称(如item-info-container、ellipsis)已不再对应当前页面的元素,同时基础反爬机制可能导致curl获取的页面内容不完整。

解决方案

1. 更新XPath匹配当前页面结构

通过查看目标页面DOM,当前房源相关元素的正确定位如下,替换原脚本中的XPath:

//div[contains(@class, "listing-item-container")] ! string-join(
  (
    "price=" || normalize-space(.//span[contains(@class, "item-price")]/text()[1]),
    "size="  || normalize-space(.//span[contains(@class, "item-detail") and contains(text(), "m²")]),
    "link="  || normalize-space(.//a[contains(@class, "item-link")]/@href),
    "desc="  || normalize-space(.//p[contains(@class, "item-description")])
  ),
  codepoints-to-string(9)
)

2. 确保curl能获取完整页面内容

  • 升级UA头到最新浏览器版本,避免被反爬识别:
    html_content=$(curl -s -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:120.0) Gecko/20100101 Firefox/120.0" "$url")
    
  • 检查生成的htmlcompleto.txt,如果内容缺失,可尝试添加浏览器会话cookie(先在浏览器访问页面,导出cookie后用--cookie参数传入curl)。

3. 修复循环终止逻辑

原脚本依赖"Successiva"文本判断下一页存在性,该文本可能已被修改或隐藏,改为检查Xidel输出是否为空来终止循环:

# 替换原有的break判断
if [[ -z "$xidel_output" ]]; then
    break
fi

4. 修复SQL注入风险

原脚本直接拼接变量到SQL语句,存在安全风险,改用参数绑定:

sqlite3 "$db_file" "INSERT INTO $sGDomain (prezzo, link, descrizione, metratura, asta) VALUES (?, ?, ?, ?, ?)" "$prezzo" "$link" "$descrizione" "$size" "$asta"

内容的提问来源于stack exchange,提问作者eofficecom com

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 18:45:54