排查Bash脚本中XPath无法匹配网页元素的问题
Idealista房源抓取:XPath无法匹配元素的解决方案
问题背景
使用Bash脚本结合Xidel抓取idealista网站房源数据时,脚本中的XPath无法匹配目标页面(https://www.idealista.it/vendita-case/fucecchio-firenze/lista-18.htm)的房源元素,无法提取价格、面积、链接、描述信息。尝试修改根节点XPath仍未生效。
原脚本代码
# Define variables for the URL and browser sGDomain="idealista" sGCitta="fucecchio-firenze" sGTypo="vendita-case" iGPagina=1 # Start of the loop while :; do # Build the URL with the iGPagina variable url="https://www.$sGDomain.it/$sGTypo/$sGCitta/lista-$iGPagina.htm" #echo "$url" # Get the HTML content of the page html_content=$(curl -s -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:89.0) Gecko/20100101 Firefox/89.0" "$url") echo "$html_content" > htmlcompleto.txt # Check if the error string is not present in the HTML content if [[ ! $html_content =~ "Successiva" ]]; then break # Exit the loop if the error string is not present fi # Use xidel to extract the ads xidel_output=$(xidel --silent --xpath ' //div[contains(@class, "item-info-container")] ! string-join( ( ( "price=" || normalize-space(.//span[contains(@class, "item-price")]/text()[1]) ), ( "size=" || normalize-space(.//span[contains(@class, "item-detail") and contains(text(), "m2")]) ), ( "link=" || normalize-space(.//a[contains(@class, "item-link")]/@href) ), ( "desc=" || normalize-space(.//p[contains(@class, "ellipsis")]) ) ), codepoints-to-string(9) ) ' -) # Check if the temporary file exists and delete it if present if [ -f "temp.txt" ]; then rm temp.txt fi # Replace special characters from "desc=" to the end of each line in semi.txt echo "$xidel_output" | sed -e "s/desc=\(.*\)\(['\"]\)/desc=\1 /g" > semi.txt sed -i 's/\([0-9]\{1,\}\)\.\([0-9]\{1,\}\),[0-9]\{2\}/\1\2/g' semi.txt sed -i 's/m²//g' semi.txt # Concatenate semi.txt with debugtxt.txt for debugging purposes cat semi.txt >> debugtxt.txt # Connect to the SQLite database db_file="immo.db" # Loop through the lines and insert them into the SQLite database while IFS= read -r line; do # Extract price, size, link, and description values from the lines using awk prezzo=$(echo "$line" | awk -F 'price=' '{print $2}' | awk -F 'size=' '{print $1}') size=$(echo "$line" | awk -F 'size=' '{print $2}' | awk -F 'link=' '{print $1}') link=$(echo "$line" | awk -F 'link=' '{print $2}' | awk -F 'desc=' '{print $1}') descrizione=$(echo "$line" | awk -F 'desc=' '{print $2}') # Determine if the description contains "asta" if [[ $descrizione =~ "asta" ]]; then asta=1 else asta=0 fi # Insert the data into the SQLite database sqlite3 "$db_file" "INSERT INTO $sGDomain (prezzo, link, descrizione, metratura, asta) VALUES ('$prezzo', '$link', '$descrizione', '$size', $asta)" done < semi.txt # Increment the iGPagina variable for the next iteration ((iGPagina++)) done
尝试过的XPath表达式
//div[contains(@class, "item-info-container")] ! string-join( ( ( "price=" || normalize-space(.//span[contains(@class, "item-price")]/text()[1]) ), ( "size=" || normalize-space(.//span[contains(@class, "item-detail") and contains(text(), "m2")]) ), ( "link=" || normalize-space(.//a[contains(@class, "item-link")]/@href) ), ( "desc=" || normalize-space(.//p[contains(@class, "ellipsis")]) ) ), codepoints-to-string(9) )
以及将根节点改为//div[contains(@class, "items-container items-list")]的版本。
问题分析
核心原因是idealista的页面DOM结构已更新,原XPath依赖的class名称(如item-info-container、ellipsis)已不再对应当前页面的元素,同时基础反爬机制可能导致curl获取的页面内容不完整。
解决方案
1. 更新XPath匹配当前页面结构
通过查看目标页面DOM,当前房源相关元素的正确定位如下,替换原脚本中的XPath:
//div[contains(@class, "listing-item-container")] ! string-join( ( "price=" || normalize-space(.//span[contains(@class, "item-price")]/text()[1]), "size=" || normalize-space(.//span[contains(@class, "item-detail") and contains(text(), "m²")]), "link=" || normalize-space(.//a[contains(@class, "item-link")]/@href), "desc=" || normalize-space(.//p[contains(@class, "item-description")]) ), codepoints-to-string(9) )
2. 确保curl能获取完整页面内容
- 升级UA头到最新浏览器版本,避免被反爬识别:
html_content=$(curl -s -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:120.0) Gecko/20100101 Firefox/120.0" "$url") - 检查生成的
htmlcompleto.txt,如果内容缺失,可尝试添加浏览器会话cookie(先在浏览器访问页面,导出cookie后用--cookie参数传入curl)。
3. 修复循环终止逻辑
原脚本依赖"Successiva"文本判断下一页存在性,该文本可能已被修改或隐藏,改为检查Xidel输出是否为空来终止循环:
# 替换原有的break判断 if [[ -z "$xidel_output" ]]; then break fi
4. 修复SQL注入风险
原脚本直接拼接变量到SQL语句,存在安全风险,改用参数绑定:
sqlite3 "$db_file" "INSERT INTO $sGDomain (prezzo, link, descrizione, metratura, asta) VALUES (?, ?, ?, ?, ?)" "$prezzo" "$link" "$descrizione" "$size" "$asta"
内容的提问来源于stack exchange,提问作者eofficecom com
相关产品推荐
相关产品推荐

