You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网页子标签信息返回空列表问题排查

问题根因

你的代码返回空列表的核心原因是CSS选择器的类名归属匹配错误:

  • 类名ergov3-txtannonce属于包裹单条房源信息的外层<div>标签,并非内部<span>、<cite>标签的类名
  • 你写的soup.find_all("span", class_='ergov3-txtannonce')逻辑是查找带ergov3-txtannonce类名的span标签,页面中不存在这类元素,因此返回空结果
修正方案

正确的提取逻辑是先定位所有class为ergov3-txtannonce的房源外层容器,再逐个从容器内提取对应子标签的内容,同时对文本中多余的换行、连续空白做规整处理,完整代码如下:

from bs4 import BeautifulSoup
import requests

# 此处替换为你实际请求得到的页面响应文本
# 例:html = requests.get(target_url).text
soup = BeautifulSoup(html, "html.parser")

ads = []
cities = []

# 遍历所有房源卡片容器
for card in soup.find_all("div", class_="ergov3-txtannonce"):
    # 提取span内的房源描述,合并多余空白字符
    ad_content = " ".join(card.find("span").get_text(strip=True).split())
    ads.append(ad_content)
    # 提取cite内的城市信息
    city_content = card.find("cite").get_text(strip=True)
    cities.append(city_content)
运行结果

执行上述代码后将得到符合预期的输出:

["House 3 pièces, 74 m²", "Appartement 3 pièces, 64 m²", "House 4 pièces, 81 m²"]
["New York (11111)", "Los Angeles (22222)", "Chicago (33333)"]

内容的提问来源于stack exchange,提问作者ladybug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 11:18:16