使用BeautifulSoup提取Zillow网页span标签中$1,773数值的问题问询
问题原因
- 你使用的
Text-c11n-8-48-0__sc-aiai24-0 jLucLe类名是Zillow前端框架自动生成的动态类名,会跟随版本迭代、页面渲染场景变化,同时会被多个金额类元素复用,无法准确定位目标元素。 - 部分预估月度费用是页面加载后通过JS动态计算渲染的,直接用requests请求拿到的静态HTML源码中不包含该数值,所以匹配不到目标结果。
解决方法
方法1:优化元素定位逻辑(优先尝试)
不要依赖动态类名定位,改为通过文本内容锚定目标模块,再提取对应数值:
from bs4 import BeautifulSoup import requests url = 'https://www.zillow.com/homedetails/4651-Genoa-St-Denver-CO-80249/13274183_zpid/' headers = {"User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:91.0) Gecko/20100101 Firefox/91.0"} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, "html.parser") # 先定位"Estimated monthly cost"标题span title_span = soup.find("span", string="Estimated monthly cost") if title_span: # 找同级的下一个span就是数值 cost_span = title_span.find_next_sibling("span") print(cost_span.text.strip()) else: print("静态页面中未找到预估月度费用模块,需使用JS渲染方案")
方法2:处理JS动态渲染数据
如果方法1返回未找到,说明该数值是JS动态加载的,你可以用selenium模拟浏览器加载完整页面后再提取:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options url = 'https://www.zillow.com/homedetails/4651-Genoa-St-Denver-CO-80249/13274183_zpid/' chrome_options = Options() chrome_options.add_argument("--headless") # 无头模式不弹出浏览器 chrome_options.add_argument("user-agent=Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:91.0) Gecko/20100101 Firefox/91.0") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面加载完成,可根据需要调整等待时长 driver.implicitly_wait(10) soup = BeautifulSoup(driver.page_source, "html.parser") title_span = soup.find("span", string="Estimated monthly cost") if title_span: cost_span = title_span.find_next_sibling("span") print(cost_span.text.strip()) driver.quit()
内容的提问来源于stack exchange,提问作者Batool
相关产品推荐
相关产品推荐

