You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyCharm中Python爬虫返回码0却无控制台输出问题求助

Zillow爬虫无输出问题排查与修复

核心原因

Zillow有严格的反爬机制,直接用requests.get()请求会被拦截,返回的页面并非正常房源列表页,导致BeautifulSoup找不到目标元素,循环无法执行,自然没有输出。另外,目标元素的class名称可能已随页面结构更新而失效,也是常见诱因。

修复步骤

  1. 添加请求头模拟浏览器访问
    给requests.get()补充User-Agent等请求头,让服务器识别为正常浏览器访问。可以从自己浏览器的开发者工具(F12)中复制真实的User-Agent值。

  2. 验证请求有效性
    先打印请求状态码和页面片段,确认请求是否成功、返回内容是否正常:

    print(page.status_code)
    print(page.text[:500])  # 打印前500字符查看页面内容
    

    若状态码为403、503,说明被反爬拦截;若返回验证码页面,需进一步处理(如使用代理或验证码识别工具)。

  3. 更新元素选择器
    Zillow页面结构频繁更新,原代码中的list-card-info、list-card-addr等class大概率已失效。通过浏览器开发者工具定位当前房源卡片的真实DOM结构,获取最新的选择器。

修改后的示例代码

from bs4 import BeautifulSoup
import requests

# 替换为你自己浏览器的User-Agent
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

url = "https://www.zillow.com/philadelphia-pa/rentals/?searchQueryState=%7B%22pagination%22%3A%7B%7D%2C%22usersSearchTerm%22%3A%22Philadelphia%2C%20PA%22%2C%22mapBounds%22%3A%7B%22west%22%3A-75.2058476090698%2C%22east%22%3A-75.17623602154539%2C%22south%22%3A39.9520661821946%2C%22north%22%3A39.97380838759173%7D%2C%22regionSelection%22%3A%5B%7B%22regionId%22%3A13271%2C%22regionType%22%3A6%7D%5D%2C%22isMapVisible%22%3Afalse%2C%22filterState%22%3A%7B%22fsba%22%3A%7B%22value%22%3Afalse%7D%2C%22fsbo%22%3A%7B%22value%22%3Afalse%7D%2C%22nc%22%3A%7B%22value%22%3Afalse%7D%2C%22fore%22%3A%7B%22value%22%3Afalse%7D%2C%22cmsn%22%3A%7B%22value%22%3Afalse%7D%2C%22auc%22%3A%7B%22value%22%3Afalse%7D%2C%22fr%22%3A%7B%22value%22%3Atrue%7D%2C%22ah%22%3A%7B%22value%22%3Atrue%7D%7D%2C%22isListVisible%22%3Atrue%2C%22mapZoom%22%3A15%7D"

page = requests.get(url, headers=headers)

# 检查请求状态
print(f"请求状态码: {page.status_code}")

if page.status_code == 200:
    soup = BeautifulSoup(page.content, 'html.parser')
    # 注意:以下class为示例,需根据当前页面实际结构修改
    lists = soup.find_all('div', class_="StyledPropertyCardDataWrapper-c11n-8-84-3__sc-1omp4c3-0")
    
    if not lists:
        print("未找到目标元素,请检查页面结构或选择器")
    else:
        for item in lists:
            title = item.find('a', class_="StyledPropertyCardDataArea-c11n-8-84-3__sc-yipmu-0")
            price = item.find('span', class_="PropertyCardWrapper__StyledPriceLine-srp__sc-16e8gqd-1")
            # 处理元素可能为None的情况
            info = [title.text.strip() if title else "无地址", price.text.strip() if price else "无价格"]
            print(info)
else:
    print(f"请求失败,状态码: {page.status_code}")

额外提示

  • 若添加请求头后仍被拦截,可补充Accept-Language、Referer等请求头,或使用代理IP分散请求来源。
  • 频繁请求易导致IP被封禁,建议添加请求间隔(time.sleep(1))控制请求频率。

内容的提问来源于stack exchange,提问作者7896

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 00:06:25