Python请求Zillow网站异常:配置HTTP Header仍未获取实际页面内容
问题解决:请求Zillow返回URL而非页面内容的修复方案
问题核心是初始化BeautifulSoup时传入了Zillow的URL字符串,而非requests请求返回的响应文本,导致BeautifulSoup无法解析实际页面内容,最终输出的是你传入的URL本身。
错误写法示例
import requests from bs4 import BeautifulSoup url = "https://www.zillow.com/" headers = { # 你的完整HTTP Header配置 "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 错误:传入URL而非响应文本 soup = BeautifulSoup(url, "html.parser") print(soup) # 输出结果为URL字符串
正确写法示例
import requests from bs4 import BeautifulSoup url = "https://www.zillow.com/" headers = { # 保留你的完整HTTP Header配置 "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 正确:传入requests响应的文本内容 soup = BeautifulSoup(response.text, "html.parser") # 可通过打印页面标题验证是否获取到实际内容 print(soup.title.string)
额外注意点
- 先验证请求是否成功:通过
response.status_code检查状态码,200表示请求正常;若返回403/503等状态码,可能需要调整Header或处理网站的反爬机制。 - BeautifulSoup的第一个参数必须是HTML/XML文本内容,URL仅用于发起requests请求,不能直接传给BeautifulSoup做解析。
内容的提问来源于stack exchange,提问作者Ibukunoluwa Omosehin
相关产品推荐
相关产品推荐

