使用Python从Redfin页面下载CSV失败,请求排查原因
问题排查与解决办法
1. 反爬拦截导致页面缺失下载链接
Redfin会识别非浏览器发起的请求,默认requests.get的请求头过于简单,会被网站拦截,返回的页面里不会包含下载按钮元素。解决方法是添加模拟浏览器的请求头:
修改请求部分代码,补充完整请求头信息:
from bs4 import BeautifulSoup import requests import csv user_input = input("Enter a zipcode: ") url = f"https://www.redfin.com/zipcode/{user_input}/filter/include=sold-3mo" # 添加模拟浏览器的请求头,避免被反爬识别 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3' } response = requests.get(url, headers=headers) # 先校验请求是否成功,状态码200为正常 if response.status_code != 200: print(f"Error: 请求失败,状态码{response.status_code}") exit() soup = BeautifulSoup(response.content, "html.parser") # 调整下载链接的定位逻辑:Redfin的下载按钮通常带有class="downloadLink"标识 download_link = soup.find("a", class_="downloadLink") # 也可以通过aria-label属性辅助查找 # download_link = soup.find("a", attrs={"aria-label": lambda x: x and "Download" in x}) if download_link: href = download_link["href"] # 处理相对路径,拼接完整URL if not href.startswith('http'): href = "https://www.redfin.com" + href # 下载CSV时同样带上请求头 csv_response = requests.get(href, headers=headers) with open("data.csv", "wb") as f: f.write(csv_response.content) print("CSV文件下载成功") else: print("Error: Unable to find link to CSV file")
2. 动态内容加载导致静态HTML无链接
如果添加请求头后仍找不到链接,说明下载按钮是通过JavaScript动态生成的,静态HTML里没有对应元素。此时可以用Selenium模拟浏览器完整加载页面:
先安装依赖:pip install selenium,并下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配)
示例代码:
from selenium import webdriver from selenium.webdriver.common.by import By import time import requests user_input = input("Enter a zipcode: ") url = f"https://www.redfin.com/zipcode/{user_input}/filter/include=sold-3mo" # 初始化Chrome浏览器驱动(需确保驱动路径正确,或已配置到系统环境变量) driver = webdriver.Chrome() driver.get(url) # 等待页面完全加载 time.sleep(5) try: # 定位下载按钮 download_button = driver.find_element(By.CLASS_NAME, "downloadLink") href = download_button.get_attribute("href") # 下载CSV文件 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get(href, headers=headers) with open("data.csv", "wb") as f: f.write(response.content) print("CSV文件下载成功") except: print("Error: Unable to find link to CSV file") finally: # 关闭浏览器 driver.quit()
3. 注意事项
- Redfin有严格的反爬机制,频繁请求可能会被封禁IP,建议添加请求间隔,避免短时间内大量爬取。
- 请遵守Redfin的网站使用条款,不要将爬取的数据用于商业用途或违规场景。
内容的提问来源于stack exchange,提问作者ImBadAtThis
相关产品推荐
相关产品推荐

