请求返回JSON而非HTML,无法用BeautifulSoup的.find方法解析求助
问题解决方法
你请求的/Providers/List是一个API接口,它返回的是JSON格式的结构化数据,而非渲染完成的HTML页面,所以用BeautifulSoup解析肯定无法使用.find()或.find_all()方法。
方案1:直接解析返回的JSON数据(推荐)
既然接口返回的是JSON,直接提取数据即可,不需要BeautifulSoup:
import requests import json api_url ='https://seniorcarefinder.com/Providers/List' headers= { "Content-Type":"application/json; charset=utf-8", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:108.0) Gecko/20100101 Firefox/108.0"} body_first_page={"Services":["Independent Living","Assisted Living","Long-Term Care / Skilled Nursing","Home Care (Non-Medical)","Home Health Care (Medicare-Certified)","Hospice","Adult Day Services","Active Adult Living"],"StarRatings":[],"PageNumber":1,"Location":"Colorado Springs, CO","Geography":{"Latitude":38.833882,"Longitude":-104.821363},"ProximityInMiles":30,"SortBy":"Verified"} res = requests.post(api_url, data=json.dumps(body_first_page), headers=headers) # 解析JSON响应 data = res.json() # 提取所需数据,比如机构列表 providers = data.get('Providers', []) for provider in providers: print(f"名称:{provider['Name']}") print(f"地址:{provider['Address']}") print(f"评分:{provider.get('StarRating', '无')}") # 根据需求提取其他字段
方案2:获取渲染后的HTML页面(若需完整页面结构)
如果一定要获取HTML来用BeautifulSoup解析,需要请求前端展示搜索结果的页面URL(而非API接口)。由于这类网站通常用JS动态加载数据,直接用requests可能无法获取完整渲染的HTML,需要使用浏览器自动化工具(如Selenium或Playwright)模拟用户操作:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 初始化浏览器(需提前安装对应浏览器驱动) driver = webdriver.Firefox() driver.get("https://seniorcarefinder.com") # 模拟搜索Colorado Springs, CO wait = WebDriverWait(driver, 10) location_input = wait.until(EC.presence_of_element_located((By.ID, "location-input"))) location_input.send_keys("Colorado Springs, CO") search_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.search-btn"))) search_btn.click() # 等待页面加载完成 wait.until(EC.presence_of_element_located((By.CLASS_NAME, "provider-card"))) # 获取页面HTML并解析 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 用BeautifulSoup提取内容 provider_cards = soup.find_all(class_="provider-card") for card in provider_cards: name = card.find(class_="provider-name").text.strip() address = card.find(class_="provider-address").text.strip() print(f"名称:{name}\n地址:{address}\n") driver.quit()
注意:使用浏览器自动化工具时,需遵守网站的
robots.txt规则和使用条款,避免过度请求导致IP被封禁。
内容的提问来源于stack exchange,提问作者Mohamed Hedeya
相关产品推荐
相关产品推荐

