You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests模块抓取网页指定城市链接失败求助

问题:抓取网页中「By City」下的城市链接失败

我尝试使用requests模块抓取某网页中「By City」标题下的城市链接,请求返回状态码为200,但脚本执行后控制台无任何输出。

我的代码:

import requests
from bs4 import BeautifulSoup

link = 'https://www.nursinghomes.com/'

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36',
    'Referer': 'https://www.nursinghomes.com/',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
    'Accept-Encoding': 'gzip, deflate, br',
    'Accept-Language': 'en-US,en;q=0.9'
}

res = requests.get(link,headers=headers)
soup = BeautifulSoup(res.text,"html.parser")
for item in soup.select("li > strong > a[class^='text-link']"):
    print(item.get("href"))

预期部分结果:

/tx/austin/
/md/baltimore/
/ny/bronx/
/ny/brooklyn/
/il/chicago/

解决方案:

  • 问题根源:原选择器li > strong > a[class^='text-link']与当前页面的HTML结构不匹配,导致无法定位到目标链接。

  • 修复步骤:

    1. 先定位「By City」标题所在元素,再匹配其下方的城市列表,代码调整如下:
      import requests
      from bs4 import BeautifulSoup
      
      link = 'https://www.nursinghomes.com/'
      
      headers = {
          'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36',
          'Referer': 'https://www.nursinghomes.com/',
          'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
          'Accept-Encoding': 'gzip, deflate, br',
          'Accept-Language': 'en-US,en;q=0.9'
      }
      
      res = requests.get(link, headers=headers)
      soup = BeautifulSoup(res.text, "html.parser")
      
      # 定位By City标题
      by_city_title = soup.find('h3', string='By City')
      if by_city_title:
          # 获取标题下方的城市列表容器
          city_list_container = by_city_title.find_next_sibling('ul')
          if city_list_container:
              # 抓取列表中的所有城市链接
              for city_link in city_list_container.select('li > a.text-link'):
                  print(city_link.get('href'))
      
    2. 验证页面结构:若上述代码仍无输出,可先打印soup.prettify()查看返回的HTML内容,确认「By City」区块是否存在,以及链接的实际嵌套结构,再调整选择器。
    3. 动态加载排查:如果页面内容是通过JavaScript动态渲染的,requests无法获取动态加载的内容,此时需改用Selenium等工具模拟浏览器加载页面。

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:33:41