You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup爬虫首次运行正常 后续无输出问题求助

问题根源
  • 目标站点配置了基础反爬策略,会识别requests默认的请求头标识python-requests/版本号,首次访问未触发拦截规则,后续访问被判定为爬虫后返回拦截页面,页面中不存在你指定的li.apotheke类节点,因此选择器匹配不到任何内容,无输出。
  • 短时间频繁请求也会触发站点临时访问限制。
修复代码
from bs4 import BeautifulSoup
import requests
import time

# 模拟正常Chrome浏览器的请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

URL = "https://www.medizinfuchs.de/?params%5Bsearch%5D=11484834&params%5Bsearch_cat%5D=1"
# 添加超时和请求头
page = requests.get(URL, headers=headers, timeout=10)
# 校验请求是否成功
page.raise_for_status()

soup = BeautifulSoup(page.content, "html.parser")

prices = []
names = []
for price in soup.select('li.apotheke div.price'):
    prices.append(float(price.text.strip(' \t\n€').replace(',', '.')))
for name in soup.select('li.apotheke a.name'):
    names.append(name.text.strip(' \t\n'))

# 按预期格式输出
for p in prices:
    print(p)
print(' '.join(map(str, prices)) + ' ' + ' '.join(names))

# 两次请求间隔3秒以上,避免触发频率限制
time.sleep(3)
关键修改说明
  • 新增浏览器UA请求头,绕过基础爬虫识别规则
  • 增加请求状态码校验,请求失败时直接抛出异常便于定位问题
  • 调整数据收集和输出逻辑,完全匹配你需要的输出格式
  • 增加请求间隔配置,避免短时间多次访问触发站点频率限制

内容的提问来源于stack exchange,提问作者Justin Hansen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 12:06:05