You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何限制Python网页爬虫仅抓取含/review/前缀的特定URL

过滤以/review/开头的URL解决方案

你只需要在循环里加个判断,检查抓取到的URL是否以/review/开头,同时要注意处理href为空的情况(避免报错)。

修改后的完整代码:

import requests
from bs4 import BeautifulSoup

req = requests.get("https://extracteddomain-unsureifitbreaksrules.com")

soup = BeautifulSoup(req.text, 'html.parser')

urls = []
for link in soup.find_all('a'):
    href = link.get('href')
    # 先判断href不为空,再检查是否以/review/开头
    if href and href.startswith('/review/'):
        urls.append(href)
        print(href)

# 最后可以查看过滤后的完整列表
print("过滤后的URL列表:", urls)

关键说明:

  • link.get('href')可能返回None(部分a标签没有href属性),所以先判断href是否存在,再用startswith('/review/')检查开头
  • 符合条件的URL会被添加到urls列表并打印,实现列表净化的需求

如果想让代码更简洁,也可以用列表推导式一行完成URL收集:

urls = [link.get('href') for link in soup.find_all('a') if link.get('href') and link.get('href').startswith('/review/')]

内容的提问来源于stack exchange,提问作者MHPYTHONQ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 21:59:58