基于JSON响应提取href属性及同链接多页面爬取方法
Python爬虫问题解决方案
一、提取a标签内的href路径值
两种实现方案可直接选用:
- 正则匹配(无额外依赖,适配固定HTML结构)
直接通过正则提取href属性的分组内容,无需安装第三方库,针对当前固定结构的a标签匹配准确率足够。
import re href_re = re.compile(r'href="([^"]+)"') for test in sample: action_html = test['actions'] match_ret = href_re.search(action_html) if match_ret: detail_path = match_ret.group(1) print(detail_path)
- HTML解析提取(结构兼容性更强)
使用BeautifulSoup解析HTML片段提取属性,可避免HTML属性顺序变动导致的正则匹配失效,先执行pip install beautifulsoup4安装依赖,代码如下:
from bs4 import BeautifulSoup for test in sample: action_html = test['actions'] soup = BeautifulSoup(action_html, 'html.parser') a_tag = soup.find('a') if a_tag: detail_path = a_tag.get('href') print(detail_path)
二、同URL分页爬取实现方案
分页请求URL完全相同的场景,分页标识一般藏在请求参数、请求体、请求头或Cookie中,针对目标站点的实现逻辑如下:
- 先通过浏览器开发者工具抓包定位分页规则
打开目标列表页,调出开发者工具的网络面板,依次切换第2、3页,对比不同分页的请求差异:
- 优先排查GET查询参数:这类后端表格接口通常会带
page、pageNum、start、length这类分页参数,分别代表当前页码、每页数据条数 - 若GET参数无变化,再检查请求体:部分接口会把分页参数放在POST表单或JSON请求体中,对应加到请求的
data或json参数中即可 - 注意接口依赖的
_csrf-frontendCookie存在有效期,爬取时如果返回400/403状态码,直接从浏览器复制最新的Cookie替换即可
- 分页爬取代码框架
import requests import re import time headers = { 'Accept': 'text/html, */*; q=0.01', 'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8,pt;q=0.7', 'Cache-Control': 'no-cache', 'Connection': 'keep-alive', 'Cookie': '_csrf-frontend=ccc4c9069d6ad3816ea693a980ecbebda2770e9448ffe9fed17cdf397a5e2851a%3A2%3A%7Bi%3A0%3Bs%3A14%3A%22_csrf-frontend%22%3Bi%3A1%3Bs%3A32%3A%22J3N0AJG6xybnGl91dfrlt-qMOk3hfbQ6%22%3B%7D', 'Pragma': 'no-cache', 'Referer': 'https://baroul-timis.ro/tabloul-avocatilor/', 'Sec-Fetch-Dest': 'empty', 'Sec-Fetch-Mode': 'cors', 'Sec-Fetch-Site': 'same-origin', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/102.0.0.0 Safari/537.36', 'X-Requested-With': 'XMLHttpRequest', 'sec-ch-ua': '" Not A;Brand";v="99", "Chromium";v="102", "Google Chrome";v="102"', 'sec-ch-ua-mobile': '?0', 'sec-ch-ua-platform': '"Windows"' } base_url = "https://baroul-timis.ro/get-av-data?param=toti-avocatii" href_re = re.compile(r'href="([^"]+)"') all_paths = [] # 爬取前先确认总页数,避免无效请求 max_page = 20 for page_num in range(1, max_page + 1): # 此处参数名根据抓包实际结果调整 req_params = { "page": page_num } resp = requests.get(base_url, headers=headers, params=req_params, timeout=10) resp.raise_for_status() resp_json = resp.json() page_data = resp_json.get('data', []) # 当前页无数据时终止爬取 if not page_data: break for item in page_data: action_html = item.get('actions', '') match_ret = href_re.search(action_html) if match_ret: all_paths.append(match_ret.group(1)) # 添加延时,降低被封禁概率 time.sleep(1.5) print(f"累计获取详情页路径{len(all_paths)}条")
若抓包发现分页参数不在URL查询参数中,直接将参数传入requests的
data(表单格式)或json(JSON格式)参数即可,核心逻辑为循环传入递增的页码参数,直到返回数据为空时终止爬取。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

