You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于JSON响应提取href属性及同链接多页面爬取方法

Python爬虫问题解决方案

一、提取a标签内的href路径值

两种实现方案可直接选用:

  • 正则匹配(无额外依赖,适配固定HTML结构)
    直接通过正则提取href属性的分组内容,无需安装第三方库,针对当前固定结构的a标签匹配准确率足够。
import re

href_re = re.compile(r'href="([^"]+)"')
for test in sample:
    action_html = test['actions']
    match_ret = href_re.search(action_html)
    if match_ret:
        detail_path = match_ret.group(1)
        print(detail_path)
  • HTML解析提取(结构兼容性更强)
    使用BeautifulSoup解析HTML片段提取属性,可避免HTML属性顺序变动导致的正则匹配失效,先执行pip install beautifulsoup4安装依赖,代码如下:
from bs4 import BeautifulSoup

for test in sample:
    action_html = test['actions']
    soup = BeautifulSoup(action_html, 'html.parser')
    a_tag = soup.find('a')
    if a_tag:
        detail_path = a_tag.get('href')
        print(detail_path)

二、同URL分页爬取实现方案

分页请求URL完全相同的场景,分页标识一般藏在请求参数、请求体、请求头或Cookie中,针对目标站点的实现逻辑如下:

  1. 先通过浏览器开发者工具抓包定位分页规则
    打开目标列表页,调出开发者工具的网络面板,依次切换第2、3页,对比不同分页的请求差异:
  • 优先排查GET查询参数:这类后端表格接口通常会带page、pageNum、start、length这类分页参数,分别代表当前页码、每页数据条数
  • 若GET参数无变化,再检查请求体:部分接口会把分页参数放在POST表单或JSON请求体中,对应加到请求的data或json参数中即可
  • 注意接口依赖的_csrf-frontendCookie存在有效期,爬取时如果返回400/403状态码,直接从浏览器复制最新的Cookie替换即可
  1. 分页爬取代码框架
import requests
import re
import time

headers = {
  'Accept': 'text/html, */*; q=0.01',
  'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8,pt;q=0.7',
  'Cache-Control': 'no-cache',
  'Connection': 'keep-alive',
  'Cookie': '_csrf-frontend=ccc4c9069d6ad3816ea693a980ecbebda2770e9448ffe9fed17cdf397a5e2851a%3A2%3A%7Bi%3A0%3Bs%3A14%3A%22_csrf-frontend%22%3Bi%3A1%3Bs%3A32%3A%22J3N0AJG6xybnGl91dfrlt-qMOk3hfbQ6%22%3B%7D',
  'Pragma': 'no-cache',
  'Referer': 'https://baroul-timis.ro/tabloul-avocatilor/',
  'Sec-Fetch-Dest': 'empty',
  'Sec-Fetch-Mode': 'cors',
  'Sec-Fetch-Site': 'same-origin',
  'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/102.0.0.0 Safari/537.36',
  'X-Requested-With': 'XMLHttpRequest',
  'sec-ch-ua': '" Not A;Brand";v="99", "Chromium";v="102", "Google Chrome";v="102"',
  'sec-ch-ua-mobile': '?0',
  'sec-ch-ua-platform': '"Windows"'
}

base_url = "https://baroul-timis.ro/get-av-data?param=toti-avocatii"
href_re = re.compile(r'href="([^"]+)"')
all_paths = []

# 爬取前先确认总页数,避免无效请求
max_page = 20
for page_num in range(1, max_page + 1):
    # 此处参数名根据抓包实际结果调整
    req_params = {
        "page": page_num
    }
    resp = requests.get(base_url, headers=headers, params=req_params, timeout=10)
    resp.raise_for_status()
    resp_json = resp.json()
    page_data = resp_json.get('data', [])
    # 当前页无数据时终止爬取
    if not page_data:
        break
    for item in page_data:
        action_html = item.get('actions', '')
        match_ret = href_re.search(action_html)
        if match_ret:
            all_paths.append(match_ret.group(1))
    # 添加延时,降低被封禁概率
    time.sleep(1.5)

print(f"累计获取详情页路径{len(all_paths)}条")

若抓包发现分页参数不在URL查询参数中,直接将参数传入requests的data(表单格式)或json(JSON格式)参数即可,核心逻辑为循环传入递增的页码参数,直到返回数据为空时终止爬取。

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 00:39:03