You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python批量抓取网页中律师个人链接的详情数据?

批量抓取Chambers亚太区律师信息的实现方案

核心思路

先从列表页批量提取所有律师的基础信息(姓名、所属律所、排名)和个人主页链接,再通过循环遍历这些链接批量抓取详情页内容,同时通过优化请求方式提升效率。


步骤1:批量提取列表页的律师链接与基础信息

一次性爬取列表页的所有律师条目,把需要的基础信息和详情页链接收集起来,避免重复请求列表页。以下是Python示例代码(需根据页面实际HTML结构调整选择器):

import requests
from bs4 import BeautifulSoup

def extract_lawyer_list():
    list_url = "https://chambers.com/all-lawyers-asia-pacific-8"
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    resp = requests.get(list_url, headers=headers)
    soup = BeautifulSoup(resp.text, "html.parser")
    
    lawyer_data = []
    # 替换为页面实际的律师条目选择器
    for item in soup.select(".search-result-item"):
        name = item.select_one(".lawyer-name").get_text(strip=True)
        firm = item.select_one(".lawyer-firm").get_text(strip=True)
        ranking = item.select_one(".rank-badge").get_text(strip=True)
        profile_href = item.select_one(".profile-link")["href"]
        # 补全绝对链接
        profile_url = f"https://chambers.com{profile_href}" if not profile_href.startswith("http") else profile_href
        
        lawyer_data.append({
            "name": name,
            "firm": firm,
            "ranking": ranking,
            "profile_url": profile_url
        })
    return lawyer_data

步骤2:循环批量抓取详情页

拿到所有律师的详情页链接后,用循环遍历请求每个链接,提取详情页信息。加入异常处理避免单个请求失败中断整个流程,同时加入延迟避免触发反爬:

import csv
import time
from concurrent.futures import ThreadPoolExecutor

def scrape_profile(lawyer):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    try:
        resp = requests.get(lawyer["profile_url"], headers=headers)
        resp.raise_for_status()
        soup = BeautifulSoup(resp.text, "html.parser")
        
        # 提取详情页信息,示例为联系方式和执业领域,需根据页面调整
        contact = soup.select_one(".lawyer-contact-info").get_text(strip=True) if soup.select_one(".lawyer-contact-info") else "无"
        practice_areas = [area.get_text(strip=True) for area in soup.select(".practice-area-tag")]
        
        lawyer.update({
            "contact": contact,
            "practice_areas": ", ".join(practice_areas)
        })
        time.sleep(0.8)  # 延迟控制请求频率
        return lawyer
    except Exception as e:
        print(f"抓取失败:{lawyer['name']} - {str(e)}")
        # 记录失败链接以便重试
        with open("failed_profiles.txt", "a", encoding="utf-8") as f:
            f.write(f"{lawyer['profile_url']}\n")
        return None

def batch_scrape(lawyer_list):
    # 用多线程提升效率,线程数根据网站反爬强度调整
    with ThreadPoolExecutor(max_workers=15) as executor:
        results = executor.map(scrape_profile, lawyer_list)
    
    # 过滤失败的条目,写入CSV
    valid_results = [res for res in results if res]
    with open("chambers_lawyers.csv", "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=["name", "firm", "ranking", "profile_url", "contact", "practice_areas"])
        writer.writeheader()
        writer.writerows(valid_results)

# 执行流程
if __name__ == "__main__":
    lawyers = extract_lawyer_list()
    batch_scrape(lawyers)

效率优化建议

  • 多线程/多进程:用ThreadPoolExecutor替代单线程循环,同时处理多个请求,能把抓取速度提升数倍(注意线程数不要超过20,避免被封)。
  • 请求缓存:用requests-cache库缓存已请求的页面,中途中断后无需重复请求已完成的页面。
  • 代理轮换:如果遇到IP限制,使用代理池轮换IP,避免被网站封禁。

注意事项

  • 先查看网站的robots.txt文件,确认允许抓取这类内容,避免违反网站规则。
  • 模拟正常用户行为:使用真实的User-Agent,避免固定间隔请求,可随机调整延迟时间(比如0.5-1.5秒)。
  • 数据分批存储:如果数据量太大,分批写入文件或数据库,避免内存占用过高。

内容的提问来源于stack exchange,提问作者Psychedelique23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 06:01:00