You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取法国国民议会网页数据并生成目标DataFrame?

解决法国国民议会网页抓取并生成DataFrame的问题

问题概述

需抓取法国国民议会特定格式的网页(示例:https://www.assemblee-nationale.fr/13/cri/2006-2007/20070152.asp),生成包含name(发言者姓名)和text(发言内容)列的Pandas DataFrame。此前尝试遇到两个关键错误:

  1. 使用BeautifulSoup时,误将ResultSet对象当作单个元素处理,触发AttributeError;
  2. 尝试以XML方式解析网页,出现「not well-formed (invalid token)」错误。

解决方案

1. 核心修正方向

  • 放弃XML解析:目标网页为HTML格式,存在XML不兼容的语法(如未闭合标签、特殊字符),改用HTML解析器即可规避格式错误;
  • 正确处理ResultSet:soup.find_all()返回的是元素集合,必须遍历每个元素提取数据,不能直接对集合调用元素级方法(如get_text())。

2. 可运行代码示例

import requests
from bs4 import BeautifulSoup
import pandas as pd

def scrape_assemblee_speeches(url):
    # 获取并解析网页
    response = requests.get(url)
    response.encoding = 'utf-8'  # 强制指定UTF-8编码避免乱码
    soup = BeautifulSoup(response.text, 'html.parser')  # 使用HTML解析器

    # 定位所有发言块:根据目标网页结构,发言内容包裹在class为"speech"的div中(需根据实际调整)
    speech_blocks = soup.find_all('div', class_='speech')
    
    speech_data = []
    for block in speech_blocks:
        # 提取发言者姓名:假设姓名在<b>标签内(实际需匹配网页标签)
        name_element = block.find('b')
        if not name_element:
            continue  # 跳过无姓名的发言块
        
        name = name_element.get_text(strip=True)
        # 提取发言内容:移除姓名后保留剩余文本
        full_text = block.get_text(strip=True)
        speech_text = full_text.replace(name, '', 1).strip()
        
        speech_data.append({'name': name, 'text': speech_text})
    
    # 转换为DataFrame
    return pd.DataFrame(speech_data)

# 测试示例网页
sample_url = "https://www.assemblee-nationale.fr/13/cri/2006-2007/20070152.asp"
speech_df = scrape_assemblee_speeches(sample_url)
print(speech_df.head())

3. 适配多同类网页的优化

要批量处理同结构的国民议会网页,可添加以下优化:

  • 灵活选择器:使用CSS选择器(soup.select())替代标签查找,适配微小的结构差异,例如将find_all('div', class_='speech')改为select('div.speech, p.speech-block');
  • 异常处理:添加HTTP错误、元素缺失的捕获逻辑,避免单个网页失败中断批量任务:
def scrape_assemblee_speeches(url):
    try:
        response = requests.get(url, timeout=10)
        response.raise_for_status()  # 抛出4xx/5xx HTTP错误
        response.encoding = 'utf-8'
        soup = BeautifulSoup(response.text, 'html.parser')

        speech_blocks = soup.select('div.speech')
        speech_data = []
        
        for block in speech_blocks:
            name_element = block.select_one('b, span.nom')  # 兼容多种姓名标签
            name = name_element.get_text(strip=True) if name_element else "匿名发言"
            full_text = block.get_text(strip=True)
            speech_text = full_text.replace(name, '', 1).strip()
            
            speech_data.append({'name': name, 'text': speech_text})
        
        return pd.DataFrame(speech_data)
    except Exception as e:
        print(f"处理网页 {url} 失败: {str(e)}")
        return pd.DataFrame()

关键错误排查

  • AttributeError原因:直接对find_all()返回的ResultSet调用get_text(),正确做法是遍历集合中的每个元素再调用方法;
  • XML解析错误原因:HTML允许非严格语法(如未闭合标签),而XML要求严格格式,因此必须使用HTML解析器。

内容的提问来源于stack exchange,提问作者MG Fern

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 07:15:36