如何抓取法国国民议会网页数据并生成目标DataFrame?
解决法国国民议会网页抓取并生成DataFrame的问题
问题概述
需抓取法国国民议会特定格式的网页(示例:https://www.assemblee-nationale.fr/13/cri/2006-2007/20070152.asp),生成包含name(发言者姓名)和text(发言内容)列的Pandas DataFrame。此前尝试遇到两个关键错误:
- 使用BeautifulSoup时,误将
ResultSet对象当作单个元素处理,触发AttributeError; - 尝试以XML方式解析网页,出现「not well-formed (invalid token)」错误。
解决方案
1. 核心修正方向
- 放弃XML解析:目标网页为HTML格式,存在XML不兼容的语法(如未闭合标签、特殊字符),改用HTML解析器即可规避格式错误;
- 正确处理ResultSet:
soup.find_all()返回的是元素集合,必须遍历每个元素提取数据,不能直接对集合调用元素级方法(如get_text())。
2. 可运行代码示例
import requests from bs4 import BeautifulSoup import pandas as pd def scrape_assemblee_speeches(url): # 获取并解析网页 response = requests.get(url) response.encoding = 'utf-8' # 强制指定UTF-8编码避免乱码 soup = BeautifulSoup(response.text, 'html.parser') # 使用HTML解析器 # 定位所有发言块:根据目标网页结构,发言内容包裹在class为"speech"的div中(需根据实际调整) speech_blocks = soup.find_all('div', class_='speech') speech_data = [] for block in speech_blocks: # 提取发言者姓名:假设姓名在<b>标签内(实际需匹配网页标签) name_element = block.find('b') if not name_element: continue # 跳过无姓名的发言块 name = name_element.get_text(strip=True) # 提取发言内容:移除姓名后保留剩余文本 full_text = block.get_text(strip=True) speech_text = full_text.replace(name, '', 1).strip() speech_data.append({'name': name, 'text': speech_text}) # 转换为DataFrame return pd.DataFrame(speech_data) # 测试示例网页 sample_url = "https://www.assemblee-nationale.fr/13/cri/2006-2007/20070152.asp" speech_df = scrape_assemblee_speeches(sample_url) print(speech_df.head())
3. 适配多同类网页的优化
要批量处理同结构的国民议会网页,可添加以下优化:
- 灵活选择器:使用CSS选择器(
soup.select())替代标签查找,适配微小的结构差异,例如将find_all('div', class_='speech')改为select('div.speech, p.speech-block'); - 异常处理:添加HTTP错误、元素缺失的捕获逻辑,避免单个网页失败中断批量任务:
def scrape_assemblee_speeches(url): try: response = requests.get(url, timeout=10) response.raise_for_status() # 抛出4xx/5xx HTTP错误 response.encoding = 'utf-8' soup = BeautifulSoup(response.text, 'html.parser') speech_blocks = soup.select('div.speech') speech_data = [] for block in speech_blocks: name_element = block.select_one('b, span.nom') # 兼容多种姓名标签 name = name_element.get_text(strip=True) if name_element else "匿名发言" full_text = block.get_text(strip=True) speech_text = full_text.replace(name, '', 1).strip() speech_data.append({'name': name, 'text': speech_text}) return pd.DataFrame(speech_data) except Exception as e: print(f"处理网页 {url} 失败: {str(e)}") return pd.DataFrame()
关键错误排查
- AttributeError原因:直接对
find_all()返回的ResultSet调用get_text(),正确做法是遍历集合中的每个元素再调用方法; - XML解析错误原因:HTML允许非严格语法(如未闭合标签),而XML要求严格格式,因此必须使用HTML解析器。
内容的提问来源于stack exchange,提问作者MG Fern
相关产品推荐
相关产品推荐

