Python提取HTML中Executives与Analysts间内容及标题的技术求助
批量处理HTML文件:提取高管信息与页面标题的完整实现
嘿,我看你已经用BeautifulSoup搭好基础的文件遍历框架了,接下来咱们一步步完善代码,搞定你需要的两个提取任务,还能把结果批量保存到CSV里方便后续分析~
首先先修正你现有代码里的小坑:
你代码里
page=f.read()之后,文件指针已经到末尾了,再调用f.read()给BeautifulSoup的话,拿到的是空内容。咱们直接把文件内容传给BeautifulSoup就行,不用额外存page变量。
任务1:提取Executives与Analysts之间的高管信息
咱们的核心思路是:
- 先定位到
id="article_participants"的目标div - 在这个div里找到
<strong>Executives</strong>作为起始节点 - 遍历它后面的所有
<p>标签,直到遇到<strong>Analysts</strong>为止,收集这些p标签的文本
对应的代码片段:
# 定位高管所在的div participants_div = soup.find('div', class_='content_part hid', id='article_participants') executives = [] if participants_div: # 筛选出Executives和Analysts的strong标签 strong_tags = participants_div.find_all('strong') exec_start = None exec_end = None for tag in strong_tags: tag_text = tag.text.strip() if tag_text == 'Executives': exec_start = tag elif tag_text == 'Analysts': exec_end = tag break # 找到Analysts就停止遍历strong标签 # 遍历起始和结束节点之间的p标签 if exec_start and exec_end: current_node = exec_start.next_sibling while current_node != exec_end: # 只提取非空的p标签文本 if current_node.name == 'p' and current_node.text.strip(): executives.append(current_node.text.strip()) current_node = current_node.next_sibling
任务2:提取页面标题
页面标题藏在id="page_header"的div下的span[itemprop="headline"]里,直接定位提取即可:
# 提取页面标题 page_header = soup.find('div', id='page_header') title = '' if page_header: headline_span = page_header.find('span', itemprop='headline') if headline_span: title = headline_span.text.strip()
完整批量处理代码(含CSV保存)
把上面的逻辑整合到文件遍历循环里,再加上CSV保存功能,所有文件的结果会自动汇总到表格:
from bs4 import BeautifulSoup import os import csv # 指定目标目录 directory = 'C:/Research syntheses - Meta analysis/SeekingAlpha' # 存储所有文件的处理结果 results = [] for filename in os.listdir(directory): if filename.endswith('.html'): fname = os.path.join(directory, filename) print(f"正在处理文件: {filename}") try: with open(fname, 'r', encoding='utf-8') as f: soup = BeautifulSoup(f.read(), 'html.parser') # 任务1:提取高管信息 participants_div = soup.find('div', class_='content_part hid', id='article_participants') executives = [] if participants_div: strong_tags = participants_div.find_all('strong') exec_start = None exec_end = None for tag in strong_tags: tag_text = tag.text.strip() if tag_text == 'Executives': exec_start = tag elif tag_text == 'Analysts': exec_end = tag break if exec_start and exec_end: current_node = exec_start.next_sibling while current_node != exec_end: if current_node.name == 'p' and current_node.text.strip(): executives.append(current_node.text.strip()) current_node = current_node.next_sibling # 任务2:提取页面标题 page_header = soup.find('div', id='page_header') title = '' if page_header: headline_span = page_header.find('span', itemprop='headline') if headline_span: title = headline_span.text.strip() # 将当前文件结果加入列表 results.append({ '文件名': filename, '页面标题': title, '高管信息': '; '.join(executives) # 用分号分隔多个高管信息 }) except Exception as e: print(f"处理文件 {filename} 时出错: {str(e)}") continue # 单个文件出错不影响整体流程 # 保存结果到CSV with open('高管信息汇总.csv', 'w', newline='', encoding='utf-8-sig') as csvfile: fieldnames = ['文件名', '页面标题', '高管信息'] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() writer.writerows(results) print("所有文件处理完成,结果已保存到 高管信息汇总.csv")
关键细节说明
- 编码处理:打开文件用
utf-8编码,CSV保存用utf-8-sig,避免中文乱码问题 - 异常处理:添加
try-except捕获错误,单个文件解析失败不会导致整个程序中断 - 空值兼容:每个提取步骤都加了存在性判断,确保标签不存在时不会报错,而是返回空值
内容的提问来源于stack exchange,提问作者Jose
相关产品推荐
相关产品推荐

