You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python提取HTML中Executives与Analysts间内容及标题的技术求助

批量处理HTML文件:提取高管信息与页面标题的完整实现

嘿,我看你已经用BeautifulSoup搭好基础的文件遍历框架了,接下来咱们一步步完善代码,搞定你需要的两个提取任务,还能把结果批量保存到CSV里方便后续分析~

首先先修正你现有代码里的小坑:

你代码里page=f.read()之后,文件指针已经到末尾了,再调用f.read()给BeautifulSoup的话,拿到的是空内容。咱们直接把文件内容传给BeautifulSoup就行,不用额外存page变量。


任务1:提取Executives与Analysts之间的高管信息

咱们的核心思路是:

  • 先定位到id="article_participants"的目标div
  • 在这个div里找到<strong>Executives</strong>作为起始节点
  • 遍历它后面的所有<p>标签,直到遇到<strong>Analysts</strong>为止,收集这些p标签的文本

对应的代码片段:

# 定位高管所在的div
participants_div = soup.find('div', class_='content_part hid', id='article_participants')
executives = []
if participants_div:
    # 筛选出Executives和Analysts的strong标签
    strong_tags = participants_div.find_all('strong')
    exec_start = None
    exec_end = None
    for tag in strong_tags:
        tag_text = tag.text.strip()
        if tag_text == 'Executives':
            exec_start = tag
        elif tag_text == 'Analysts':
            exec_end = tag
            break  # 找到Analysts就停止遍历strong标签
    # 遍历起始和结束节点之间的p标签
    if exec_start and exec_end:
        current_node = exec_start.next_sibling
        while current_node != exec_end:
            # 只提取非空的p标签文本
            if current_node.name == 'p' and current_node.text.strip():
                executives.append(current_node.text.strip())
            current_node = current_node.next_sibling

任务2:提取页面标题

页面标题藏在id="page_header"的div下的span[itemprop="headline"]里,直接定位提取即可:

# 提取页面标题
page_header = soup.find('div', id='page_header')
title = ''
if page_header:
    headline_span = page_header.find('span', itemprop='headline')
    if headline_span:
        title = headline_span.text.strip()

完整批量处理代码(含CSV保存)

把上面的逻辑整合到文件遍历循环里,再加上CSV保存功能,所有文件的结果会自动汇总到表格:

from bs4 import BeautifulSoup
import os
import csv

# 指定目标目录
directory = 'C:/Research syntheses - Meta analysis/SeekingAlpha'
# 存储所有文件的处理结果
results = []

for filename in os.listdir(directory):
    if filename.endswith('.html'):
        fname = os.path.join(directory, filename)
        print(f"正在处理文件: {filename}")
        try:
            with open(fname, 'r', encoding='utf-8') as f:
                soup = BeautifulSoup(f.read(), 'html.parser')
                
                # 任务1:提取高管信息
                participants_div = soup.find('div', class_='content_part hid', id='article_participants')
                executives = []
                if participants_div:
                    strong_tags = participants_div.find_all('strong')
                    exec_start = None
                    exec_end = None
                    for tag in strong_tags:
                        tag_text = tag.text.strip()
                        if tag_text == 'Executives':
                            exec_start = tag
                        elif tag_text == 'Analysts':
                            exec_end = tag
                            break
                    if exec_start and exec_end:
                        current_node = exec_start.next_sibling
                        while current_node != exec_end:
                            if current_node.name == 'p' and current_node.text.strip():
                                executives.append(current_node.text.strip())
                            current_node = current_node.next_sibling
                
                # 任务2:提取页面标题
                page_header = soup.find('div', id='page_header')
                title = ''
                if page_header:
                    headline_span = page_header.find('span', itemprop='headline')
                    if headline_span:
                        title = headline_span.text.strip()
                
                # 将当前文件结果加入列表
                results.append({
                    '文件名': filename,
                    '页面标题': title,
                    '高管信息': '; '.join(executives)  # 用分号分隔多个高管信息
                })
        except Exception as e:
            print(f"处理文件 {filename} 时出错: {str(e)}")
            continue  # 单个文件出错不影响整体流程

# 保存结果到CSV
with open('高管信息汇总.csv', 'w', newline='', encoding='utf-8-sig') as csvfile:
    fieldnames = ['文件名', '页面标题', '高管信息']
    writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
    writer.writeheader()
    writer.writerows(results)

print("所有文件处理完成,结果已保存到 高管信息汇总.csv")

关键细节说明

  • 编码处理:打开文件用utf-8编码,CSV保存用utf-8-sig,避免中文乱码问题
  • 异常处理:添加try-except捕获错误,单个文件解析失败不会导致整个程序中断
  • 空值兼容:每个提取步骤都加了存在性判断,确保标签不存在时不会报错,而是返回空值

内容的提问来源于stack exchange,提问作者Jose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 14:29:10