You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup提取指定段落并生成目标MCQ文档求助

问题描述
  • 有一份杂乱的旧MCQ选择题Word文档,已转换为HTML格式,需提取其中的MCQ并整理成规范格式,方便制作Microsoft Forms。
  • 目标格式要求:题目单独成段且加粗,选项以- 开头列在题目下方,清晰区分题目与选项。
  • 现有代码会提取表格内的无效段落,干扰题目和选项的正确整理,无法得到符合需求的题目列表及对应选项。

现有代码

from bs4 import BeautifulSoup
import os
from nltk.tokenize import RegexpTokenizer

# Read .docx file in the CWD
file=[x for x in os.listdir() if '.htm' in x][0]

# Create a soup to parse information
soup = BeautifulSoup(open(file), "html.parser")

# Find all paragraph elements that contains required information
results = soup.find_all("p", class_="MsoNormal")

# Check number of words
tokenizer = RegexpTokenizer(r'\w+')

# Extract questions
Extract_questions=[x.text for x in results if len(tokenizer.tokenize(x.text))>1]

解决方案

步骤1:过滤表格内的无效段落

HTML中表格内的段落通常嵌套在<table>标签下,先排除这些内容:

# 获取所有不在表格内的p.MsoNormal元素
non_table_paragraphs = soup.find_all("p", class_="MsoNormal")
# 过滤掉表格中的段落
filtered_paragraphs = [p for p in non_table_paragraphs if not p.find_parent("table")]

步骤2:区分题目与选项

通过正则匹配识别题目(数字编号开头)和选项(A-D字母编号开头):

import re

# 匹配题目格式(如1.、2.开头)
question_pattern = re.compile(r'^\d+\.')
# 匹配选项格式(如A.、B.开头)
option_pattern = re.compile(r'^[A-D]\.')

mcqs = []
current_question = None

for para in filtered_paragraphs:
    text = para.text.strip()
    if not text:
        continue
    # 识别题目并初始化当前题目对象
    if question_pattern.match(text):
        if current_question:
            mcqs.append(current_question)
        current_question = {
            "question": text,
            "options": []
        }
    # 识别选项并添加到当前题目中
    elif option_pattern.match(text) and current_question:
        current_question["options"].append(text)
# 加入最后一个题目
if current_question:
    mcqs.append(current_question)

步骤3:生成规范DOCX文件

使用python-docx库生成符合要求的文档:

from docx import Document
from docx.shared import Pt

doc = Document()
# 设置基础字体样式
style = doc.styles['Normal']
font = style.font
font.name = 'Arial'
font.size = Pt(12)

for mcq in mcqs:
    # 添加题目并设置加粗
    q_paragraph = doc.add_paragraph()
    q_run = q_paragraph.add_run(mcq["question"])
    q_run.bold = True
    
    # 添加选项,替换字母编号为短横线
    for option in mcq["options"]:
        opt_text = re.sub(r'^[A-D]\.', '-', option).strip()
        doc.add_paragraph(opt_text, style='List Bullet')
    
    # 题目间添加空行分隔
    doc.add_paragraph()

# 保存文档
doc.save("规范MCQ文档.docx")

完整代码

整合所有步骤的完整代码:

from bs4 import BeautifulSoup
import os
import re
from docx import Document
from docx.shared import Pt

# 读取HTML文件
file = [x for x in os.listdir() if '.htm' in x][0]
soup = BeautifulSoup(open(file), "html.parser")

# 过滤表格内的段落
non_table_paragraphs = soup.find_all("p", class_="MsoNormal")
filtered_paragraphs = [p for p in non_table_paragraphs if not p.find_parent("table")]

# 识别题目和选项
question_pattern = re.compile(r'^\d+\.')
option_pattern = re.compile(r'^[A-D]\.')

mcqs = []
current_question = None

for para in filtered_paragraphs:
    text = para.text.strip()
    if not text:
        continue
    if question_pattern.match(text):
        if current_question:
            mcqs.append(current_question)
        current_question = {
            "question": text,
            "options": []
        }
    elif option_pattern.match(text) and current_question:
        current_question["options"].append(text)
if current_question:
    mcqs.append(current_question)

# 生成DOCX文档
doc = Document()
style = doc.styles['Normal']
font = style.font
font.name = 'Arial'
font.size = Pt(12)

for mcq in mcqs:
    q_paragraph = doc.add_paragraph()
    q_run = q_paragraph.add_run(mcq["question"])
    q_run.bold = True
    
    for option in mcq["options"]:
        opt_text = re.sub(r'^[A-D]\.', '-', option).strip()
        doc.add_paragraph(opt_text, style='List Bullet')
    
    doc.add_paragraph()

doc.save("规范MCQ文档.docx")

注意事项

  1. 先安装依赖库:pip install beautifulsoup4 python-docx
  2. 若题目/选项格式有细微差异,需调整正则表达式的匹配规则
  3. 确保HTML文件存放在当前工作目录下

内容的提问来源于stack exchange,提问作者rsc05

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 04:45:37