基于BeautifulSoup提取指定段落并生成目标MCQ文档求助
问题描述
- 有一份杂乱的旧MCQ选择题Word文档,已转换为HTML格式,需提取其中的MCQ并整理成规范格式,方便制作Microsoft Forms。
- 目标格式要求:题目单独成段且加粗,选项以
-开头列在题目下方,清晰区分题目与选项。 - 现有代码会提取表格内的无效段落,干扰题目和选项的正确整理,无法得到符合需求的题目列表及对应选项。
现有代码
from bs4 import BeautifulSoup import os from nltk.tokenize import RegexpTokenizer # Read .docx file in the CWD file=[x for x in os.listdir() if '.htm' in x][0] # Create a soup to parse information soup = BeautifulSoup(open(file), "html.parser") # Find all paragraph elements that contains required information results = soup.find_all("p", class_="MsoNormal") # Check number of words tokenizer = RegexpTokenizer(r'\w+') # Extract questions Extract_questions=[x.text for x in results if len(tokenizer.tokenize(x.text))>1]
解决方案
步骤1:过滤表格内的无效段落
HTML中表格内的段落通常嵌套在<table>标签下,先排除这些内容:
# 获取所有不在表格内的p.MsoNormal元素 non_table_paragraphs = soup.find_all("p", class_="MsoNormal") # 过滤掉表格中的段落 filtered_paragraphs = [p for p in non_table_paragraphs if not p.find_parent("table")]
步骤2:区分题目与选项
通过正则匹配识别题目(数字编号开头)和选项(A-D字母编号开头):
import re # 匹配题目格式(如1.、2.开头) question_pattern = re.compile(r'^\d+\.') # 匹配选项格式(如A.、B.开头) option_pattern = re.compile(r'^[A-D]\.') mcqs = [] current_question = None for para in filtered_paragraphs: text = para.text.strip() if not text: continue # 识别题目并初始化当前题目对象 if question_pattern.match(text): if current_question: mcqs.append(current_question) current_question = { "question": text, "options": [] } # 识别选项并添加到当前题目中 elif option_pattern.match(text) and current_question: current_question["options"].append(text) # 加入最后一个题目 if current_question: mcqs.append(current_question)
步骤3:生成规范DOCX文件
使用python-docx库生成符合要求的文档:
from docx import Document from docx.shared import Pt doc = Document() # 设置基础字体样式 style = doc.styles['Normal'] font = style.font font.name = 'Arial' font.size = Pt(12) for mcq in mcqs: # 添加题目并设置加粗 q_paragraph = doc.add_paragraph() q_run = q_paragraph.add_run(mcq["question"]) q_run.bold = True # 添加选项,替换字母编号为短横线 for option in mcq["options"]: opt_text = re.sub(r'^[A-D]\.', '-', option).strip() doc.add_paragraph(opt_text, style='List Bullet') # 题目间添加空行分隔 doc.add_paragraph() # 保存文档 doc.save("规范MCQ文档.docx")
完整代码
整合所有步骤的完整代码:
from bs4 import BeautifulSoup import os import re from docx import Document from docx.shared import Pt # 读取HTML文件 file = [x for x in os.listdir() if '.htm' in x][0] soup = BeautifulSoup(open(file), "html.parser") # 过滤表格内的段落 non_table_paragraphs = soup.find_all("p", class_="MsoNormal") filtered_paragraphs = [p for p in non_table_paragraphs if not p.find_parent("table")] # 识别题目和选项 question_pattern = re.compile(r'^\d+\.') option_pattern = re.compile(r'^[A-D]\.') mcqs = [] current_question = None for para in filtered_paragraphs: text = para.text.strip() if not text: continue if question_pattern.match(text): if current_question: mcqs.append(current_question) current_question = { "question": text, "options": [] } elif option_pattern.match(text) and current_question: current_question["options"].append(text) if current_question: mcqs.append(current_question) # 生成DOCX文档 doc = Document() style = doc.styles['Normal'] font = style.font font.name = 'Arial' font.size = Pt(12) for mcq in mcqs: q_paragraph = doc.add_paragraph() q_run = q_paragraph.add_run(mcq["question"]) q_run.bold = True for option in mcq["options"]: opt_text = re.sub(r'^[A-D]\.', '-', option).strip() doc.add_paragraph(opt_text, style='List Bullet') doc.add_paragraph() doc.save("规范MCQ文档.docx")
注意事项
- 先安装依赖库:
pip install beautifulsoup4 python-docx - 若题目/选项格式有细微差异,需调整正则表达式的匹配规则
- 确保HTML文件存放在当前工作目录下
内容的提问来源于stack exchange,提问作者rsc05
相关产品推荐
相关产品推荐

