如何修复将docx试题文件转为Python字典时返回空字典的问题
修复docx转试题字典的空字典问题
你要把带项目符号的docx试题转为指定结构的Python字典,但运行现有代码得到空字典,问题出在几个关键逻辑上,以下是修复方案:
试题格式参考
- Question
a. first answer
b. second answer
c. third answer
d. fourth answer
e. fifth answer
目标字典结构
{ '1': { 'question': 'the question text', 'answer': ['first answer','second answer','third answer','fourth answer','fifth answer'], 'correct_answer': 2 }, # 其他题目... }
原代码存在的问题及修复点
1. 题目编号判断逻辑错误
原代码通过text[1] == '.'判断,但如果题目是多位数(比如10. xxx),这个判断会失效。应该用.分割编号和题目内容,确保兼容单/多位数编号。
2. 选项遍历方式错误
python-docx的Paragraph对象没有next_paragraph属性,原代码的遍历逻辑根本无法获取后续选项段落,需要先把所有段落存入列表,通过索引来遍历后续段落。
3. 加粗判断逻辑不严谨
- 原代码只检查选项段落的第一个run,但若加粗是整个选项文本或部分run,会漏判;
run.bold可能返回None(默认样式继承自段落),需要明确判断run.bold is True。
4. 选项文本截取逻辑错误
原代码用next_text[3:]截取,假设前缀是a. (字母+点+空格),但如果没有空格(比如a.xxx),会错误截断文本,应该用. 分割前缀和内容。
修复后的完整代码
from docx import Document def has_bold_run(paragraph): # 检查段落中是否有任意run是加粗状态 for run in paragraph.runs: if run.bold is True: return True return False # 打开文档 doc = Document('sample.docx') paragraphs = list(doc.paragraphs) # 把段落转为列表方便按索引访问 questions_and_answers = {} # 遍历所有段落 for idx, paragraph in enumerate(paragraphs): text = paragraph.text.strip() if not text: continue # 判断是否是题目(以数字+点开头) if text[0].isdigit() and '.' in text: # 分割编号和题目内容 question_part = text.split('.', 1) if len(question_part) < 2: continue question_number = question_part[0].strip() question_text = question_part[1].strip() answer_choices = [] correct_answer_index = None # 遍历后续段落获取选项 for next_idx in range(idx + 1, len(paragraphs)): next_paragraph = paragraphs[next_idx] next_text = next_paragraph.text.strip() if not next_text: continue # 判断是否是下一个题目(终止选项遍历) if next_text[0].isdigit() and '.' in next_text: break # 判断是否是选项(字母+点开头) if next_text[0].isalpha() and '.' in next_text: option_part = next_text.split('.', 1) if len(option_part) < 2: continue option_text = option_part[1].strip() answer_choices.append(option_text) # 检查该选项是否是正确答案(加粗) if has_bold_run(next_paragraph): correct_answer_index = len(answer_choices) - 1 # 将题目加入字典,对齐目标结构的键名 questions_and_answers[question_number] = { 'question': question_text, 'answer': answer_choices, 'correct_answer': correct_answer_index } # 打印结果 for number, data in questions_and_answers.items(): print(f"{number}: {data['question']}") print("Answers:") for idx, ans in enumerate(data['answer']): print(f"- {ans} {'(correct)' if idx == data['correct_answer'] else ''}") print()
内容的提问来源于stack exchange,提问作者Simone Fildi
相关产品推荐
相关产品推荐

