You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复将docx试题文件转为Python字典时返回空字典的问题

修复docx转试题字典的空字典问题

你要把带项目符号的docx试题转为指定结构的Python字典,但运行现有代码得到空字典,问题出在几个关键逻辑上,以下是修复方案:

试题格式参考

  1. Question
    a. first answer
    b. second answer
    c. third answer
    d. fourth answer
    e. fifth answer

目标字典结构

{
  '1': {  
    'question': 'the question text',
    'answer': ['first answer','second answer','third answer','fourth answer','fifth answer'],
    'correct_answer': 2
   },
   # 其他题目...
}

原代码存在的问题及修复点

1. 题目编号判断逻辑错误

原代码通过text[1] == '.'判断,但如果题目是多位数(比如10. xxx),这个判断会失效。应该用.分割编号和题目内容,确保兼容单/多位数编号。

2. 选项遍历方式错误

python-docx的Paragraph对象没有next_paragraph属性,原代码的遍历逻辑根本无法获取后续选项段落,需要先把所有段落存入列表,通过索引来遍历后续段落。

3. 加粗判断逻辑不严谨

  • 原代码只检查选项段落的第一个run,但若加粗是整个选项文本或部分run,会漏判;
  • run.bold可能返回None(默认样式继承自段落),需要明确判断run.bold is True。

4. 选项文本截取逻辑错误

原代码用next_text[3:]截取,假设前缀是a. (字母+点+空格),但如果没有空格(比如a.xxx),会错误截断文本,应该用. 分割前缀和内容。

修复后的完整代码

from docx import Document

def has_bold_run(paragraph):
    # 检查段落中是否有任意run是加粗状态
    for run in paragraph.runs:
        if run.bold is True:
            return True
    return False

# 打开文档
doc = Document('sample.docx')
paragraphs = list(doc.paragraphs)  # 把段落转为列表方便按索引访问
questions_and_answers = {}

# 遍历所有段落
for idx, paragraph in enumerate(paragraphs):
    text = paragraph.text.strip()
    if not text:
        continue
    
    # 判断是否是题目(以数字+点开头)
    if text[0].isdigit() and '.' in text:
        # 分割编号和题目内容
        question_part = text.split('.', 1)
        if len(question_part) < 2:
            continue
        question_number = question_part[0].strip()
        question_text = question_part[1].strip()
        
        answer_choices = []
        correct_answer_index = None
        
        # 遍历后续段落获取选项
        for next_idx in range(idx + 1, len(paragraphs)):
            next_paragraph = paragraphs[next_idx]
            next_text = next_paragraph.text.strip()
            if not next_text:
                continue
            
            # 判断是否是下一个题目(终止选项遍历)
            if next_text[0].isdigit() and '.' in next_text:
                break
            
            # 判断是否是选项(字母+点开头)
            if next_text[0].isalpha() and '.' in next_text:
                option_part = next_text.split('.', 1)
                if len(option_part) < 2:
                    continue
                option_text = option_part[1].strip()
                answer_choices.append(option_text)
                
                # 检查该选项是否是正确答案(加粗)
                if has_bold_run(next_paragraph):
                    correct_answer_index = len(answer_choices) - 1
        
        # 将题目加入字典,对齐目标结构的键名
        questions_and_answers[question_number] = {
            'question': question_text,
            'answer': answer_choices,
            'correct_answer': correct_answer_index
        }

# 打印结果
for number, data in questions_and_answers.items():
    print(f"{number}: {data['question']}")
    print("Answers:")
    for idx, ans in enumerate(data['answer']):
        print(f"- {ans} {'(correct)' if idx == data['correct_answer'] else ''}")
    print()

内容的提问来源于stack exchange,提问作者Simone Fildi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 01:40:15