Python实现OCR书页文本按关键词分组为章节及索引遍历问题排查
解决列表按章节关键词分组的索引问题
我来帮你搞定这个分组问题!你遇到的索引跳过和逻辑错误,主要是因为原代码里的两个坑:list.index()的局限性,以及迭代器使用不当。咱们一步步来解决:
先分析原代码的问题
test_list.index(f)的隐患:这个方法只会返回第一个匹配元素的索引,如果有多个页面包含相同的关键词文本,会导致起始索引记录错误。比如如果你的列表里有两页都是'hi',index()只会返回第一个'hi'的位置,后面的就都错了。- 迭代器使用逻辑错误:你初始化的
indices列表会包含重复的起始索引(比如例子里的[0,0,3]),用iter遍历的时候,每次调用next(it)会直接移动指针,导致切片范围完全不符合预期,甚至最后会抛出StopIteration异常。
正确的解决方案:先收集准确的起始索引,再分组
我们先准确记录所有包含关键词的页面索引,然后根据这些索引把列表切成连续的章节组,完全避免索引跳过的问题:
代码示例
# 你的测试列表 test_list = ['is the', 'best', 'and', 'is so', 'popular'] # 章节起始页的关键词 keyword = 'is' # 第一步:收集所有章节起始页的索引(用enumerate准确获取每个匹配项的位置) start_indices = [] for idx, page_text in enumerate(test_list): if keyword in page_text: start_indices.append(idx) # 第二步:根据起始索引拆分章节 chapters = [] for i in range(len(start_indices)): current_start = start_indices[i] # 下一个章节的起始位置就是当前章节的结束位置,最后一组直接到列表末尾 current_end = start_indices[i+1] if (i+1) < len(start_indices) else len(test_list) chapters.append(test_list[current_start:current_end]) print(chapters) # 输出:[['is the', 'best', 'and'], ['is so', 'popular']]
验证你的另一个测试案例
针对关键词'hi'的情况,同样适用:
test_list = ['hi','yes', 'hi you', 'me'] keyword = 'hi' start_indices = [idx for idx, page in enumerate(test_list) if keyword in page] chapters = [] for i in range(len(start_indices)): start = start_indices[i] end = start_indices[i+1] if i+1 < len(start_indices) else len(test_list) chapters.append(test_list[start:end]) print(chapters) # 输出:[['hi', 'yes'], ['hi you', 'me']]
为什么这个方法靠谱?
- 准确记录起始位置:用
enumerate遍历每个页面的索引和内容,确保每个包含关键词的页面位置都被正确记录,不会遗漏或重复。 - 清晰的分组逻辑:通过遍历起始索引列表,每个章节的范围明确是从当前起始页到下一个起始页(最后一组到列表末尾),完全符合你要保留起始页并分组的需求。
内容的提问来源于stack exchange,提问作者steca
相关产品推荐
相关产品推荐

