Python如何删除带编号段落中不符合连续顺序的段落内容?
实现思路
- 第一步:匹配所有带编号的段落,兼容开头带引号的格式,同时提取编号和对应内容
- 第二步:用字典存储编号和内容,天然解决重复编号的问题(相同编号仅保留一条)
- 第三步:从编号1开始依次查找,只要存在对应编号的内容就保留,直到连续中断或者达到最大预期编号221为止
代码实现
import re def filter_continuous_numbered_paragraphs(text, max_expected=221): # 正则兼容行首带引号的编号格式,分组提取编号和内容 pattern = re.compile(r'^\"?(\d+)\.\s+(.*)$', flags=re.MULTILINE) num_content_map = {} for match in pattern.finditer(text): num = int(match.group(1)) content = match.group(2) # 重复编号默认保留第一个匹配结果,要保留最后一个可去掉if判断直接赋值 if num not in num_content_map: num_content_map[num] = content # 收集从1开始连续的段落 result = [] current_num = 1 while current_num <= max_expected and current_num in num_content_map: result.append(f"{current_num}. {num_content_map[current_num]}") current_num += 1 return '\n'.join(result) # 测试示例 text = """1. Shares of Paras Defence and Space Technologies gained 2.85 times. 2. The company, engaged in manufacturing and testing of defence and space engineering products. "3. Its stock ended at Rs 499 versus issue price of Rs 175 per share. 42. On July 23, Zomato NSE 0.00 % Ltd. listed on the Indian stock exchanges. 43. That was exactly a week after the food-delivery and restaurant discovery platform's initial public offering went live. 4. Paras Defence’s IPO, which closed on September 23, had generated bids worth Rs 38,021 crore. 5. It surpassed the previous record of Salasar Technologies’ IPO. 14. NBFCs are betting big time on the IPO. 6. Paras Defence is one of the few players having an edge in defence deals.""" relevant_text = filter_continuous_numbered_paragraphs(text) print(relevant_text)
方案说明
- 正则里的
\"?配置了行首可选的引号,解决了之前匹配不到"3.格式编号的问题 - 用字典存储键值对的方式自动处理重复编号,可根据需求调整保留第一个还是最后一个重复内容
- 过滤逻辑自动跳过42、43、14这类非连续编号,最终输出的段落按编号顺序排列,和原始文本的段落位置无关
- 可通过调整
max_expected参数控制最大匹配的编号上限,适配你1到221的编号范围需求
内容的提问来源于stack exchange,提问作者Ashraya Singh
相关产品推荐
相关产品推荐

