如何将含break的Python for循环转为无break的while循环?
问题描述
这段代码里的break语句让我头疼,研究了半天,想问问有没有符合Python风格的方法把它改成while循环,核心是要移除break。我正在练习设计更高效的循环。
原代码如下:
import re file = open('parse.txt', 'r') html = file.readlines() def cleanup(): result = [] for line in html: if "<li" and "</li>" in line: stripped = re.sub(r'[\n\t]*<[^<]+?>', '', line).rstrip() quoted = f'"{stripped}"' result.append(quoted) elif "INSTRUCTIONS" in line: break return ",\n".join(result)
附parse.txt内容:
<p style="text-align:justify"><strong><span style="background-color:#ecf0f1">INGREDIENTS</span></strong></p> <li style="text-align:justify"><span style="background-color:#ecf0f1">3 lb ground beef (80/20)</span></li> <ul> <li style="text-align:justify"><span style="background-color:#ecf0f1">1 large onion, chopped</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">2-3 cloves garlic, minced</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">2 jalapeño peppers, roasted, peeled, de-seeded, chopped</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">4-5 roma tomatoes, roasted peeled, chopped</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">1 15 oz can kidney beans, strained and washed</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">2 tsp salt</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">2 tsp black pepper</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">2 tsp cumin</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">¼ - ½ tsp cayenne pepper</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">1 tsp garlic powder</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">1 tsp Mexican oregano</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">1 tsp paprika</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">1 tsp smoked paprika</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">3 cups chicken broth</span></li> <li style="text-align:justify"><span style="background-color:#ecf0f1">2 tbsp tomato paste</span></li> </ul> <p style="text-align:justify"><strong>INSTRUCTIONS</strong></p> <ol> <li style="text-align:justify">Heat a large put or Dutch oven over medium-high heat and brown the beef, while stirring to break it up. Cook until no longer pink. Drain out the liquid.</li> <li style="text-align:justify">Stir in onions and cook for about 5 minutes until they are pale and soft. Add in minced garlic and jalapeño peppers, stirring for another minute.</li> <li style="text-align:justify">Stir in the chopped tomatoes, all the spices, and tomato paste until well-distributed and tomato paste has broken up, then follow with the broth. Allow the pot to come to a gentle boil over medium heat, uncovered for about 20 minutes.</li> <li style="text-align:justify">Reduce heat to low, cover and simmer for at least 3 hours, until liquid has reduced.</li> <li style="text-align:justify">During the last 20-30 minutes of cook time, add in the kidney beans; uncover and allow liquid to reduce further during this time.</li> <li style="text-align:justify">Serve hot with jalapeño cornbread muffins, shredded cheese, avocado chunks, chopped cilantro, chopped green onion, tortilla chips.</li> </ol>
解决方案
方法一:转成while循环(彻底移除break)
你可以用索引控制循环,通过while遍历到"INSTRUCTIONS"行就停下,完全不用break:
import re file = open('parse.txt', 'r') html = file.readlines() def cleanup(): result = [] index = 0 # 只要索引没超范围,且当前行不含INSTRUCTIONS,就继续循环 while index < len(html) and "INSTRUCTIONS" not in html[index]: line = html[index] # 原代码的条件判断有bug,这里修正成两个标签都要在该行里 if "<li" in line and "</li>" in line: stripped = re.sub(r'[\n\t]*<[^<]+?>', '', line).rstrip() quoted = f'"{stripped}"' result.append(quoted) index += 1 return ",\n".join(result)
另外提一句:原代码里if "<li" and "</li>" in line的逻辑有问题,"<li"本身是非空字符串,会被当成True,实际只会判断"</li>"是否在该行,这里帮你修正成两个条件都满足的正确逻辑。
方法二:更Pythonic的写法(不用while也能优雅截断)
如果不是非要用while,Python里更地道的方式是用迭代器截断,比如itertools.takewhile,直接取到目标行之前的所有内容,代码更简洁:
import re from itertools import takewhile file = open('parse.txt', 'r') html = file.readlines() def cleanup(): # 只保留INSTRUCTIONS行之前的所有行 lines_before_instructions = takewhile(lambda line: "INSTRUCTIONS" not in line, html) result = [] for line in lines_before_instructions: if "<li" in line and "</li>" in line: stripped = re.sub(r'[\n\t]*<[^<]+?>', '', line).rstrip() quoted = f'"{stripped}"' result.append(quoted) return ",\n".join(result)
这种写法符合Python的迭代思维,不用手动管理索引,可读性更高。
额外优化:文件处理的正确姿势
原代码直接open文件没关闭,建议用with上下文管理器,确保文件自动关闭,还能逐行读取省内存:
import re from itertools import takewhile def cleanup(): result = [] with open('parse.txt', 'r') as file: # 逐行读取,碰到INSTRUCTIONS就停止 lines_before_instructions = takewhile(lambda line: "INSTRUCTIONS" not in line, file) for line in lines_before_instructions: if "<li" in line and "</li>" in line: stripped = re.sub(r'[\n\t]*<[^<]+?>', '', line).rstrip() quoted = f'"{stripped}"' result.append(quoted) return ",\n".join(result)
这样处理大文件也不会占太多内存,还避免了资源泄漏。
内容的提问来源于stack exchange,提问作者Snerd
相关产品推荐
相关产品推荐

