技术问询:实现列表文本追加至特定字符串及辩论脚本预处理
Alright, let's tackle these two Python tasks step by step. I'll walk you through practical, easy-to-follow solutions for each requirement:
需求1:向新列表追加文本直到出现特定字符串
This is simple to implement with a loop that adds elements to your new list and stops the moment it hits your target string. Here's a working example:
# 示例源列表 source_list = ["Start", "Keep adding", "Almost there", "STOP", "Don't include this"] target_string = "STOP" new_list = [] for item in source_list: # 把当前元素追加到新列表 new_list.append(item) # 检查是否是目标字符串,是的话立刻终止循环 if item == target_string: break print(new_list) # 输出: ['Start', 'Keep adding', 'Almost there', 'STOP']
小提示:
- If the target string isn't in the source list, the loop will automatically add all elements to the new list.
- The
breakstatement ensures we don't process any elements that come after the target string—perfect for stopping exactly when you need to.
需求2:拆分特朗普-希拉里辩论脚本到三个人物的发言列表
Assuming your debate script entries follow the format SPEAKER: Speech content (with a colon + space separating the speaker and their words), we can split each entry and sort the content into dedicated lists. Let's assume the three speakers are TRUMP, CLINTON, and a MODERATOR (adjust the speaker names if your script uses different labels):
# 初始化三个空列表,分别存储三位人物的发言 trump_speeches = [] clinton_speeches = [] moderator_speeches = [] # 遍历加载好的脚本列表(loaded_txt) for line in loaded_txt: # 先检查当前条目是否符合 "发言人: 内容" 的格式 if ": " in line: # 只按第一个 ": " 拆分,避免内容里的冒号(比如时间、引用)破坏拆分逻辑 speaker, raw_content = line.split(": ", 1) # 清理内容前后的空白字符(换行、多余空格等) cleaned_content = raw_content.strip() # 根据发言人分类存储 if speaker == "TRUMP": trump_speeches.append(cleaned_content) elif speaker == "CLINTON": clinton_speeches.append(cleaned_content) elif speaker == "MODERATOR": moderator_speeches.append(cleaned_content) else: # 处理未知发言人的情况,避免程序崩溃 print(f"Warning: Unknown speaker '{speaker}' in line: {line}") else: # 处理格式不符合的条目 print(f"Warning: Invalid line format (no ': ' found): {line}") # 验证结果(可选) print(f"Total Trump speeches: {len(trump_speeches)}") print(f"Total Clinton speeches: {len(clinton_speeches)}") print(f"Total Moderator speeches: {len(moderator_speeches)}")
关键细节:
- Using
split(": ", 1)is essential here—it ensures we only split once at the first:, so speech content with colons (like "It's 2:30 AM") won't break the code. - The
strip()method cleans up messy whitespace that often comes with raw text data, making your speech lists cleaner. - We added error handling for unknown speakers and invalid lines to keep the code robust, especially with a large dataset like 1046 entries.
内容的提问来源于stack exchange,提问作者kakanaldo
相关产品推荐
相关产品推荐

