如何用Python读取文本文件并排除含DM/RT的行以统计推文数
嘿,我来帮你搞定这个统计推文的问题!
首先,咱们先优化下原代码的文件读取方式——用with语句打开文件会更安全,它能自动帮你处理文件关闭的操作,避免资源泄漏。核心要做的就是添加过滤逻辑,把包含“DM”或“RT”的行排除掉。
方案一:直接统计有效行数(最省内存)
如果你的需求只是统计数量,不需要保存每条推文内容,这个方法最高效:
with open('stream.txt', 'r') as file: tweet_count = 0 for line in file: # 检查当前行是否不包含"DM"和"RT" if 'DM' not in line and 'RT' not in line: tweet_count += 1 print(f"符合要求的推文数量:{tweet_count}")
方案二:保存有效推文并统计(方便后续处理)
如果你之后还要对推文内容做拆分或其他操作,可以在过滤后保存每行内容(或拆分后的单词列表):
with open('stream.txt', 'r') as file: # 过滤掉包含"DM"或"RT"的行,同时去掉每行末尾的换行符 valid_tweets = [line.strip() for line in file if 'DM' not in line and 'RT' not in line] # 要是你需要把每行拆分成单词列表,改成下面这行: # valid_tweets = [line.strip().split() for line in file if 'DM' not in line and 'RT' not in line] tweet_count = len(valid_tweets) print(f"符合要求的推文数量:{tweet_count}")
额外说明:如果要排除"DM"/"RT"作为单独单词的行
如果你的需求是只有当"DM"或"RT"是单独单词时才排除该行(比如不排除包含"DMxxx"这类字符串的行),那可以先把每行拆分成单词再检查:
with open('stream.txt', 'r') as file: valid_tweets = [] for line in file: words = line.strip().split() if 'DM' not in words and 'RT' not in words: valid_tweets.append(words) tweet_count = len(valid_tweets) print(f"符合要求的推文数量:{tweet_count}")
根据你的实际需求选对应的方案就好啦!
内容的提问来源于stack exchange,提问作者kmurp62rulz
相关产品推荐
相关产品推荐

