如何将Python爬虫的Counter统计结果保存至TXT文件?
问题分析与修改后的代码
原代码存在两个核心问题:
- 统计与保存逻辑缩进错误,每次分页点击后就执行统计保存,会重复覆盖文件且未爬完所有页面
Counter对象无法直接写入文件,需先转换为字符串格式
以下是修复后的完整代码:
from selenium import webdriver from bs4 import BeautifulSoup from selenium.common.exceptions import ElementClickInterceptedException from selenium.webdriver.common.by import By from collections import Counter driver = webdriver.Chrome() driver.get("URL") # 替换为你的目标网页地址 city = [] while True: driver.implicitly_wait(10) page_source = driver.page_source soup = BeautifulSoup(page_source, 'lxml') # 提取当前页面城市信息并去重空白字符,直接扩展列表更简洁 cities = [x.get_text().strip() for x in soup.find_all('span', attrs={'class': 'region d-inline-block mr-5'})] city.extend(cities) try: driver.find_element(by=By.LINK_TEXT, value='»').click() except ElementClickInterceptedException: break # 爬完所有页面后统一统计 count = Counter(city) print(count) # 将统计结果写入文件,提供两种格式可选 with open('example.txt', 'w', encoding='utf-8') as f: # 方式1:直接转为字符串写入(与控制台输出格式一致) f.write(str(count)) # 方式2:逐行写入,更易读(注释掉上面一行,启用下面代码) # for city_name, num in count.items(): # f.write(f"{city_name}: {num}\n") driver.quit()
关键修改说明
- 将统计和保存代码移至
while循环外部,确保爬完所有页面后一次性处理,避免重复覆盖文件 - 把
f.write(count)改为f.write(str(count)),解决非字符串类型无法写入文件的问题 - 用
extend替代循环append优化列表添加逻辑,同时用strip()清理城市名称的空白字符 - 移除了无意义的
result = driver.get(...)语句,driver.get()无返回值
内容的提问来源于stack exchange,提问作者3hmd
相关产品推荐
相关产品推荐

