You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python爬虫的Counter统计结果保存至TXT文件?

问题分析与修改后的代码

原代码存在两个核心问题:

  1. 统计与保存逻辑缩进错误,每次分页点击后就执行统计保存,会重复覆盖文件且未爬完所有页面
  2. Counter对象无法直接写入文件,需先转换为字符串格式

以下是修复后的完整代码:

from selenium import webdriver
from bs4 import BeautifulSoup
from selenium.common.exceptions import ElementClickInterceptedException
from selenium.webdriver.common.by import By
from collections import Counter

driver = webdriver.Chrome()
driver.get("URL")  # 替换为你的目标网页地址

city = []

while True:
    driver.implicitly_wait(10)
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, 'lxml')

    # 提取当前页面城市信息并去重空白字符,直接扩展列表更简洁
    cities = [x.get_text().strip() for x in soup.find_all('span', attrs={'class': 'region d-inline-block mr-5'})]
    city.extend(cities)

    try:
        driver.find_element(by=By.LINK_TEXT, value='»').click()
    except ElementClickInterceptedException:
        break

# 爬完所有页面后统一统计
count = Counter(city)
print(count)

# 将统计结果写入文件,提供两种格式可选
with open('example.txt', 'w', encoding='utf-8') as f:
    # 方式1:直接转为字符串写入(与控制台输出格式一致)
    f.write(str(count))
    
    # 方式2:逐行写入,更易读(注释掉上面一行,启用下面代码)
    # for city_name, num in count.items():
    #     f.write(f"{city_name}: {num}\n")

driver.quit()

关键修改说明

  • 将统计和保存代码移至while循环外部,确保爬完所有页面后一次性处理,避免重复覆盖文件
  • 把f.write(count)改为f.write(str(count)),解决非字符串类型无法写入文件的问题
  • 用extend替代循环append优化列表添加逻辑,同时用strip()清理城市名称的空白字符
  • 移除了无意义的result = driver.get(...)语句,driver.get()无返回值

内容的提问来源于stack exchange,提问作者3hmd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 01:50:20