You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在BeautifulSoup中使用多replace函数处理CSV输出问题求助

解决BeautifulSoup爬取新冠数据时的链式replace问题

你的需求完全可以通过链式调用replace()实现,Python字符串的replace()方法本身会返回新字符串,天然支持链式操作。以下是针对你两个问题的具体修改方案:

1. 数值字段处理(替换N/A为"0"、去除千分位逗号)

对于确诊数、康复数这类数值字段,直接按你设想的链式写法即可,建议把strip()放在最后清理首尾的空白字符(比如换行、制表符):

# 以总康复数为例
total_recovered = cols[6].text.replace(",", "").replace("N/A", "0").strip()

如果字段中存在数值中间的多余空格,可以在链式调用中加入replace(" ", ""),但如果只是首尾空格,strip()已经足够处理。

2. 国家名称处理(去除所有空格)

针对包含空格的国家名称(比如North Macedonia),单独处理时只需去掉所有空格即可:

# 假设国家名称在第1个td列(索引0)
country = cols[0].text.replace(" ", "").strip()

完整代码示例

from bs4 import BeautifulSoup
import requests
import csv

# 示例:获取目标网页内容
url = "你的新冠数据目标网页URL"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# 写入CSV文件
with open('covid_data.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    # 写入表头
    writer.writerow(["Country", "Total Cases", "Total Recovered"])
    
    # 遍历表格行(跳过表头行)
    for row in soup.find_all('tr')[1:]:
        cols = row.find_all('td')
        if len(cols) < 7:  # 跳过无效行
            continue
        
        # 处理国家名称:去除所有空格
        country = cols[0].text.replace(" ", "").strip()
        # 处理总确诊数:去逗号、替换N/A为0
        total_cases = cols[1].text.replace(",", "").replace("N/A", "0").strip()
        # 处理总康复数:链式replace操作
        total_recovered = cols[6].text.replace(",", "").replace("N/A", "0").strip()
        
        writer.writerow([country, total_cases, total_recovered])

注意事项

  • 如果遇到非标准空白字符(比如&nbsp;对应的\xa0),可以修改strip()为strip('\xa0 ')来同时清理这类特殊空白
  • 链式调用replace()的顺序不影响最终结果,但建议先处理格式符号(比如逗号),再替换特殊值(比如N/A),最后清理空白

内容的提问来源于stack exchange,提问作者Ömer Sevban Tümer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 00:20:40