You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫写入CSV文件时移除字段度符号的实现方法

解决方案

只需要在提取字段原始文本后、写入CSV前,对字段值做符号清理即可,无需改动原有爬虫的请求、解析、文件写入逻辑,以下是两种可行实现,按需选择即可:

  • 方法1:字符串直接替换(最简方案)
    用Python字符串自带的replace()方法直接移除所有°符号,额外加strip()清理首尾可能存在的空白字符(空格、换行、制表符等),适配当前固定带°符号的场景,性能最优。
    修改原有字段提取代码为:

    # Extract the coordinates
    Longitude = data[23].get_text().replace('°', '').strip()
    Latitude = data[24].get_text().replace('°', '').strip()
    
    # Extract heading
    Heading = data[27].get_text().replace('°', '').strip()
    
  • 方法2:正则提取数字(高兼容性方案)
    如果后续字段值除了°之外,还可能混入方向标识(N/E/S/W)、多余空格、特殊符号等非数字内容,可以用正则直接提取字段中的数字内容(支持整数、小数、负数格式),你代码中已经提前导入了re模块,无需额外安装依赖。
    修改原有字段提取代码为:

    # Extract the coordinates
    Longitude = re.search(r'-?\d+\.?\d*', data[23].get_text()).group()
    Latitude = re.search(r'-?\d+\.?\d*', data[24].get_text()).group()
    
    # Extract heading
    Heading = re.search(r'-?\d+\.?\d*', data[27].get_text()).group()
    

    说明:如果你的坐标是度分秒格式,可以对应调整正则匹配规则即可。

修改后完整可运行代码

替换原有字段提取段的代码后,完整代码如下,运行后导出的CSV文件三个字段均为纯数字格式,不会再携带°符号:

#import modules
import requests
import urllib.request
from bs4 import BeautifulSoup
from datetime import datetime
import time
import csv
import os
import re
from selenium import webdriver
import schedule

try:      
    def retrieve_website():
        # Create header
        headers = {'user-agent': 'Mozilla/5.0 (X11; Linux i686) AppleWebKit/537.17 (KHTML, like Gecko) Chrome/24.0.1312.27 Safari/537.17'}

        # URL of the ship you want to track, execute the request and parse it to the variable 'soup'
        url = 'https://website-'
        reqs = requests.get(url, headers=headers)
        soup = BeautifulSoup(reqs.text, 'lxml')

        # Save file to local disk
        with open("output1.html", "w", encoding='utf-8') as file:
            file.write(str(soup))

        # open file to local disk
        with open("output1.html", "r", encoding='utf-8') as file:
            soup = BeautifulSoup(file, 'lxml')

        # All td tags are read into a list
        data = soup.find_all('td')

        # Extract fields and remove degree symbol
        Longitude = data[23].get_text().replace('°', '').strip()
        Latitude = data[24].get_text().replace('°', '').strip()
        Heading = data[27].get_text().replace('°', '').strip()
        
        #save as location
        dwnpath = r'S:\location'
        
        # Write data to a csv file with comma as seperator    
        with open(os.path.join(dwnpath, 'Track.csv'), 'w', newline='') as csv_file:
            fieldnames = ['Longitude', 'Latitude', 'Heading']
            writer = csv.DictWriter(csv_file, fieldnames=fieldnames, delimiter=',')
            writer.writeheader()
            writer.writerow({'Longitude': Longitude, 
                             'Latitude': Latitude,
                              'Heading': Heading})
    
    # Start the funtion the first time when the program starts
    retrieve_website()
    
except Exception as error:
    print(error)
    
print('Script Complete!')

内容的提问来源于stack exchange,提问作者iqbaltriputra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 09:03:04