Python爬虫写入CSV文件时移除字段度符号的实现方法
解决方案
只需要在提取字段原始文本后、写入CSV前,对字段值做符号清理即可,无需改动原有爬虫的请求、解析、文件写入逻辑,以下是两种可行实现,按需选择即可:
方法1:字符串直接替换(最简方案)
用Python字符串自带的replace()方法直接移除所有°符号,额外加strip()清理首尾可能存在的空白字符(空格、换行、制表符等),适配当前固定带°符号的场景,性能最优。
修改原有字段提取代码为:# Extract the coordinates Longitude = data[23].get_text().replace('°', '').strip() Latitude = data[24].get_text().replace('°', '').strip() # Extract heading Heading = data[27].get_text().replace('°', '').strip()方法2:正则提取数字(高兼容性方案)
如果后续字段值除了°之外,还可能混入方向标识(N/E/S/W)、多余空格、特殊符号等非数字内容,可以用正则直接提取字段中的数字内容(支持整数、小数、负数格式),你代码中已经提前导入了re模块,无需额外安装依赖。
修改原有字段提取代码为:# Extract the coordinates Longitude = re.search(r'-?\d+\.?\d*', data[23].get_text()).group() Latitude = re.search(r'-?\d+\.?\d*', data[24].get_text()).group() # Extract heading Heading = re.search(r'-?\d+\.?\d*', data[27].get_text()).group()说明:如果你的坐标是度分秒格式,可以对应调整正则匹配规则即可。
修改后完整可运行代码
替换原有字段提取段的代码后,完整代码如下,运行后导出的CSV文件三个字段均为纯数字格式,不会再携带°符号:
#import modules import requests import urllib.request from bs4 import BeautifulSoup from datetime import datetime import time import csv import os import re from selenium import webdriver import schedule try: def retrieve_website(): # Create header headers = {'user-agent': 'Mozilla/5.0 (X11; Linux i686) AppleWebKit/537.17 (KHTML, like Gecko) Chrome/24.0.1312.27 Safari/537.17'} # URL of the ship you want to track, execute the request and parse it to the variable 'soup' url = 'https://website-' reqs = requests.get(url, headers=headers) soup = BeautifulSoup(reqs.text, 'lxml') # Save file to local disk with open("output1.html", "w", encoding='utf-8') as file: file.write(str(soup)) # open file to local disk with open("output1.html", "r", encoding='utf-8') as file: soup = BeautifulSoup(file, 'lxml') # All td tags are read into a list data = soup.find_all('td') # Extract fields and remove degree symbol Longitude = data[23].get_text().replace('°', '').strip() Latitude = data[24].get_text().replace('°', '').strip() Heading = data[27].get_text().replace('°', '').strip() #save as location dwnpath = r'S:\location' # Write data to a csv file with comma as seperator with open(os.path.join(dwnpath, 'Track.csv'), 'w', newline='') as csv_file: fieldnames = ['Longitude', 'Latitude', 'Heading'] writer = csv.DictWriter(csv_file, fieldnames=fieldnames, delimiter=',') writer.writeheader() writer.writerow({'Longitude': Longitude, 'Latitude': Latitude, 'Heading': Heading}) # Start the funtion the first time when the program starts retrieve_website() except Exception as error: print(error) print('Script Complete!')
内容的提问来源于stack exchange,提问作者iqbaltriputra
相关产品推荐
相关产品推荐

