如何将类矩阵行列预制列表转换为指定格式CSV文件
问题与解决方案
需求
抓取美国劳工统计局网页数据,生成符合以下要求的CSV文件:
- 首行为标题行
- 后续为1985-2022年共38行数据,每行对应12个月度数据,形成38×12的矩阵格式
修正后的代码
from bs4 import BeautifulSoup import requests import pandas as pd url = "https://www.bls.gov/web/ximpim/beaimp.htm" page = requests.get(url) doc = BeautifulSoup(page.text, "html.parser") # 提取页面标题作为CSV首行 title = doc.find("h1", class_="title").text.strip() # 定位目标数据:1985-2022年共38年×12月=456个数据点 data_points = doc.find_all("span", class_="datavalue") target_data = [point.text.strip() for point in data_points[651:651+456]] # 按年份拆分数据,每组12个月度值 yearly_groups = [target_data[i*12 : (i+1)*12] for i in range(38)] years = list(range(1985, 2023)) # 构造DataFrame,设置年份为索引,月度为列名 df = pd.DataFrame( yearly_groups, index=years, columns=["1月", "2月", "3月", "4月", "5月", "6月", "7月", "8月", "9月", "10月", "11月", "12月"] ) # 写入CSV:先写标题,再写结构化数据 with open("imports_excluding_petroleum.csv", "w", encoding="utf-8") as f: f.write(f"{title}\n") df.to_csv(f, sep=",", index_label="年份") print("文件生成完成:imports_excluding_petroleum.csv")
代码说明
- 数据定位:精准截取对应"All imports excluding petroleum"的1985-2022年数据段,共456个数据点,匹配38年×12月的需求
- 结构整理:将一维数据列表按年份拆分为38组,每组12个月度数据,直接对应CSV的行结构
- CSV输出:先写入标题行,再用pandas的
to_csv写入结构化数据,保留年份索引和月度列名,确保输出格式符合要求 - 编码处理:使用UTF-8编码写入,避免中文或特殊字符乱码
内容的提问来源于stack exchange,提问作者CaptainRyan1
相关产品推荐
相关产品推荐

