You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取维基百科化学元素列表表格存入pandas DataFrame导出CSV问题求助

修正后可实现需求的完整代码

from bs4 import BeautifulSoup
import requests as r
import pandas as pd

response = r.get('https://en.wikipedia.org/wiki/List_of_chemical_elements')
wiki_text = response.text
soup = BeautifulSoup(wiki_text, 'html.parser')

table = soup.select_one('table.wikitable')
table_body = table.find('tbody')
rows = table_body.find_all('tr')

# 提取表头
cols = [c.text.replace('\n', '') for c in rows[1].find_all('th')]

# 初始化空列表存储所有行数据
all_records = []
for row in rows:
    records = row.find_all('td')
    if records:
        records = [r.text.strip() for r in records]
        all_records.append(records)

# 一次性将所有数据转为DataFrame
df2a = pd.DataFrame(all_records, columns=cols)

# 导出为CSV文件,指定编码避免乱码
df2a.to_csv('化学元素列表.csv', index=False, encoding='utf-8-sig')

关键改动说明

  • 没有逐行往空DataFrame里加数据,而是先把所有行存入列表后一次性生成DataFrame,运行效率更高,也不会出现行索引对齐的异常
  • 最后添加了to_csv方法直接导出为CSV文件
  • 如果你更习惯逐行添加的写法,也可以把循环部分替换为如下内容:
df2a = pd.DataFrame(columns = cols)
for row in rows:
    records = row.find_all('td')
    if records:
        records = [r.text.strip() for r in records]
        # 用loc指定行位置添加数据
        df2a.loc[len(df2a)] = records

内容的提问来源于stack exchange,提问作者Jeff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 02:15:02