You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拼接三个Pandas DataFrame避免出现行不匹配问题

问题原因

你拼接错位的核心原因是三个DataFrame的行索引不一致,pd.concat(axis=1)是按照索引值对齐行的:

  • 你构造的df0是默认从0开始的整数索引,唯一一行的索引值是0
  • 你从网页读取的table1、table2都是典型的key-value两列表格,用index_col=0读取后转置得到的df1、df2,唯一一行的索引值是原表格的列名,不是0
  • 拼接时索引不匹配就会生成两行,分别对应索引0和原列名,就出现了你截图里df0和df1、df2错位的情况

修复方法

只需要在转置df1、df2之后重置索引,丢弃原来的索引值,让三个DataFrame的索引统一即可,修改后代码如下:

import requests
import pandas as pd
from bs4 import BeautifulSoup

List = ['LU0526609390:EUR', 'IE00BHBX0Z19:EUR', 'LU1076093779:EUR', 'LU1116896363:EUR']
df = pd.DataFrame(List, columns=['List'])
urls = 'https://markets.ft.com/data/funds/tearsheet/summary?s='+ df['List']

dfs =[]
for url in urls:
    print(url)
    r = requests.get(url).content
    soup = BeautifulSoup(r, 'html.parser')
    elemList = soup.find('title')
    df0 = pd.DataFrame(elemList, columns = ['Fund Name'])
    df0["Fund Name"] = df0["Fund Name"].str.replace("summary - FT.com", "", regex=True)
    table1 = soup.find_all('table')[0]
    table2 = soup.find_all('table')[1]
    # 转置后重置索引,和df0对齐
    df1 = pd.read_html(str(table1), index_col=0)[0].T.reset_index(drop=True)
    df2 = pd.read_html(str(table2), index_col=0)[0].T.reset_index(drop=True)
    df = pd.concat([df0, df1, df2], axis=1)
    dfs.append(df)

pd.concat(dfs).to_csv(r'/Users/Test.csv', index=False)

内容的提问来源于stack exchange,提问作者Tcs106

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 17:27:03