You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas extractall函数时遇索引层级不一致拼接错误求助

解决pandas分组提取日期时的“Cannot concat indices that do not have the same number of levels”错误

我正在编写一个读取Outlook邮件并提取信息到pandas DataFrame的脚本,在提取邮件正文中的日期时遇到了错误。

我的操作代码如下:

# 按主题分组邮件,将单个邮件归为线程组
dfgroup = df.groupby('Subject')
# 尝试提取邮件正文中提到的所有日期
temp = dfgroup['Message'].apply(lambda x: x.str.extractall(r'(?P<extract>(?P<month>(January|February|March|April|May|June|July|August|September|October|November|December))\s(?P<date>\d{2})\,\s(?P<year>\d{4})\s(?P<time>\d{1,2}\:\d{2}\s(PM|AM)))'))

收到的错误信息:

File "C:\Users\tioxr\AppData\Local\Continuum\Anaconda3\lib\site-packages\pandas\core\reshape\concat.py", line 573, in _make_concat_multiindex
raise AssertionError("Cannot concat indices that do"
AssertionError: Cannot concat indices that do not have the same number of levels


问题原因

这个错误的核心是str.extractall()方法会返回一个带有双重索引的DataFrame:一层是原DataFrame的行索引,另一层是匹配项的编号(每个行里的第N个匹配)。当你在分组apply这个函数时,不同分组返回的结果在索引层级上可能出现不一致,导致pandas在拼接这些结果时抛出索引层级不匹配的错误。

解决方案

方案一:移除extractall生成的匹配索引层级

修改apply里的lambda函数,在extractall之后调用reset_index()移除多余的索引层级,确保每个分组返回的结果索引结构一致:

dfgroup = df.groupby('Subject')
temp = dfgroup['Message'].apply(
    lambda x: x.str.extractall(r'(?P<extract>(?P<month>(January|February|March|April|May|June|July|August|September|October|November|December))\s(?P<date>\d{2})\,\s(?P<year>\d{4})\s(?P<time>\d{1,2}\:\d{2}\s(PM|AM)))')
    .reset_index(level=1, drop=True)  # 移除匹配项的索引层级
)

这里reset_index(level=1, drop=True)专门移除extractall新增的匹配编号索引,只保留原行索引,这样所有分组的结果索引层级统一,拼接时就不会报错了。

方案二:先全局提取再分组(更高效)

如果不需要在分组内逐个处理,推荐先对整个Message列提取所有日期,再关联原数据进行分组,这种方式逻辑更清晰,性能也更好:

# 全局提取所有日期,保留原行索引
extracted_dates = df['Message'].str.extractall(r'(?P<extract>(?P<month>(January|February|March|April|May|June|July|August|September|October|November|December))\s(?P<date>\d{2})\,\s(?P<year>\d{4})\s(?P<time>\d{1,2}\:\d{2}\s(PM|AM)))')

# 关联原DataFrame的Subject列,再按主题分组
temp = extracted_dates.join(df['Subject']).groupby('Subject')

这种方式绕过了分组apply时的索引拼接问题,同时也避免了apply带来的性能损耗,适合处理较大的数据集。


内容的提问来源于stack exchange,提问作者Sidney Tio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:35:37