You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Regex从Pandas句子列提取单词用于网络分析的报错问题求助

报错原因
  1. 直接遍历base_network时默认取的是DataFrame的列名,而非逐行读取数据
  2. re.split()第二个参数传入了整个Body列的Series对象,而非单个字符串,不符合方法入参要求
  3. 即使提前转换了Body列为字符串类型,调用方法时传入的还是整列数据,依然会触发类型报错
最优解决方案

Pandas 0.25及以上版本可直接用str.split()+explode()的组合实现需求,无需手动写循环,执行效率更高:

import pandas as pd
import re

# 先确保Body列全为字符串类型,避免空值/非字符串值导致报错
base_network['Body'] = base_network['Body'].astype(str)

# 按空格拆分句子为单词列表,再将列表炸开为每行一个单词,自动保留对应Rating
result_df = base_network.assign(
    单词 = base_network['Body'].str.split(r'\s+')
).explode('单词').reset_index(drop=True)

# 可根据需要删除原Body列:result_df = result_df.drop(columns=['Body'])

# 查看结果
print(result_df.head())
原循环写法的修正(仅做错误演示,不推荐生产使用)

如果要沿用循环逻辑,需要改为逐行遍历的写法,仅适合小数据量场景:

spaces = r"\s+"
df = pd.DataFrame()

# 用iterrows逐行遍历
for idx, row in base_network.iterrows():
    # 拆分当前行的Body字符串
    words = re.split(spaces, str(row['Body']))
    # 构造临时DataFrame
    temp_df = pd.DataFrame({
        '单词': words,
        'Rating': [row['Rating']]*len(words)
    })
    df = pd.concat([df, temp_df], ignore_index=True)

print(df.head())

注意:append方法已在Pandas 2.0+中废弃,建议使用pd.concat做数据合并。

内容的提问来源于stack exchange,提问作者Binne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 20:54:05