You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas str.findall返回全NA 如何统计新闻文本对应LIWC类别词频

问题原因

原有代码返回全NA主要有三个错误:

  • 正则拼接对象错误:'|'.join(df2)会拼接df2的列名而非词汇列表,应该取df2['Word']列拼接
  • 多余的单引号匹配:正则里加了'包裹匹配规则,但新闻文本中的词汇没有被单引号包裹,完全匹配不到结果
  • 统计逻辑错误:直接对匹配到的词汇做Counter后,用分类名当索引,无法对应上词汇所属的分类,统计结果必然错位

Pandas 实现代码

import pandas as pd
import re
from collections import Counter

# 1. 构建词汇到LIWC分类的映射字典
word_to_cat = dict(zip(df2['Word'], df2['Category']))
# 2. 构建匹配正则,加\b单词边界避免部分匹配,加re.escape处理词汇中的特殊字符
pattern = re.compile(r'\b(' + '|'.join(re.escape(word) for word in df2['Word']) + r')\b')
# 3. 定义单篇文本的分类统计函数
def count_liwc(text):
    matched_words = pattern.findall(text)
    # 词汇转分类后统计频次
    cat_counts = Counter(word_to_cat[word] for word in matched_words)
    return pd.Series(cat_counts)

# 4. 统计结果和原文本合并,空值填0
liwc_result = df1.join(df1['articles'].apply(count_liwc)) \
                 .fillna(0) \
                 .astype({cat:int for cat in df2['Category'].unique()})

R 实现参考

library(dplyr)
library(stringr)
library(tidyr)

# 1. 构建匹配正则
pattern <- str_c("\\b(", str_c(df2$Word, collapse = "|"), ")\\b")
# 2. 逐篇统计分类频次
liwc_result <- df1 %>%
  mutate(matched_word = str_extract_all(articles, pattern)) %>%
  unnest(matched_word) %>%
  left_join(df2, by = c("matched_word" = "Word")) %>%
  count(articles, Category) %>%
  pivot_wider(names_from = Category, values_from = n, values_fill = 0) %>%
  right_join(df1, by = "articles") %>%
  mutate(across(all_of(unique(df2$Category)), ~replace_na(., 0)))

内容的提问来源于stack exchange,提问作者user17143533

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 15:45:08