You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python re库统计DataFrame匹配词数时遇TypeError问题求助

Fix: TypeError: expected string or bytes-like object when counting pattern matches per row in DataFrame

What's Causing the Error?

I dug into your code and found two main issues triggering that TypeError:

  1. You're passing an entire Pandas Series (from df[dict[elt][0]].str.lower()) directly to re.findall — but re.findall only works with individual string/byte objects, not entire Series.
  2. Your current logic calculates a total match count for the whole column and assigns that single number to every row in the new column, which isn't what you want for per-row counting.

Core Issues in Your Original Code

  • Feeding a Series to re.findall instead of processing each row's string value one by one
  • Aggregating counts across the entire column instead of computing them per row

Corrected Working Code

We'll use apply to process each row individually, add safeguards for non-string values (like NaNs), and fix the counting logic. Here's the full version:

First, import required libraries:

import pandas as pd
import re
from collections import Counter

Define your pattern lists, configuration (note: don't use dict as a variable name — it overrides Python's built-in dict type!), and sample data:

list_1 = ['Apple', 'Mango' ,'Orange', 'pr[éeêè]t[s]?' ]
list_2 = ['weather', 'r[ea]d' ,'p[wr]iority', 'pr[éeêè]t[s]?' ]
list_3 = ['n[eéè]d','snow[s]?', 'pr[éeêè]t[s]?' ]

# Renamed from 'dict' to 'match_config' to avoid conflicts
match_config = {
    "s1": ['Column_1', list_1],
    "s2": ['Column_1', list_3],
    "s3": ['Column_2', list_2],
    "s4": ['Column_3', list_3],
    "s5": ['Column_2', 'Column_3', list_1],
}

# Sample DataFrame
d = {
    'Column_1': ['mango pret Orange No manner', ' préts No scan'],
    'Column_2': ['read priority No', 'This is a priority'],
    'Column_3': ['No add', 'yep']
}
df = pd.DataFrame(data=d)

Create a helper function to count matches for a single row, then apply it across the DataFrame:

def count_row_matches(row, target_columns, pattern_list):
    # Precompile regex with case-insensitive matching for better performance
    combined_pattern = re.compile(r'|'.join(pattern_list), re.IGNORECASE)
    total_matches = 0
    
    for col in target_columns:
        # Ensure cell value is a string to avoid errors with NaNs/non-strings
        cell_text = str(row[col])
        # Count all matches in the current cell
        matches = combined_pattern.findall(cell_text)
        total_matches += len(matches)
    
    return total_matches

# Iterate through config to add each result column
for result_col, config_vals in match_config.items():
    if len(config_vals) == 2:
        # Single column + pattern list scenario
        target_col, patterns = config_vals
        df[result_col] = df.apply(lambda row: count_row_matches(row, [target_col], patterns), axis=1)
    else:
        # Multiple columns + pattern list scenario (first N items are columns, last is patterns)
        target_cols = config_vals[:-1]
        patterns = config_vals[-1]
        df[result_col] = df.apply(lambda row: count_row_matches(row, target_cols, patterns), axis=1)

Expected Output

After running the code, your DataFrame will match your expected results (and include all missing columns):

Column_1               Column_2 Column_3  s1  s2  s3  s4  s5
0  mango pret Orange No manner       read priority No     No add     3  1  2  0  1
1                préts No scan  This is a priority        yep       1  1  1  0  0

Key Improvements

  • Renamed dict to match_config to prevent overriding Python's built-in dictionary type
  • Used re.compile to precompile regex patterns for faster repeated use
  • Added str(row[col]) to handle non-string cell values (like NaNs) without errors
  • Used apply(axis=1) to process each row individually, ensuring accurate per-row counts

内容的提问来源于stack exchange,提问作者This

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:38:28