使用Python re库统计DataFrame匹配词数时遇TypeError问题求助
Fix: TypeError: expected string or bytes-like object when counting pattern matches per row in DataFrame
What's Causing the Error?
I dug into your code and found two main issues triggering that TypeError:
- You're passing an entire Pandas Series (from
df[dict[elt][0]].str.lower()) directly tore.findall— butre.findallonly works with individual string/byte objects, not entire Series. - Your current logic calculates a total match count for the whole column and assigns that single number to every row in the new column, which isn't what you want for per-row counting.
Core Issues in Your Original Code
- Feeding a Series to
re.findallinstead of processing each row's string value one by one - Aggregating counts across the entire column instead of computing them per row
Corrected Working Code
We'll use apply to process each row individually, add safeguards for non-string values (like NaNs), and fix the counting logic. Here's the full version:
First, import required libraries:
import pandas as pd import re from collections import Counter
Define your pattern lists, configuration (note: don't use dict as a variable name — it overrides Python's built-in dict type!), and sample data:
list_1 = ['Apple', 'Mango' ,'Orange', 'pr[éeêè]t[s]?' ] list_2 = ['weather', 'r[ea]d' ,'p[wr]iority', 'pr[éeêè]t[s]?' ] list_3 = ['n[eéè]d','snow[s]?', 'pr[éeêè]t[s]?' ] # Renamed from 'dict' to 'match_config' to avoid conflicts match_config = { "s1": ['Column_1', list_1], "s2": ['Column_1', list_3], "s3": ['Column_2', list_2], "s4": ['Column_3', list_3], "s5": ['Column_2', 'Column_3', list_1], } # Sample DataFrame d = { 'Column_1': ['mango pret Orange No manner', ' préts No scan'], 'Column_2': ['read priority No', 'This is a priority'], 'Column_3': ['No add', 'yep'] } df = pd.DataFrame(data=d)
Create a helper function to count matches for a single row, then apply it across the DataFrame:
def count_row_matches(row, target_columns, pattern_list): # Precompile regex with case-insensitive matching for better performance combined_pattern = re.compile(r'|'.join(pattern_list), re.IGNORECASE) total_matches = 0 for col in target_columns: # Ensure cell value is a string to avoid errors with NaNs/non-strings cell_text = str(row[col]) # Count all matches in the current cell matches = combined_pattern.findall(cell_text) total_matches += len(matches) return total_matches # Iterate through config to add each result column for result_col, config_vals in match_config.items(): if len(config_vals) == 2: # Single column + pattern list scenario target_col, patterns = config_vals df[result_col] = df.apply(lambda row: count_row_matches(row, [target_col], patterns), axis=1) else: # Multiple columns + pattern list scenario (first N items are columns, last is patterns) target_cols = config_vals[:-1] patterns = config_vals[-1] df[result_col] = df.apply(lambda row: count_row_matches(row, target_cols, patterns), axis=1)
Expected Output
After running the code, your DataFrame will match your expected results (and include all missing columns):
Column_1 Column_2 Column_3 s1 s2 s3 s4 s5 0 mango pret Orange No manner read priority No No add 3 1 2 0 1 1 préts No scan This is a priority yep 1 1 1 0 0
Key Improvements
- Renamed
dicttomatch_configto prevent overriding Python's built-in dictionary type - Used
re.compileto precompile regex patterns for faster repeated use - Added
str(row[col])to handle non-string cell values (like NaNs) without errors - Used
apply(axis=1)to process each row individually, ensuring accurate per-row counts
内容的提问来源于stack exchange,提问作者This
相关产品推荐
相关产品推荐

