如何在Python Pandas中过滤DataFrame列转换列表中的非英文句子?
Got it, let's work through filtering out non-English sentences from your Q8 list. I've got two practical methods for you, depending on how strict you need the filtering to be:
Method 1: Regex-Based Filtering (Strict Match)
If you want to only keep sentences made up entirely of English characters and common punctuation, regex is a quick, no-external-library way to go. Here's how to implement it:
import re import pandas as pd # 先读取Excel并处理空值再转列表,避免后续报错 df = pd.read_excel('your_file.xlsx') df_lst = df['Q8 Why do you say so ?'].dropna().values.tolist() # 定义判断英文的函数 def is_english(text): # 匹配英文字母、空格和常见标点符号,可根据需求调整正则 english_regex = re.compile(r'^[a-zA-Z\s.!?\',;-]+$') return bool(english_regex.match(str(text))) # 过滤得到纯英文句子列表 filtered_english_lst = [sentence for sentence in df_lst if is_english(sentence)]
Note: If your valid English responses might include numbers or other symbols, just add those characters to the regex pattern (e.g., add 0-9 inside the square brackets).
Method 2: Language Detection (Smarter Match)
For a more flexible approach—where sentences might have minor non-English characters but are primarily English—use the langdetect library. It’s better at identifying actual language context instead of just character sets.
First, install the library if you haven't already:
pip install langdetect
Then use this code:
from langdetect import detect, LangDetectException import pandas as pd df = pd.read_excel('your_file.xlsx') df_lst = df['Q8 Why do you say so ?'].dropna().values.tolist() def is_english(text): try: # 检测语言,返回'en'表示英文 return detect(str(text)) == 'en' except LangDetectException: # 处理无法识别的文本(比如空字符串、乱码) return False filtered_english_lst = [sentence for sentence in df_lst if is_english(sentence)]
This method handles edge cases better, like short phrases or sentences with occasional non-English symbols. Just keep in mind that very short or gibberish text might throw a detection error, hence the try-except block to skip those entries.
Quick Tip
Always start by dropping null values from your column (like I did with dropna())—this avoids unnecessary errors when processing empty cells later on.
内容的提问来源于stack exchange,提问作者explorer_x

