You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Pandas中过滤DataFrame列转换列表中的非英文句子?

Filtering Non-English Sentences from Your Pandas DataFrame List

Got it, let's work through filtering out non-English sentences from your Q8 list. I've got two practical methods for you, depending on how strict you need the filtering to be:

Method 1: Regex-Based Filtering (Strict Match)

If you want to only keep sentences made up entirely of English characters and common punctuation, regex is a quick, no-external-library way to go. Here's how to implement it:

import re
import pandas as pd

# 先读取Excel并处理空值再转列表,避免后续报错
df = pd.read_excel('your_file.xlsx')
df_lst = df['Q8 Why do you say so ?'].dropna().values.tolist()

# 定义判断英文的函数
def is_english(text):
    # 匹配英文字母、空格和常见标点符号,可根据需求调整正则
    english_regex = re.compile(r'^[a-zA-Z\s.!?\',;-]+$')
    return bool(english_regex.match(str(text)))

# 过滤得到纯英文句子列表
filtered_english_lst = [sentence for sentence in df_lst if is_english(sentence)]

Note: If your valid English responses might include numbers or other symbols, just add those characters to the regex pattern (e.g., add 0-9 inside the square brackets).

Method 2: Language Detection (Smarter Match)

For a more flexible approach—where sentences might have minor non-English characters but are primarily English—use the langdetect library. It’s better at identifying actual language context instead of just character sets.

First, install the library if you haven't already:

pip install langdetect

Then use this code:

from langdetect import detect, LangDetectException
import pandas as pd

df = pd.read_excel('your_file.xlsx')
df_lst = df['Q8 Why do you say so ?'].dropna().values.tolist()

def is_english(text):
    try:
        # 检测语言,返回'en'表示英文
        return detect(str(text)) == 'en'
    except LangDetectException:
        # 处理无法识别的文本(比如空字符串、乱码)
        return False

filtered_english_lst = [sentence for sentence in df_lst if is_english(sentence)]

This method handles edge cases better, like short phrases or sentences with occasional non-English symbols. Just keep in mind that very short or gibberish text might throw a detection error, hence the try-except block to skip those entries.

Quick Tip

Always start by dropping null values from your column (like I did with dropna())—this avoids unnecessary errors when processing empty cells later on.

内容的提问来源于stack exchange,提问作者explorer_x

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:35:23