You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python移除Pandas中含非英文单词的行

移除Pandas DataFrame中含非英文单词的行

需求:移除Pandas DataFrame中包含非英文单词的行(每行是已分词的句子列表)。

示例数据集

data = {
  "col1": [['apartment', 'expectations', 'insinuate', 'welcome'], ['très', 'réactive', 'arrangeante', 'notre','place'],['buena', 'ubicación','you']]
}

# 加载数据到DataFrame
df = pd.DataFrame(data)

print(df)

尝试的错误代码及报错

尝试了以下仅支持Python3.7+的代码,但触发错误:

# 仅支持Python3.7及以上版本
df[df.col1.map(lambda x: x.isascii())] 

报错信息:

AttributeError: 'list' object has no attribute 'isascii'

解决方法

方法1:检查列表中所有单词均为ASCII字符

col1的每个元素是列表,不能直接调用isascii(),需遍历列表内每个单词,确认所有单词都符合ASCII标准:

# 过滤出所有单词都是ASCII的行
filtered_df = df[df['col1'].apply(lambda x: all(word.isascii() for word in x))]
print(filtered_df)

方法2:用正则匹配纯英文单词(仅含a-z/A-Z)

若需严格匹配仅由英文字母组成的单词(不含特殊字符、数字),可使用正则表达式:

import re

def is_all_english_words(word_list):
    pattern = re.compile(r'^[a-zA-Z]+$')
    return all(pattern.match(word) for word in word_list)

filtered_df = df[df['col1'].apply(is_all_english_words)]
print(filtered_df)

预期输出

执行上述代码后得到预期结果:

col1
0  [apartment, expectations, insinuate, welcome]

内容的提问来源于stack exchange,提问作者Todd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 03:48:24