如何检查DataFrame/ndarray中的邮箱是否含数字并提取数字?
提取邮箱地址中的数字到数组的解决方案
我来帮你搞定这个问题!其实用 pandas 的字符串处理工具结合正则表达式就能轻松解决,不用折腾复杂的 numpy 操作~
步骤1:先清理数据(和你已有的操作衔接)
你已经在做 dropna 了,先把空值清理干净,避免后续报错:
# 假设你的邮箱列是 mail_addresses 的第0列(和你的代码对应) mail_addresses = mail_addresses.dropna(axis=0)
步骤2:提取邮箱中的数字
根据你的需求,分两种情况处理:
情况A:提取邮箱里的连续数字序列(比如 roberto123@example.com 提取出 "123")
用 str.extract 配合正则 (\d+),它会匹配第一个连续数字串,没有数字的会返回 NaN:
# 提取数字,expand=False 让结果返回 Series 而不是 DataFrame extracted_numbers = mail_addresses[0].str.extract(r'(\d+)', expand=False)
情况B:提取邮箱里的所有单个数字(比如 ro1be2rto3@test.com 提取出 ["1","2","3"])
用 str.findall 配合正则 \d,它会返回每个邮箱对应的数字列表:
extracted_numbers_list = mail_addresses[0].str.findall(r'\d')
步骤3:把数字转成你需要的数组
如果是情况A,先去掉空值,再转成列表或 numpy 数组:
# 转成 Python 列表 numbers_list = extracted_numbers.dropna().tolist() # 转成 numpy ndarray import numpy as np numbers_ndarray = np.array(numbers_list)
如果是情况B,要把所有数字合并成一个数组的话,可以这样:
# 把所有子列表里的数字合并成一个大列表 all_numbers = [num for sublist in extracted_numbers_list for num in sublist] numbers_ndarray = np.array(all_numbers)
举个完整例子
假设你的 mail_addresses 数据是这样的:
| 0 |
|---|
| roberto123@example.com |
| alice@test.com |
| bob4567@domain.org |
| charlie89@xyz.net |
运行情况A的代码后,numbers_list 会是 ["123", "4567", "89"],完全符合你的需求~
之前没成功的原因大概率是没用到 pandas 专门的字符串处理方法(str.xxx),或者正则表达式没写对,试试上面的方法应该就能解决啦!
内容的提问来源于stack exchange,提问作者robsanna
相关产品推荐
相关产品推荐

