You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/PySpark获取列表中包含指定子串的完整字符串

Python与PySpark实现搜索包含指定子串的字符串

针对列表 ['abc/aed','bcd/eac','rtf/reew','opee/rew'],以下是搜索包含子串'aed'的字符串并返回完整匹配项的实现方法:

Python实现

方法1:列表推导式筛选

直接通过列表推导式过滤出所有包含目标子串的元素,再按需返回结果:

target_list = ['abc/aed','bcd/eac','rtf/reew','opee/rew']
# 筛选所有包含'aed'的字符串
matched = [item for item in target_list if 'aed' in item]

# 判断并输出结果
if matched:
    print(matched[0])  # 取第一个匹配项,输出 'abc/aed'
    # 若要返回所有匹配项,直接使用matched即可
else:
    print("不存在包含子串'aed'的字符串")

方法2:循环遍历查找

通过遍历列表逐个检查,找到第一个匹配项后终止循环:

target_list = ['abc/aed','bcd/eac','rtf/reew','opee/rew']
found_item = None

for item in target_list:
    if 'aed' in item:
        found_item = item
        break

if found_item:
    print(found_item)  # 输出 'abc/aed'
else:
    print("不存在包含子串'aed'的字符串")

PySpark实现

通过Spark DataFrame的contains方法筛选匹配项:

from pyspark.sql import SparkSession
from pyspark.sql.functions import col

# 初始化Spark会话
spark = SparkSession.builder.appName("SubstringSearch").getOrCreate()

# 构造数据源并创建DataFrame
data = [('abc/aed',), ('bcd/eac',), ('rtf/reew',), ('opee/rew',)]
df = spark.createDataFrame(data, schema=["str_value"])

# 筛选包含'aed'的行
result_df = df.filter(col("str_value").contains("aed"))

# 收集结果并输出
result = result_df.collect()
if result:
    print(result[0][0])  # 输出 'abc/aed'
else:
    print("不存在包含子串'aed'的字符串")

# 关闭Spark会话
spark.stop()

内容的提问来源于stack exchange,提问作者Swati B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:05:31