如何在PySpark DataFrame中展开文本中的常见英语缩写?
PySpark 展开英语缩写实现方案
方法一:使用内置正则替换函数(高效推荐)
通过regexp_replace结合reduce批量处理无歧义的缩写,适合大多数场景,性能优于自定义UDF。
- 定义缩写映射字典(用
\b匹配单词边界,避免误替换) - 利用
reduce迭代应用所有替换规则
from pyspark.sql import SparkSession from pyspark.sql.functions import regexp_replace from functools import reduce # 初始化Spark会话 spark = SparkSession.builder.appName("AbbrevExpand").getOrCreate() # 创建示例DataFrame temp = spark.createDataFrame([ (0, "Julia isn't awesome"), (1, "I wish Java-DL couldn't use case-classes"), (2, "Data-science wasn't my subject"), (3, "Machine") ], ["id", "words"]) # 缩写映射:键为正则匹配模式,值为展开后的文本 abbrev_mapping = { r"\bisn't\b": "is not", r"\bcouldn't\b": "could not", r"\bwasn't\b": "was not", # 可根据需求添加更多缩写,比如r"\bdon't\b": "do not" } # 批量替换函数 def apply_abbrev_replacements(df, mapping): return reduce(lambda current_df, (pattern, replacement): current_df.withColumn("words", regexp_replace(current_df["words"], pattern, replacement)), mapping.items(), df) # 生成展开后的DataFrame expanded_df = apply_abbrev_replacements(temp, abbrev_mapping) # 查看结果 expanded_df.show(truncate=False)
执行后输出:
+---+-----------------------------------------+ |id |words | +---+-----------------------------------------+ |0 |Julia is not awesome | |1 |I wish Java-DL could not use case-classes| |2 |Data-science was not my subject | |3 |Machine | +---+-----------------------------------------+
方法二:自定义UDF(处理歧义缩写)
如果遇到像I'd这种存在两种展开可能(I would/I had)的歧义缩写,可以用自定义UDF结合上下文判断,灵活性更高。
from pyspark.sql.functions import udf from pyspark.sql.types import StringType import re def expand_ambiguous_abbrevs(text): # 先处理无歧义缩写 basic_mapping = { r"\bisn't\b": "is not", r"\bcouldn't\b": "could not", r"\bwasn't\b": "was not" } for pattern, repl in basic_mapping.items(): text = re.sub(pattern, repl, text) # 处理歧义缩写I'd:根据上下文判断是I had还是I would def replace_id(match): # 简单判断上下文是否存在had/have/has等词,决定展开形式 if re.search(r"\b(had|have|has)\b", text, flags=re.IGNORECASE): return "I had" else: return "I would" text = re.sub(r"\bI'd\b", replace_id, text) return text # 注册UDF expand_abbrev_udf = udf(expand_ambiguous_abbrevs, StringType()) # 应用UDF expanded_df = temp.withColumn("words", expand_abbrev_udf(temp["words"])) expanded_df.show(truncate=False)
注意事项
- 内置函数方案性能更优,优先用于大数据量场景
- 正则模式中的
\b是单词边界,避免替换嵌入在其他单词中的类似字符串 - 歧义缩写的上下文判断逻辑可根据实际需求优化
内容的提问来源于stack exchange,提问作者merkle
相关产品推荐
相关产品推荐

