You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark DataFrame中展开文本中的常见英语缩写?

PySpark 展开英语缩写实现方案

方法一:使用内置正则替换函数(高效推荐)

通过regexp_replace结合reduce批量处理无歧义的缩写,适合大多数场景,性能优于自定义UDF。

  1. 定义缩写映射字典(用\b匹配单词边界,避免误替换)
  2. 利用reduce迭代应用所有替换规则
from pyspark.sql import SparkSession
from pyspark.sql.functions import regexp_replace
from functools import reduce

# 初始化Spark会话
spark = SparkSession.builder.appName("AbbrevExpand").getOrCreate()

# 创建示例DataFrame
temp = spark.createDataFrame([
    (0, "Julia isn't awesome"),
    (1, "I wish Java-DL couldn't use case-classes"),
    (2, "Data-science wasn't my subject"),
    (3, "Machine")
], ["id", "words"])

# 缩写映射:键为正则匹配模式,值为展开后的文本
abbrev_mapping = {
    r"\bisn't\b": "is not",
    r"\bcouldn't\b": "could not",
    r"\bwasn't\b": "was not",
    # 可根据需求添加更多缩写,比如r"\bdon't\b": "do not"
}

# 批量替换函数
def apply_abbrev_replacements(df, mapping):
    return reduce(lambda current_df, (pattern, replacement): 
                  current_df.withColumn("words", regexp_replace(current_df["words"], pattern, replacement)),
                  mapping.items(), df)

# 生成展开后的DataFrame
expanded_df = apply_abbrev_replacements(temp, abbrev_mapping)

# 查看结果
expanded_df.show(truncate=False)

执行后输出:

+---+-----------------------------------------+
|id |words                                    |
+---+-----------------------------------------+
|0  |Julia is not awesome                     |
|1  |I wish Java-DL could not use case-classes|
|2  |Data-science was not my subject          |
|3  |Machine                                  |
+---+-----------------------------------------+

方法二:自定义UDF(处理歧义缩写)

如果遇到像I'd这种存在两种展开可能(I would/I had)的歧义缩写,可以用自定义UDF结合上下文判断,灵活性更高。

from pyspark.sql.functions import udf
from pyspark.sql.types import StringType
import re

def expand_ambiguous_abbrevs(text):
    # 先处理无歧义缩写
    basic_mapping = {
        r"\bisn't\b": "is not",
        r"\bcouldn't\b": "could not",
        r"\bwasn't\b": "was not"
    }
    for pattern, repl in basic_mapping.items():
        text = re.sub(pattern, repl, text)
    
    # 处理歧义缩写I'd:根据上下文判断是I had还是I would
    def replace_id(match):
        # 简单判断上下文是否存在had/have/has等词,决定展开形式
        if re.search(r"\b(had|have|has)\b", text, flags=re.IGNORECASE):
            return "I had"
        else:
            return "I would"
    
    text = re.sub(r"\bI'd\b", replace_id, text)
    return text

# 注册UDF
expand_abbrev_udf = udf(expand_ambiguous_abbrevs, StringType())

# 应用UDF
expanded_df = temp.withColumn("words", expand_abbrev_udf(temp["words"]))
expanded_df.show(truncate=False)

注意事项

  • 内置函数方案性能更优,优先用于大数据量场景
  • 正则模式中的\b是单词边界,避免替换嵌入在其他单词中的类似字符串
  • 歧义缩写的上下文判断逻辑可根据实际需求优化

内容的提问来源于stack exchange,提问作者merkle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 04:18:13