You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP:合并多数组数据为单个数组并过滤无效哈希标签

解决方案

步骤1:重写哈希标签提取函数,返回有效数组

原函数用explode拆分的方式容易生成无效标签,改用正则匹配更精准,同时让函数返回数组而非直接输出,方便后续合并操作。

function fetchHashtags($getData) {
    $validHashtags = [];
    // 匹配规则:#开头,后面必须紧跟字母/数字,允许中间加下划线(但不能只有下划线)
    preg_match_all('/#(?=[\p{L}\p{N}])[\p{L}\p{N}_]+/u', $getData, $matches);
    
    if (!empty($matches[0])) {
        // 过滤掉纯下划线、单个非字母数字的无效标签(比如#_、#۰)
        foreach ($matches[0] as $tag) {
            $tagContent = substr($tag, 1);
            if ($tagContent !== '_' && preg_match('/[\p{L}\p{N}]/u', $tagContent)) {
                $validHashtags[] = $tag;
            }
        }
    }
    
    return $validHashtags;
}

正则规则说明

  • #:匹配哈希标签的前缀
  • (?=[\p{L}\p{N}]):确保#后面紧跟字母或数字,直接排除#_、#۰这类无效开头
  • [\p{L}\p{N}_]+:支持匹配Unicode字母、数字和下划线,适配非英文标签(比如波斯语、葡萄牙语)
  • u修饰符:开启Unicode字符支持,避免乱码

步骤2:合并多个数组为一维数组

如果多次调用fetchHashtags得到了多个独立数组,用array_merge把它们合并成一个一维数组,还可以按需用array_unique去重:

// 示例:多个待处理的文本内容
$text1 = "Residência #architecture #_ 测试#۰";
$text2 = "#We #ascaspenvswheaton mountainhouse";
$text3 = "#شجریان_بشنویم ... #interiores";

// 分别提取每个文本的有效标签
$tags1 = fetchHashtags($text1);
$tags2 = fetchHashtags($text2);
$tags3 = fetchHashtags($text3);

// 合并为单个一维数组
$mergedTags = array_merge($tags1, $tags2, $tags3);

// 可选:去除重复标签
$mergedTags = array_unique($mergedTags);

// 查看最终结果
print_r($mergedTags);

效果验证

针对你提供的待过滤数组:

Array
(
    [0] => #Residência
    [1] => architecture // 需移除
    [2] => #casanaserra
    [3] => mountainhouse // 需移除
    [4] => #interiores
)

处理后会得到纯净的有效标签数组:

Array
(
    [0] => #Residência
    [1] => #casanaserra
    [2] => #interiores
)

同时#_、#۰这类无效标签会被过滤,多次调用生成的独立数组也会合并成一个一维数组,不会再出现零散输出的情况。

内容的提问来源于stack exchange,提问作者Sonny's Kitchen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 20:51:13