You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup对HTML文本内容进行查找替换?

HTML文本内容的查找替换(Perl Mojo::DOM 修复方案)

需求

对HTML元素内的纯文本部分执行查找替换,例如将foo替换为<b>bar</b>:

  • 输入:
<div id="foo">foo <i>foo</i> hi foo hi</div>
  • 预期输出:
<div id="foo"><b>bar</b> <i><b>bar</b></i> hi <b>bar</b> hi</div>

原代码问题

你提供的Mojo::DOM实现会将替换后的HTML标签转义,导致无法生成预期的嵌套元素。问题出在replace方法直接传入HTML字符串时,Mojo::DOM会把它当作纯文本处理:

#!/usr/bin/env perl
##
use strict;
use warnings;
use v5.34.0;

use Mojo::DOM;
##
my $input = do { local $/; <STDIN> };

my $dom = Mojo::DOM->new($input);

$dom->descendant_nodes->grep(sub { $_->type eq 'text' })
    ->each(sub{
        $_->replace(s/(sth)/<span class="todo at_tag">$1<\/span>/gr)
           });

say $dom;

修复后的代码

要正确插入HTML元素,需要拆分文本节点并创建对应的元素节点,而非直接替换为字符串:

#!/usr/bin/env perl
use strict;
use warnings;
use v5.34.0;

use Mojo::DOM;

my $input = do { local $/; <STDIN> };
my $dom = Mojo::DOM->new($input);

# 遍历所有文本节点
$dom->descendant_nodes->grep(sub { $_->type eq 'text' })->each(sub {
    my $text_node = shift;
    my $content = $text_node->content;
    return unless $content =~ /foo/;

    my @new_nodes;
    # 拆分文本,分离出匹配项
    while ($content =~ /(.*?)(foo)(.*)/s) {
        # 添加匹配前的文本(如果非空)
        push @new_nodes, Mojo::DOM::TextNode->new($1) if length $1;
        # 创建替换后的<b>bar</b>元素节点
        push @new_nodes, Mojo::DOM->new('<b>bar</b>')->[0];
        $content = $3;
    }
    # 添加剩余的文本(如果非空)
    push @new_nodes, Mojo::DOM::TextNode->new($content) if length $content;

    # 用新节点列表替换原文本节点
    $text_node->replace(@new_nodes);
});

print $dom->to_string;

关键改进点

  • 避免直接用HTML字符串替换文本节点,而是拆分文本为多个片段和元素节点
  • 使用Mojo::DOM::TextNode创建纯文本节点,用Mojo::DOM->new()创建元素节点
  • 确保替换后的元素被正确解析为DOM节点,而非转义的文本

内容的提问来源于stack exchange,提问作者HappyFace

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:25:21