You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup结合CSV的anchor_text与URL为HTML添加超链接

基于Python BeautifulSoup实现批量为HTML文本添加指定超链接

需求说明

根据包含anchor_text和hyperlink两列的CSV文件,为HTML文本中的目标内容批量添加超链接,需满足以下规则:

  • 跳过已包含在<a>标签内的文本(不修改已有链接)
  • 支持大小写不敏感匹配(如CSV中的"Bing"可匹配HTML里的"bing")
  • 优先匹配多词锚文本(如"Active Campaign",避免被拆分为单个词匹配)
  • 保留原文本的大小写格式

实现代码

1. 依赖安装

首先安装所需库:

pip install beautifulsoup4 pandas

2. 完整代码

from bs4 import BeautifulSoup, NavigableString
import pandas as pd
import re
from io import StringIO

def load_link_mappings(csv_source):
    """读取并整理CSV中的锚文本-链接映射"""
    # 支持读取CSV文件路径或字符串内容
    if isinstance(csv_source, str) and '\n' not in csv_source:
        df = pd.read_csv(csv_source)
    else:
        df = pd.read_csv(StringIO(csv_source))
    
    # 从hyperlink列的a标签中提取真实URL(兼容带标签和纯URL两种格式)
    df['hyperlink'] = df['hyperlink'].apply(
        lambda x: BeautifulSoup(x, 'html.parser').a['href'] if '<a' in x else x
    )
    
    # 按锚文本长度降序排序,优先匹配长文本,避免多词内容被拆分
    df = df.sort_values(by='anchor_text', key=lambda x: x.str.len(), ascending=False)
    
    # 构建正则匹配规则,支持大小写不敏感+完整单词匹配
    mappings = []
    for _, row in df.iterrows():
        anchor = row['anchor_text']
        # 使用单词边界\b确保匹配完整文本,避免部分匹配(如"Google"不会匹配"Google123")
        pattern = re.compile(rf'\b{re.escape(anchor)}\b', re.IGNORECASE)
        mappings.append((pattern, row['hyperlink'], anchor))
    return mappings

def add_links_to_html(html_content, mappings):
    """为HTML文本中的目标内容添加超链接"""
    soup = BeautifulSoup(html_content, 'html.parser')
    
    # 递归遍历节点,仅处理非a标签内的文本节点
    def traverse(node):
        if isinstance(node, NavigableString):
            text = node.string
            if not text:
                return
            
            # 遍历所有映射规则,匹配成功则替换
            for pattern, url, _ in mappings:
                matches = list(pattern.finditer(text))
                if not matches:
                    continue
                
                # 从后往前替换,避免索引偏移问题
                new_nodes = []
                last_pos = 0
                for match in reversed(matches):
                    start, end = match.span()
                    # 添加匹配前的文本
                    new_nodes.append(NavigableString(text[last_pos:start]))
                    # 创建a标签并保留原文本大小写
                    a_tag = soup.new_tag('a', href=url)
                    a_tag.string = text[start:end]
                    new_nodes.append(a_tag)
                    last_pos = end
                
                # 添加剩余未匹配的文本
                if last_pos < len(text):
                    new_nodes.append(NavigableString(text[last_pos:]))
                
                # 反转节点列表(因从后往前处理)并替换原节点
                new_nodes.reverse()
                node.replace_with(*new_nodes)
                break  # 避免同一文本被多次匹配
        else:
            # 跳过a标签,不处理其内部内容
            if node.name != 'a':
                for child in list(node.children):
                    traverse(child)
    
    traverse(soup)
    return soup.prettify(formatter='html')

# 使用示例
if __name__ == '__main__':
    # 示例CSV内容(实际使用时可替换为文件路径,如"links.csv")
    sample_csv = """anchor_text,hyperlink
Google,<a href="https://www.google.com" rel="nofollow noreferrer">https://www.google.com</a>
Bing,<a href="https://bing.com" rel="nofollow noreferrer">https://bing.com</a>
Yahoo,<a href="https://yahoo.com" rel="nofollow noreferrer">https://yahoo.com</a>
Active Campaign,<a href="https://activecampaign.com" rel="nofollow noreferrer">https://activecampaign.com</a>
"""
    
    # 示例HTML内容
    sample_html = """<!-- wp:paragraph -->
<p>This is a existing link <a class="test" href="https://yahoo.com/">Yahoo</a> Text</p>
<!-- /wp:paragraph -->

<!-- wp:paragraph -->
<p>This is another Google Text</p>
<!-- /wp:paragraph -->

<!-- wp:paragraph -->
<p>This is another lowercase bing Text</p>
<!-- /wp:paragraph -->

<!-- wp:paragraph -->
<p>This is another multi word Active Campaign Text</p>
<!-- /wp:paragraph -->
"""
    
    # 加载映射规则并处理HTML
    link_mappings = load_link_mappings(sample_csv)
    processed_html = add_links_to_html(sample_html, link_mappings)
    
    # 输出处理后的HTML
    print(processed_html)

代码关键点说明

  1. CSV处理逻辑:自动提取a标签中的真实URL,按锚文本长度排序确保多词内容优先匹配
  2. 文本匹配规则:使用正则单词边界\b避免部分匹配,配合re.IGNORECASE实现大小写不敏感匹配
  3. 节点遍历与替换:递归跳过<a>标签内部内容,从后往前替换文本避免索引偏移,保留原文本的大小写格式

内容的提问来源于stack exchange,提问作者Abhishek R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 00:50:16