You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup匹配<td>标签中的子字符串以筛选目标URL

解决BeautifulSoup子字符串匹配标签的问题

没问题,我来帮你搞定这个子字符串匹配的问题!你当前代码里的text=the_word是精确匹配——只有当的文本和the_word完全一模一样时才会被找到,这就是为什么子字符串匹配失败的原因。

下面给你两种简单有效的解决方案:

方法1:使用Lambda函数判断

Lambda函数可以自定义匹配逻辑,我们只需要检查the_word是否是文本的子字符串即可。注意要先判断文本不为None,避免报错:

import requests
from bs4 import BeautifulSoup

the_word = "你要找的子字符串"  # 替换成你的目标子串
urls = ["url1", "url2", "url3"]  # 你的URL列表

for url in urls:
    r = requests.get(url, allow_redirects=False)
    soup = BeautifulSoup(r.content, 'lxml')
    # 用lambda函数判断子字符串是否存在
    words = soup.find_all("td", text=lambda text: text and the_word in text)
    
    if words:
        print(f"找到匹配项: {words}")
        print(f"对应URL: {url}\n")

方法2:使用正则表达式

如果需要更灵活的匹配(比如忽略大小写、匹配特殊格式等),正则表达式是更好的选择:

import requests
from bs4 import BeautifulSoup
import re  # 别忘了导入正则模块

the_word = "你要找的子字符串"
urls = ["url1", "url2", "url3"]

for url in urls:
    r = requests.get(url, allow_redirects=False)
    soup = BeautifulSoup(r.content, 'lxml')
    # 用正则表达式匹配子字符串
    words = soup.find_all("td", text=re.compile(the_word))
    # 如果需要忽略大小写,改成:re.compile(the_word, re.IGNORECASE)
    
    if words:
        print(f"找到匹配项: {words}")
        print(f"对应URL: {url}\n")

两种方法都能实现子字符串匹配,你可以根据自己的需求选择:

  • Lambda函数适合简单的包含判断,逻辑直观;
  • 正则表达式适合复杂匹配场景,比如模糊匹配、多条件匹配等。

内容的提问来源于stack exchange,提问作者subtleseeker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:07:03