You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Requests、BeautifulSoup提取HTML中<br>后的aaa、bbb内容

解决提取内
后文本为空的问题
  • 问题根源:你写的soup.find_all('td', {'colspan': '2', 'strong': True})是在找带有strong属性的<td>标签,但实际目标<td>里是包含<strong>子标签,不是自身有这个属性,所以这一步根本没匹配到任何元素,自然输出为空。另外请求URL缺少http/https协议头,会导致请求失败,这也是潜在问题。

  • 修正方案:

    1. 先筛选所有colspan="2"的<td>,再过滤出内部包含<strong>子标签的元素;
    2. 提取<br>标签后的文本,可通过next_sibling直接获取,也能结合正则精准匹配。

修正后的代码(纯BeautifulSoup实现)

import requests
from bs4 import BeautifulSoup

params = {
    'api_key': 'APIKEY', 
    'custom_cookies': 'PHPSESSID=SESSIONID,domain=DOMAIN.com;',
}

response = requests.get(
    url='http://www.example.com',  # 补全http/https协议头
    params=params,
    timeout=120,
)

soup = BeautifulSoup(response.content, 'html.parser')
# 先找colspan=2的td,再过滤出包含strong子标签的元素
results = [td for td in soup.find_all('td', {'colspan': '2'}) if td.find('strong')]

for result in results:
    br_tag = result.find('br')
    if br_tag:
        text_content = br_tag.next_sibling.strip()
        print(text_content)

结合正则的实现方案

import requests
import re
from bs4 import BeautifulSoup

params = {
    'api_key': 'APIKEY', 
    'custom_cookies': 'PHPSESSID=SESSIONID,domain=DOMAIN.com;',
}

response = requests.get(
    url='http://www.example.com',
    params=params,
    timeout=120,
)

soup = BeautifulSoup(response.content, 'html.parser')
results = [td for td in soup.find_all('td', {'colspan': '2'}) if td.find('strong')]

# 正则匹配<br>到</td>之间的内容
pattern = re.compile(r'<br>(.*?)(?=</td>)', re.DOTALL)
for result in results:
    td_html = str(result)
    match = pattern.search(td_html)
    if match:
        text_content = match.group(1).strip()
        print(text_content)

内容的提问来源于stack exchange,提问作者ZASE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 16:20:20