如何用Python Requests、BeautifulSoup提取HTML中<br>后的aaa、bbb内容
解决提取内
后文本为空的问题
后文本为空的问题
问题根源:你写的
soup.find_all('td', {'colspan': '2', 'strong': True})是在找带有strong属性的<td>标签,但实际目标<td>里是包含<strong>子标签,不是自身有这个属性,所以这一步根本没匹配到任何元素,自然输出为空。另外请求URL缺少http/https协议头,会导致请求失败,这也是潜在问题。修正方案:
- 先筛选所有
colspan="2"的<td>,再过滤出内部包含<strong>子标签的元素; - 提取
<br>标签后的文本,可通过next_sibling直接获取,也能结合正则精准匹配。
- 先筛选所有
修正后的代码(纯BeautifulSoup实现)
import requests from bs4 import BeautifulSoup params = { 'api_key': 'APIKEY', 'custom_cookies': 'PHPSESSID=SESSIONID,domain=DOMAIN.com;', } response = requests.get( url='http://www.example.com', # 补全http/https协议头 params=params, timeout=120, ) soup = BeautifulSoup(response.content, 'html.parser') # 先找colspan=2的td,再过滤出包含strong子标签的元素 results = [td for td in soup.find_all('td', {'colspan': '2'}) if td.find('strong')] for result in results: br_tag = result.find('br') if br_tag: text_content = br_tag.next_sibling.strip() print(text_content)
结合正则的实现方案
import requests import re from bs4 import BeautifulSoup params = { 'api_key': 'APIKEY', 'custom_cookies': 'PHPSESSID=SESSIONID,domain=DOMAIN.com;', } response = requests.get( url='http://www.example.com', params=params, timeout=120, ) soup = BeautifulSoup(response.content, 'html.parser') results = [td for td in soup.find_all('td', {'colspan': '2'}) if td.find('strong')] # 正则匹配<br>到</td>之间的内容 pattern = re.compile(r'<br>(.*?)(?=</td>)', re.DOTALL) for result in results: td_html = str(result) match = pattern.search(td_html) if match: text_content = match.group(1).strip() print(text_content)
内容的提问来源于stack exchange,提问作者ZASE
相关产品推荐
相关产品推荐

