如何用Python BeautifulSoup提取HTML标签纯文本?解决string返回None问题
问题原因
第二个<td>标签内部包含文本节点、<a>子标签和<br>标签,并非单一的文本节点。BeautifulSoup的.string属性仅当标签是唯一子节点或者仅包含文本节点时才会返回文本内容,否则返回None。
解决方法
使用.get_text()方法(或等价的.text属性)提取标签下所有子节点的文本内容,它会自动拼接所有文本片段:
from bs4 import BeautifulSoup html = ''' <td style="padding-right: 10px;" valign="top">1.1</td> <td valign="top"> If applicable, do <a href="url"> link </a> to switch one to other mode.<br/> </td> ''' soup = BeautifulSoup(html, 'html.parser') # 提取第一个td的文本 print(soup.find_all("td")[0].get_text()) # 提取第二个td的文本并自动去除多余空白 print(soup.find_all("td")[1].get_text(strip=True))
执行结果:
1.1 If applicable, do link to switch one to other mode.
如果需要保留原始的缩进、换行等格式,去掉strip=True参数即可:
print(soup.find_all("td")[1].get_text())
结果会保留原有的空白格式:
If applicable, do link to switch one to other mode.
另外,也可以用.stripped_strings生成器获取所有去除了多余空白的文本片段,再自行拼接:
text = ' '.join(soup.find_all("td")[1].stripped_strings) print(text)
输出同样为:
If applicable, do link to switch one to other mode.
内容的提问来源于stack exchange,提问作者Shockyn
相关产品推荐
相关产品推荐

