使用Jsoup移除td中的sup元素文本失败,寻求解决方法
解决方案:移除td中的sup元素及引用标记
看起来你是在网页抓取时遇到了td标签里混着sup引用标记的问题,我给你几个靠谱的处理方法,分场景适用:
方法1:先移除sup元素再提取文本(最推荐,结构化处理)
如果你用的是Python的BeautifulSoup,直接从DOM结构里删掉不需要的sup标签,再取文本就干净了:
from bs4 import BeautifulSoup # 假设你已经拿到了目标td的HTML内容 td_html = """<td> L.A. Confidential (film) Curtis Hanson 138 minutes L.A. Story Mick Jackson 98 minutes<sup>[1]</sup> L.I.E. Michael Cuesta 97 minutes<sup>[1]</sup> L.O.R.D: Legend of Ravaging Dynasties Guo Jingming 117 minutes<sup>[2]</sup> </td>""" soup = BeautifulSoup(td_html, 'html.parser') target_td = soup.find('td') # 遍历删除所有sup标签 for sup_tag in target_td.find_all('sup'): sup_tag.decompose() # 提取整理后的文本 clean_text = target_td.get_text(strip=True, separator=' ') print(clean_text)
运行后会输出:
L.A. Confidential (film) Curtis Hanson 138 minutes L.A. Story Mick Jackson 98 minutes L.I.E. Michael Cuesta 97 minutes L.O.R.D: Legend of Ravaging Dynasties Guo Jingming 117 minutes
方法2:遍历子节点,跳过sup元素取文本
如果不想修改原DOM结构,可以逐个检查td的子节点,只收集非sup元素的内容:
clean_content = [] for child in target_td.contents: # 跳过sup标签 if child.name != 'sup': # 处理文本节点和其他元素的文本 content = child.strip() if isinstance(child, str) else child.get_text(strip=True) if content: clean_content.append(content) clean_text = ' '.join(clean_content) print(clean_text)
备选方案:正则替换(适合已拿到带标记的文本后处理)
如果你已经得到了带[1]这类标记的原始文本,也可以用正则直接替换掉这些引用标记:
import re raw_text = "L.A. Confidential (film) Curtis Hanson 138 minutes L.A. Story Mick Jackson 98 minutes[1] L.I.E. Michael Cuesta 97 minutes[1] L.O.R.D: Legend of Ravaging Dynasties Guo Jingming 117 minutes[2]" # 替换掉所有[数字]格式的标记 clean_text = re.sub(r'\[\d+\]', '', raw_text) # 清理多余的空格 clean_text = re.sub(r'\s+', ' ', clean_text).strip() print(clean_text)
为啥之前的方法失效?
如果之前你直接用td.get_text(),它会把所有子元素(包括sup标签里的文本)全部合并,所以才会带出[1]这类标记。上面的方法要么提前移除sup元素,要么跳过sup的内容,自然就能得到干净的结果了。
内容的提问来源于stack exchange,提问作者webscrapingtech
相关产品推荐
相关产品推荐

