You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在BeautifulSoup中获取文本节点在原始HTML字符串中的源索引?

如何在BeautifulSoup中获取文本节点在原始HTML字符串中的源索引?

我现在碰到个棘手的问题:怎么获取HTML字符串里文本节点对应的原始字符串索引?

我知道BeautifulSoup里的标签元素有sourceline和sourcepos属性能用来定位源位置,但NavigableString(也就是文本节点)好像没有直接能用的类似属性。

一开始我尝试写了这样的函数:

def get_index(text_node: NavigableString) -> int:
    return text_node.next_element.sourcepos - len(text_node)

但这个方法根本没法完美工作,因为HTML标签的闭合部分长度是完全不确定的。比如运行这个测试:

>>> get_index(BeautifulSoup('<p>hello</p><br>', 'html.parser').find(text=True))
7

结果明显是错的。而且像<p>hello</p >这种合法的HTML写法,会导致结果偏差得更离谱,我用BeautifulSoup现有的工具完全不知道该怎么处理这种情况。

除了BeautifulSoup的方案,我也想知道如果用lxml或者Python自带的html模块,有没有更简单的解决办法。

我期望的正确结果应该是这样的:

>>> get_index(BeautifulSoup('hello', 'html.parser').find(text=True))
0
>>> get_index(BeautifulSoup('<p>hello</p><br>', 'html.parser').find(text=True))
3
>>> get_index(BeautifulSoup('<!-- hi -->hello', 'html.parser').find(text=True))
11
>>> get_index(BeautifulSoup('<p></p ><p >hello<br>there</p>', 'html.parser').find(text=True))
12
>>> get_index(BeautifulSoup('<p></p ><p >hello<br>there</p>', 'html.parser').find_all(string=True)[1])
21

备注:内容来源于stack exchange,提问作者Nils

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 17:53:07