如何在BeautifulSoup中获取文本节点在原始HTML字符串中的源索引?
如何在BeautifulSoup中获取文本节点在原始HTML字符串中的源索引?
我现在碰到个棘手的问题:怎么获取HTML字符串里文本节点对应的原始字符串索引?
我知道BeautifulSoup里的标签元素有sourceline和sourcepos属性能用来定位源位置,但NavigableString(也就是文本节点)好像没有直接能用的类似属性。
一开始我尝试写了这样的函数:
def get_index(text_node: NavigableString) -> int: return text_node.next_element.sourcepos - len(text_node)
但这个方法根本没法完美工作,因为HTML标签的闭合部分长度是完全不确定的。比如运行这个测试:
>>> get_index(BeautifulSoup('<p>hello</p><br>', 'html.parser').find(text=True)) 7
结果明显是错的。而且像<p>hello</p >这种合法的HTML写法,会导致结果偏差得更离谱,我用BeautifulSoup现有的工具完全不知道该怎么处理这种情况。
除了BeautifulSoup的方案,我也想知道如果用lxml或者Python自带的html模块,有没有更简单的解决办法。
我期望的正确结果应该是这样的:
>>> get_index(BeautifulSoup('hello', 'html.parser').find(text=True)) 0 >>> get_index(BeautifulSoup('<p>hello</p><br>', 'html.parser').find(text=True)) 3 >>> get_index(BeautifulSoup('<!-- hi -->hello', 'html.parser').find(text=True)) 11 >>> get_index(BeautifulSoup('<p></p ><p >hello<br>there</p>', 'html.parser').find(text=True)) 12 >>> get_index(BeautifulSoup('<p></p ><p >hello<br>there</p>', 'html.parser').find_all(string=True)[1]) 21
备注:内容来源于stack exchange,提问作者Nils
相关产品推荐
相关产品推荐

