如何用BeautifulSoup从Goodreads的quoteText中分离名言与作者?
解决Goodreads名言与作者分离提取的问题
嘿,我来帮你搞定这个提取名言的问题!你之前用extract()时出错,是因为这个方法根本不接受class_参数哦——它是用来移除已经定位到的元素的,不能直接传筛选条件。下面给你几个靠谱的实现方法:
方法1:先移除作者标签,再提取名言
这是最直接的思路:先找到作者的<a>标签,把它从父节点里移除,剩下的文本就是名言了。代码示例:
from bs4 import BeautifulSoup # 你的HTML片段 html = '''<div class="quoteText"> “Don't cry because it's over, smile because it happened.” <br/> ― <a class="authorOrTitle" href="/author/show/61105.Dr_Seuss">Dr. Seuss</a> </div>''' soup = BeautifulSoup(html, 'html.parser') # 定位到quoteText容器 quote_container = soup.find(class_="quoteText") # 提取作者名 author_tag = quote_container.find(class_="authorOrTitle") author_name = author_tag.get_text(strip=True) # 移除作者标签(这一步是关键!) author_tag.extract() # 提取并清理名言文本 quote_content = quote_container.get_text(strip=True).strip('“”') print(f"名言:{quote_content}") print(f"作者:{author_name}")
这里的strip('“”')是为了去掉名言前后的引号,让结果更整洁,你可以根据实际情况调整。
方法2:直接筛选文本节点
如果不想修改DOM结构,也可以直接遍历quoteText下的子节点,只保留纯文本内容:
from bs4 import BeautifulSoup, NavigableString soup = BeautifulSoup(html, 'html.parser') quote_container = soup.find(class_="quoteText") # 提取作者名 author_name = quote_container.find(class_="authorOrTitle").get_text(strip=True) # 筛选出所有纯文本节点,拼接后清理 quote_content = ''.join([ node.strip() for node in quote_container.contents if isinstance(node, NavigableString) ]).strip('“”') print(f"名言:{quote_content}") print(f"作者:{author_name}")
这种方法不会改动原有的DOM结构,适合需要保留原始HTML的场景。
方法3:用decompose()删除作者标签
和extract()类似,decompose()也可以删除元素,区别是它不会返回被删除的元素,直接从DOM中移除:
quote_container = soup.find(class_="quoteText") author_tag = quote_container.find(class_="authorOrTitle") author_name = author_tag.get_text(strip=True) # 直接删除作者标签 author_tag.decompose() quote_content = quote_container.get_text(strip=True).strip('“”')
总结一下:extract()是针对单个元素的方法,你得先找到要移除的子元素,再调用它的extract(),而不是给extract()传筛选参数。上面几种方法都能帮你把名言和作者分开提取,选你觉得顺手的就行~
内容的提问来源于stack exchange,提问作者Prratek Ramchandani
相关产品推荐
相关产品推荐

