如何用XPath拼接艺术家与标题字符串,生成指定格式内容?
Beatport 艺术家与作品标题拼接解决方案
方法1:XPath 2.0+ 直接拼接
如果你的爬虫工具支持XPath 2.0及以上(如Saxon、带XPath2扩展的lxml),可以直接通过定位作品容器,在容器内完成艺术家与标题的拼接,避免跨条目匹配错误:
//div[contains(@class, 'buk-horz-release')]/concat( string-join(p[@class='buk-horz-release-artists']/a, ","), " - ", p[contains(@class,'buk-horz-release-title')]/a/text() )
这个XPath先定位单个作品的父容器,再分别提取该容器下的艺术家(用string-join合并多艺术家)和标题,最后用concat拼接成要求的格式。
方法2:工具层面拼接(兼容XPath 1.0)
若工具仅支持XPath 1.0(如默认配置的lxml、BeautifulSoup),可先分别提取艺术家和标题列表,再按索引一一配对拼接。以Python为例:
from lxml import etree # 假设html为目标HTML内容 tree = etree.HTML(html) # 提取所有艺术家(处理多艺术家情况) artist_list = [",".join(node.xpath("./a/text()")) for node in tree.xpath("//p[@class='buk-horz-release-artists']")] # 提取所有标题 title_list = [node.xpath("./a/text()")[0] for node in tree.xpath("//p[contains(@class,'buk-horz-release-title')]")] # 拼接成目标格式 final_result = [f"{artist} - {title}" for artist, title in zip(artist_list, title_list)] # 输出结果 for item in final_result: print(item)
关键提示:必须保证艺术家和标题列表的顺序完全对应,所以提取时要确保两个XPath的遍历顺序一致,最好基于作品容器的遍历顺序来提取子元素,避免错位。
内容的提问来源于stack exchange,提问作者11ph22il
相关产品推荐
相关产品推荐

