如何用Python+BeautifulSoup提取XML标签内不含属性的内容?
问题
我正在用Python结合BeautifulSoup处理XML转HTML的工作,通过Python模板传递变量字符串。现有一段XML片段,其中<interview>标签包含多个属性和子元素。我已经通过interview变量获取了完整的带属性的<interview>标签,但现在需要得到仅包含该标签内部子内容的transcript变量(示例见下方)。当前代码里直接把transcript赋值为interview,导致拿到的是整个标签,请问怎么实现需求?
XML片段
<item key="68"> <interview titleBib="22582" title="Harvey Birdman - Session III" bib="22582-3" interviewer="Phil Ken Sebben" location="Memphis, Tennessee" dateSingleRaw="1997-06-16T12:00:00Z" dateSingleRawTime="Mon, 06/16/1997 - 12:00" dateMultipleRaw="" abstractRaw="<p>Some of the topics discussed include: his childhood; </p>"> <div class="field field-name-body field-type-text-with-summary field-label-hidden"> <div class="field-items"> <div class="field-item even" property="content:encoded"> <h3>Sebben:</h3> <p>This is the text of the interview.</p> <h3>Birdman:</h3> <p>That's correct.</p> </div> </div> </div> </interview> </item>
期望的transcript变量内容
<div class="field field-name-body field-type-text-with-summary field-label-hidden"> <div class="field-items"> <div class="field-item even" property="content:encoded"> <h3>Sebben:</h3> <p>This is the text of the interview.</p> <h3>Birdman:</h3> <p>That's correct.</p> </div> </div> </div>
当前相关Python代码
interview = get_transcript(bn) interviewer="" date="" location="" transcript="" if interview is not None: interviewer = interview.get('interviewer') date = interview.get('dateSingleRawTime') location = interview.get('location') transcript = interview out = htmltemplate.substitute(title=title, interviewer=interviewer, date=date, location=location, transcript=transcript)
解决方案
要获取<interview>标签内部的子内容,可通过以下两种BeautifulSoup方法实现:
方法1:使用decode_contents()方法
该方法直接返回标签内部所有内容的HTML/XML字符串,完全匹配需求。修改transcript的赋值语句:
transcript = interview.decode_contents()
方法2:拼接子元素的字符串表示
若需更灵活处理子元素,可遍历<interview>的所有子节点并拼接其字符串形式:
transcript = ''.join(str(child) for child in interview.children)
修改后的完整代码
interview = get_transcript(bn) interviewer="" date="" location="" transcript="" if interview is not None: interviewer = interview.get('interviewer') date = interview.get('dateSingleRawTime') location = interview.get('location') # 获取标签内部子内容 transcript = interview.decode_contents() out = htmltemplate.substitute(title=title, interviewer=interviewer, date=date, location=location, transcript=transcript)
内容的提问来源于stack exchange,提问作者Chip Calhoun
相关产品推荐
相关产品推荐

