You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+BeautifulSoup提取XML标签内不含属性的内容?

问题

我正在用Python结合BeautifulSoup处理XML转HTML的工作,通过Python模板传递变量字符串。现有一段XML片段,其中<interview>标签包含多个属性和子元素。我已经通过interview变量获取了完整的带属性的<interview>标签,但现在需要得到仅包含该标签内部子内容的transcript变量(示例见下方)。当前代码里直接把transcript赋值为interview,导致拿到的是整个标签,请问怎么实现需求?

XML片段

<item key="68">
    <interview titleBib="22582" title="Harvey Birdman - Session III" bib="22582-3" interviewer="Phil Ken Sebben" location="Memphis, Tennessee" dateSingleRaw="1997-06-16T12:00:00Z" dateSingleRawTime="Mon, 06/16/1997 - 12:00" dateMultipleRaw="" abstractRaw="<p>Some of the topics discussed include: his childhood; </p>">
        <div class="field field-name-body field-type-text-with-summary field-label-hidden">
            <div class="field-items">
                <div class="field-item even" property="content:encoded">
                    <h3>Sebben:</h3>
                    <p>This is the text of the interview.</p>
                    <h3>Birdman:</h3>
                    <p>That's correct.</p>
                </div>
            </div>
        </div>
    </interview>
</item>

期望的transcript变量内容

<div class="field field-name-body field-type-text-with-summary field-label-hidden">
    <div class="field-items">
        <div class="field-item even" property="content:encoded">
            <h3>Sebben:</h3>
            <p>This is the text of the interview.</p>
            <h3>Birdman:</h3>
            <p>That's correct.</p>
        </div>
    </div>
</div>

当前相关Python代码

interview = get_transcript(bn)
interviewer=""
date=""
location=""
transcript=""
if interview is not None:
    interviewer = interview.get('interviewer')
    date = interview.get('dateSingleRawTime')
    location = interview.get('location')
    transcript = interview
out = htmltemplate.substitute(title=title, interviewer=interviewer, date=date, location=location, 
                                transcript=transcript)

解决方案

要获取<interview>标签内部的子内容,可通过以下两种BeautifulSoup方法实现:

方法1:使用decode_contents()方法

该方法直接返回标签内部所有内容的HTML/XML字符串,完全匹配需求。修改transcript的赋值语句:

transcript = interview.decode_contents()

方法2:拼接子元素的字符串表示

若需更灵活处理子元素,可遍历<interview>的所有子节点并拼接其字符串形式:

transcript = ''.join(str(child) for child in interview.children)

修改后的完整代码

interview = get_transcript(bn)
interviewer=""
date=""
location=""
transcript=""
if interview is not None:
    interviewer = interview.get('interviewer')
    date = interview.get('dateSingleRawTime')
    location = interview.get('location')
    # 获取标签内部子内容
    transcript = interview.decode_contents()
out = htmltemplate.substitute(title=title, interviewer=interviewer, date=date, location=location, 
                                transcript=transcript)

内容的提问来源于stack exchange,提问作者Chip Calhoun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 13:36:20