You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从XML文件中查找特定用户回答的对应问题帖?

提取特定用户回答对应的问题帖(XML处理方案)

核心思路

  1. 先定位目标用户的所有回答帖(PostTypeId="2"),收集这些回答对应的ParentId(即问题帖的Id);
  2. 根据收集到的ParentId,提取所有匹配的问题帖(PostTypeId="1")。

实现代码(Python)

针对不同规模的XML文件,提供三种方案:

基础版(适合中小规模XML)

两次遍历XML,逻辑清晰易理解:

import xml.etree.ElementTree as ET

def get_user_answered_questions(xml_path, target_user_id):
    parent_ids = set()
    tree = ET.parse(xml_path)
    root = tree.getroot()

    # 收集目标用户回答的问题Id
    for row in root.findall('.//row'):
        if row.get('PostTypeId') == '2' and row.get('OwnerUserId') == str(target_user_id):
            parent_id = row.get('ParentId')
            if parent_id:
                parent_ids.add(parent_id)
    
    # 提取对应问题帖
    answered_questions = []
    for row in root.findall('.//row'):
        if row.get('PostTypeId') == '1' and row.get('Id') in parent_ids:
            answered_questions.append(row.attrib)
    
    return answered_questions

# 使用示例
if __name__ == '__main__':
    xml_file = "你的XML文件路径.xml"
    target_user = 28
    questions = get_user_answered_questions(xml_file, target_user)
    
    print(f"用户ID {target_user} 回答过的问题:")
    for q in questions:
        print(f"ID: {q['Id']}, 标题: {q.get('Title', '无标题')}")

优化版(一次遍历,效率更高)

仅遍历XML一次,同时收集问题帖和目标用户的回答关联Id,适合中等规模文件:

import xml.etree.ElementTree as ET

def get_user_answered_questions_optimized(xml_path, target_user_id):
    questions_dict = {}
    parent_ids = set()
    
    tree = ET.parse(xml_path)
    root = tree.getroot()
    
    for row in root.findall('.//row'):
        post_type = row.get('PostTypeId')
        post_id = row.get('Id')
        if post_type == '1':
            questions_dict[post_id] = row.attrib
        elif post_type == '2' and row.get('OwnerUserId') == str(target_user_id):
            parent_id = row.get('ParentId')
            if parent_id:
                parent_ids.add(parent_id)
    
    # 筛选匹配的问题帖
    answered_questions = [questions_dict[pid] for pid in parent_ids if pid in questions_dict]
    return answered_questions

# 使用示例
if __name__ == '__main__':
    xml_file = "你的XML文件路径.xml"
    target_user = 28
    questions = get_user_answered_questions_optimized(xml_file, target_user)
    
    for q in questions:
        print(f"ID: {q['Id']}, 标题: {q.get('Title')}")

超大XML适配版(迭代解析,低内存占用)

采用迭代解析方式,无需加载整个XML到内存,支持万行以上的超大型文件:

import xml.etree.ElementTree as ET

def get_user_answered_questions_large_xml(xml_path, target_user_id):
    questions_dict = {}
    parent_ids = set()
    
    # 迭代处理XML元素,避免内存溢出
    for event, elem in ET.iterparse(xml_path, events=('end',)):
        if elem.tag == 'row':
            post_type = elem.get('PostTypeId')
            post_id = elem.get('Id')
            if post_type == '1':
                questions_dict[post_id] = elem.attrib.copy()
            elif post_type == '2' and elem.get('OwnerUserId') == str(target_user_id):
                parent_id = elem.get('ParentId')
                if parent_id:
                    parent_ids.add(parent_id)
            # 清理已处理元素,释放内存
            elem.clear()
    
    answered_questions = [questions_dict[pid] for pid in parent_ids if pid in questions_dict]
    return answered_questions

# 使用示例
if __name__ == '__main__':
    xml_file = "超大XML文件路径.xml"
    target_user = 28
    questions = get_user_answered_questions_large_xml(xml_file, target_user)
    
    for q in questions:
        print(f"ID: {q['Id']}, 标题: {q.get('Title')}")

注意事项

  • 确保XML中的属性值(如OwnerUserId、ParentId)都是字符串类型,代码中已做类型转换适配;
  • 若XML使用命名空间,需在findall或XPath中加入命名空间前缀;
  • 超大文件场景优先选择迭代解析方案,避免内存不足问题。

内容的提问来源于stack exchange,提问作者breath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 09:00:39