如何从XML文件中查找特定用户回答的对应问题帖?
提取特定用户回答对应的问题帖(XML处理方案)
核心思路
- 先定位目标用户的所有回答帖(
PostTypeId="2"),收集这些回答对应的ParentId(即问题帖的Id); - 根据收集到的
ParentId,提取所有匹配的问题帖(PostTypeId="1")。
实现代码(Python)
针对不同规模的XML文件,提供三种方案:
基础版(适合中小规模XML)
两次遍历XML,逻辑清晰易理解:
import xml.etree.ElementTree as ET def get_user_answered_questions(xml_path, target_user_id): parent_ids = set() tree = ET.parse(xml_path) root = tree.getroot() # 收集目标用户回答的问题Id for row in root.findall('.//row'): if row.get('PostTypeId') == '2' and row.get('OwnerUserId') == str(target_user_id): parent_id = row.get('ParentId') if parent_id: parent_ids.add(parent_id) # 提取对应问题帖 answered_questions = [] for row in root.findall('.//row'): if row.get('PostTypeId') == '1' and row.get('Id') in parent_ids: answered_questions.append(row.attrib) return answered_questions # 使用示例 if __name__ == '__main__': xml_file = "你的XML文件路径.xml" target_user = 28 questions = get_user_answered_questions(xml_file, target_user) print(f"用户ID {target_user} 回答过的问题:") for q in questions: print(f"ID: {q['Id']}, 标题: {q.get('Title', '无标题')}")
优化版(一次遍历,效率更高)
仅遍历XML一次,同时收集问题帖和目标用户的回答关联Id,适合中等规模文件:
import xml.etree.ElementTree as ET def get_user_answered_questions_optimized(xml_path, target_user_id): questions_dict = {} parent_ids = set() tree = ET.parse(xml_path) root = tree.getroot() for row in root.findall('.//row'): post_type = row.get('PostTypeId') post_id = row.get('Id') if post_type == '1': questions_dict[post_id] = row.attrib elif post_type == '2' and row.get('OwnerUserId') == str(target_user_id): parent_id = row.get('ParentId') if parent_id: parent_ids.add(parent_id) # 筛选匹配的问题帖 answered_questions = [questions_dict[pid] for pid in parent_ids if pid in questions_dict] return answered_questions # 使用示例 if __name__ == '__main__': xml_file = "你的XML文件路径.xml" target_user = 28 questions = get_user_answered_questions_optimized(xml_file, target_user) for q in questions: print(f"ID: {q['Id']}, 标题: {q.get('Title')}")
超大XML适配版(迭代解析,低内存占用)
采用迭代解析方式,无需加载整个XML到内存,支持万行以上的超大型文件:
import xml.etree.ElementTree as ET def get_user_answered_questions_large_xml(xml_path, target_user_id): questions_dict = {} parent_ids = set() # 迭代处理XML元素,避免内存溢出 for event, elem in ET.iterparse(xml_path, events=('end',)): if elem.tag == 'row': post_type = elem.get('PostTypeId') post_id = elem.get('Id') if post_type == '1': questions_dict[post_id] = elem.attrib.copy() elif post_type == '2' and elem.get('OwnerUserId') == str(target_user_id): parent_id = elem.get('ParentId') if parent_id: parent_ids.add(parent_id) # 清理已处理元素,释放内存 elem.clear() answered_questions = [questions_dict[pid] for pid in parent_ids if pid in questions_dict] return answered_questions # 使用示例 if __name__ == '__main__': xml_file = "超大XML文件路径.xml" target_user = 28 questions = get_user_answered_questions_large_xml(xml_file, target_user) for q in questions: print(f"ID: {q['Id']}, 标题: {q.get('Title')}")
注意事项
- 确保XML中的属性值(如
OwnerUserId、ParentId)都是字符串类型,代码中已做类型转换适配; - 若XML使用命名空间,需在
findall或XPath中加入命名空间前缀; - 超大文件场景优先选择迭代解析方案,避免内存不足问题。
内容的提问来源于stack exchange,提问作者breath
相关产品推荐
相关产品推荐

