Palantir Code Workbook中Python遍历DataFrame行写入XML文件问题
在Palantir Code Workbook中遍历数据集行并写入XML文件的正确方法
问题说明
我在Palantir Code Workbook里有一个包含fileName和Xml两列的数据集filtered_ds,尝试用Python遍历每行,将Xml列的内容写入对应fileName的独立文件,但当前代码遍历的是列而非行,需要修正实现方式。
尝试的错误代码:
def write_xmls(filtered_ds): output = Transforms.get_output() output_fs = output.filesystem() for row in filtered_ds: with output_fs.open(str(row[1]), 'w') as f: f.write(str(row[2])) f.close()
解决方案
错误原因
直接遍历filtered_ds时,Palantir的Dataset对象默认迭代的是列而非行,导致代码逻辑错误。此外,手动调用f.close()属于冗余操作——with语句会自动管理文件的打开与关闭。
修正代码(针对Palantir Dataset对象)
使用rows()方法获取行迭代器,同时通过列名访问数据(比索引更直观,避免位置出错):
def write_xmls(filtered_ds): output = Transforms.get_output() output_fs = output.filesystem() # 遍历数据集的每一行 for row in filtered_ds.rows(): file_name = str(row["fileName"]) xml_content = str(row["Xml"]) with output_fs.open(file_name, 'w') as f: f.write(xml_content)
备选方案(若数据集已转为Pandas DataFrame)
如果filtered_ds是Pandas DataFrame,使用iterrows()方法遍历行:
def write_xmls(filtered_ds): output = Transforms.get_output() output_fs = output.filesystem() # iterrows()返回索引与行数据的元组 for _, row in filtered_ds.iterrows(): file_name = str(row["fileName"]) xml_content = str(row["Xml"]) with output_fs.open(file_name, 'w') as f: f.write(xml_content)
内容的提问来源于stack exchange,提问作者asb
相关产品推荐
相关产品推荐

