You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取Couchbase指定collection全部文档并保存为DataFrame

实现方案

基于你已经写好的集群连接代码,可通过以下两种常用方式提取全量文档并转为pandas DataFrame:


方法1:使用N1QL查询(通用方案,适合绝大多数中小规模数据集)

# 1. 执行全量查询语句,注意scope和collection名称需和实际一致,需用反引号包裹避免特殊字符报错
query = "SELECT * FROM `B`.`C`"
result = cluster.query(query, QueryOptions())

# 2. 解析查询结果为字典列表
docs = []
for row in result.rows():
    # N1QL查询返回结果默认以collection名作为外层key,需取出内层实际文档内容
    docs.append(row["C"])

# 3. 直接转为pandas DataFrame
df = pd.DataFrame(docs)

# 可选:验证输出
print(df.shape)
print(df.head())

方法2:使用KV范围扫描(性能更优,适合超大规模数据集)

如果你的Couchbase版本为7.0及以上,可使用KV范围扫描直接遍历全文档,跳过N1QL解析开销,性能更高:

from couchbase.options import ScanOptions

# 全量扫描collection所有文档
scan_result = cb_coll.scan(ScanOptions())

# 解析扫描结果
docs = []
for item in scan_result:
    docs.append(item.content_as[dict])

# 转为DataFrame
df = pd.DataFrame(docs)

注意事项

  • 如果你的collection数据量超过10万条,建议增加分页逻辑避免单次请求占用过多内存,分页示例:
    batch_size = 1000
    offset = 0
    docs = []
    while True:
        query = f"SELECT * FROM `B`.`C` LIMIT {batch_size} OFFSET {offset}"
        res = cluster.query(query)
        batch = [row["C"] for row in res.rows()]
        if not batch:
            break
        docs.extend(batch)
        offset += batch_size
    df = pd.DataFrame(docs)
    
  • 执行全量扫描前建议确认集群资源充足,避免影响线上业务运行。

内容的提问来源于stack exchange,提问作者SoKu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 10:48:03