You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AstraDBVectorStore add_documents报错'dict'无page_content属性的解决方法

问题解决:'dict' object has no attribute 'page_content'

错误原因

LangChain的AstraDBVectorStore.add_documents()方法要求传入LangChain Document类的实例列表,而非普通字典列表。你当前构造的是字典,代码尝试通过属性访问page_content(如doc.page_content),但字典只能通过键访问(doc["page_content"]),因此触发异常。

修复步骤

  1. 导入LangChain的Document类
  2. 将字典替换为Document实例

修改后的代码

# 先导入Document类
from langchain_core.documents import Document

def store_embeddings_in_astradb(embeddings,text_chunks, metadata):

    vstore = AstraDBVectorStore(
        collection_name="test",
        embedding=embedding_model,
        token=os.getenv("ASTRA_DB_APPLICATION_TOKEN"),
        api_endpoint=os.getenv("ASTRA_DB_API_ENDPOINT"),
    )
    print("after Vstore")

    # 构造Document实例列表,而非字典
    documents = [
        Document(page_content=chunk, metadata=metadata)
        for chunk in text_chunks
    ]
    for doc in documents:
        print(f"Document structure: {doc}")
    print("after documents")

    # Add documents to AstraDB vector store
    inserted_ids = vstore.add_documents(documents)
    return inserted_ids

# 后续代码保持不变...
pdf_files = ["WhatYouNeedToKnowAboutWOMENSHEALTH.pdf", "Womens-Health-Book.pdf"]

embedding_model = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2")

for pdf_file in pdf_files:
    if not os.path.isfile(pdf_file):
        raise ValueError(f"PDF file '{pdf_file}' not found.")

    print(f"Processing file: {pdf_file}")

    text = extract_text_from_pdf(pdf_file)
    text_chunks = split_text_into_chunks(text)
    embeddings = embed_text_chunks(text_chunks, embedding_model)
    metadata = extract_metadata(pdf_file)

    try:
        inserted_ids = store_embeddings_in_astradb(embeddings,text_chunks, metadata)
        print(f"Inserted {len(inserted_ids)} embeddings from '{pdf_file}' into AstraDB.")
    except Exception as e:
        print(f"Failed to insert embeddings for '{pdf_file}': {e}")

额外说明

  • 如果需要为每个文本块添加独立的元数据(比如页码、chunk索引),可以在循环中动态生成metadata,例如:
    documents = [
        Document(
            page_content=chunk, 
            metadata={**metadata, "chunk_index": i}
        )
        for i, chunk in enumerate(text_chunks)
    ]
    
  • 若你已提前生成了embedding向量,也可以使用vstore.add_embeddings(embeddings, documents)方法,提升效率。

内容的提问来源于stack exchange,提问作者mukul Bedwa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 20:13:16