dbt Python增量模型切换报错:IndentationError问题咨询
dbt Python增量模型切换引发IndentationError问题(BigQuery环境)
问题现象
将原本运行正常的materialized = "table"类型Python模型改为materialized = "incremental"后,未添加if dbt.is_incremental:增量逻辑就触发以下错误:
File "/tmp/d87435d6-edb3-4afa-84e7-04dae648adcf/query_consistency.py", line 5 create or replace table graph-mainnet.internal_metrics.query_consistency__dbt_tmp ^ IndentationError: unexpected indent
原因分析
dbt对Python增量模型的处理逻辑与表模型不同:当设置为增量模式时,dbt会自动生成临时表创建与数据合并的SQL片段,但如果代码中缺少显式的if dbt.is_incremental:分支判断,dbt生成的临时表语句会出现缩进格式冲突,进而触发错误。即使暂时不需要实现真正的增量逻辑,也必须保留该分支结构。
解决方案
1. 添加增量逻辑分支
修改代码,在返回DataFrame前添加if dbt.is_incremental:判断(可先留空或返回全量数据,后续再补充增量过滤逻辑):
import requests import pandas as pd import json import time def model(dbt, session): dbt.config(materialized = "incremental") # ENTER THE SCHEMA TYPE YOU WANT TO GET ALL DATA FOR schema_type = 'dex-amm' # fetch the data from the deployment file response = requests.get('https://raw.githubusercontent.com/messari/subgraphs/master/deployment/deployment.json') subgraphs = response.json() # create query query = '''{ financialsDailySnapshots(orderBy: timestamp, orderDirection: desc, first: 365) { cumulativeVolumeUSD dailyProtocolSideRevenueUSD totalValueLockedUSD cumulativeTotalRevenueUSD dailyTotalRevenueUSD dailyVolumeUSD timestamp } }''' base_url = 'https://api.thegraph.com/subgraphs/name/messari/' data = [] for project in subgraphs: for deployment in subgraphs[project]['deployments']: schema = subgraphs[project]['schema'] status = subgraphs[project]['deployments'][deployment]['status'] if status != 'prod' or schema != schema_type: continue if len(data) >= 4: # check if we've reached the subgraphs limit break try: # need this because not all have hosted-service field slug = subgraphs[project]['deployments'][deployment]['services']['hosted-service']['slug'] except KeyError: print(f"KeyError: unable to extract data from '{slug}' for '{project}'") response = requests.post(base_url + slug, json={'query': query}) time.sleep(1) if response.ok: response_json = response.json() headers = response.headers timestamp_query = headers.get('Date') try: data.append((project, deployment, timestamp_query, response_json['data']['financialsDailySnapshots'])) print(f"Got data for: {slug}") except KeyError: print(f"KeyError: unable to extract data from '{slug}' for '{project}'") else: print(f'Request failed for {base_url + slug} with status {response.status_code}') if len(data) >= 4: # check if we've reached the subgraphs limit break # create dataframe df = pd.DataFrame(data, columns=['project', 'deployment', 'timestamp_query', 'data']) # adjust df df = df.explode('data') df = pd.concat([df.drop(['data'], axis=1), df['data'].apply(pd.Series)], axis=1) # convert timestamp df['timestamp_query'] = df['timestamp_query'].drop_duplicates().reset_index(drop=True) df['timestamp_query'] = pd.to_datetime(df['timestamp_query'], format='%a, %d %b %Y %H:%M:%S %Z') # identifier df['product'] = 'hosted_service' # 增量模式判断分支(必须保留) if dbt.is_incremental(): # 后续可补充增量过滤逻辑,例如: # max_timestamp = session.sql("select max(timestamp) from {{ this }}").collect()[0][0] # df = df[df['timestamp'] > max_timestamp] pass return df
2. 验证dbt版本兼容性
Python增量模型在dbt-bigquery 1.3.0及以上版本才正式支持,执行以下命令检查版本:
dbt --version
若版本过低,升级至最新稳定版:
pip install --upgrade dbt-bigquery
3. 确认BigQuery权限
确保dbt使用的服务账号拥有目标数据集的创建临时表和写入现有表权限,避免因权限不足导致的隐式格式错误。
内容的提问来源于stack exchange,提问作者Ricky
相关产品推荐
相关产品推荐

