saxpy库find_best_discord_brute_force函数参数及输入要求咨询
find_best_discord_brute_force (saxpy) Nice question! Let's break down each parameter clearly so you can use this function effectively for finding discords in your time series:
series: You don’t need to input SAX words here—this should be your raw time series numerical data (like a Python list or NumPy array of your time-stamped values). The saxpy library handles all the underlying SAX conversion and similarity calculations internally, so just feed in your original dataset directly.win_size: This is the length of the sliding window you want to use to extract sub-sequences from your time series. For example, if you’re working with hourly temperature data and want to find anomalous 24-hour periods, setwin_size=24. All candidate discords will be continuous sub-sequences of this exact length.global_registry: This is a tracking object from saxpy’sregistrymodule (specifically, aGlobalRegistryinstance) that stores distance metrics and candidate discord data during the brute-force search. You’ll need to initialize it first (e.g.,registry = GlobalRegistry()) and pass it in. Its job is to keep track of the minimum distance between each sub-sequence and all others—since discords are the sub-sequences with the largest minimum distance, this registry helps identify that top anomaly.z_threshold(optional, default=0.01): This is a filtering threshold based on Z-score, used to skip unnecessary comparisons with sub-sequences that are too similar to each other. It helps speed up the brute-force search by ignoring trivial, non-anomalous matches.
Quick Example Usage
from saxpy.registry import GlobalRegistry from saxpy.discord import find_best_discord_brute_force # Your raw time series data raw_time_series = [2.1, 2.3, 2.2, 15.7, 2.4, 2.2, ...] # truncated for brevity window_length = 5 registry = GlobalRegistry() # Find the top discord best_discord = find_best_discord_brute_force(raw_time_series, window_length, registry)
内容的提问来源于stack exchange,提问作者lord AJ

