HIVE字符串路径清理:移除特定字符间非纯字母内容
Clean Paths to Retain Only Letter-Only Segments
Got it, let's tackle this path-cleaning task head-on. The core ask is simple: we need to strip out any path segments wrapped in / that contain non-alphabetic characters (like numbers, hyphens, underscores), and keep only the segments made up entirely of letters.
Approach Breakdown
Here's a step-by-step way to implement this logic:
- Split the input path into individual segments using
/as the separator. - Filter out any segments that aren't 100% alphabetic (a-z, A-Z only).
- Reconstruct the cleaned path, making sure to preserve the leading
/and match the trailing slash behavior from your examples.
Example Implementation (Python)
def clean_target_path(input_path): # Split path into segments, skip empty strings from leading/trailing slashes segments = [seg for seg in input_path.split('/') if seg] # Keep only segments with pure letters valid_segments = [seg for seg in segments if seg.isalpha()] # Rebuild the path with leading slash, add trailing slash if needed (matches your examples) cleaned_path = '/' + '/'.join(valid_segments) # Add trailing slash only if there are valid segments and the original had a non-letter segment at the end if valid_segments and not input_path.endswith(valid_segments[-1]): cleaned_path += '/' return cleaned_path # Test with your sample paths sample_paths = [ "/keywordOneA/keywordTwoA/393r-mr49-j5n65_9e8e77g77b8", "/keywordOneA/keywordTwoA/111-4444-jjjj_1b1b1b1b1b1b1b", "/keywordOneA/keywordTwoB/393r-mr49-j5n65_9e8e77g77b8/keywordThreeA" ] for idx, path in enumerate(sample_paths, 1): print(f"字符串{idx}处理后: {clean_target_path(path)}")
Output Result
字符串1处理后: /keywordOneA/keywordTwoA/ 字符串2处理后: /keywordOneA/keywordTwoA/ 字符串3处理后: /keywordOneA/keywordTwoB/keywordThreeA
How This Works
split('/')breaks the path into parts (e.g.,["", "keywordOneA", "keywordTwoA", "393r-mr49-j5n65_9e8e77g77b8"]for the first sample).- We filter out empty strings to avoid extra slashes in the final output.
seg.isalpha()checks if a segment has only letters—this automatically rejects any segment with numbers, symbols, or mixed characters.- The trailing slash logic ensures we match your expected output: if the original path ended with a non-letter segment, we add a trailing slash to the cleaned path; if it ended with a valid letter segment, we skip the trailing slash.
If you're using a different language, the core logic translates directly. For example, in JavaScript, you'd use split('/'), filter(seg => /^[a-zA-Z]+$/.test(seg)), then join('/') with similar leading/trailing slash handling.
内容的提问来源于stack exchange,提问作者Developer
相关产品推荐
相关产品推荐

