寻求用于测试地理编码脚本的多缺陷全球地址数据集
Looking for a robust dataset to test your geocoding script? Here are some solid options that meet your requirements, plus tips to tailor them to your needs:
1. OpenAddresses Collection
This crowdsourced dataset boasts millions of address records from around the globe, covering every continent except Antarctica. It’s perfect for testing because it includes raw, unpolished entries—think typos, missing components (like incomplete street names or missing zip codes), and inconsistent formats right out of the box. You’ll find multilingual entries too, including Arabic and Chinese addresses. The data is available in CSV format, which you can easily load into a pandas or Dask DataFrame with a simple pd.read_csv() call. If you need more controlled invalid/NoData cases, you can mix in your own manually created faulty entries or filter the raw dataset to emphasize imperfect records.
2. Geonames Geographic Database
With over 10 million records, Geonames is a go-to for global geographic data that includes address-like entries. It covers all inhabited continents, supports dozens of languages, and has plenty of entries that fit your "imperfect" criteria: missing fields, misspelled place names, and ambiguous locations (fuzzy addresses). The data comes as tab-separated files, which are straightforward to import into pandas using pd.read_csv(sep='\t'). You can also use their filtering tools to extract subsets of data that focus on the error types you want to test (like NoData entries or multilingual addresses).
3. Synthetic + Real Data Hybrid Approach
If you want full control over the exact types of imperfections in your dataset, combine real address data with synthetically generated faulty entries. Use Python libraries like Faker to generate fake addresses in multiple languages, then introduce intentional errors: typos, invalid zip codes, missing city/state fields, or nonsensical strings that mimic real-world bad input. Mix these synthetic entries with a real dataset (like the ones above) to hit your 10k+ sample size. This method lets you tailor the dataset exactly to your error-handling test cases.
Quick Tips for Importing & Enhancing
- For pandas: Use
pd.read_csv()with thena_valuesparameter to explicitly mark NoData entries (e.g.,na_values=['', 'N/A', 'Invalid']). - For Dask: Use
dask.dataframe.read_csv()for larger datasets that don’t fit into memory—it works similarly to pandas but handles chunking automatically. - To add more fuzzy/invalid entries: Manually introduce typos, remove address components, or use regex to modify existing records to simulate real-world errors.
内容的提问来源于stack exchange,提问作者Rutger Hofste

